Andrej Karpathy Watch posted an update
Andrej Karpathy’s open-source autoresearch project produced an approximately 11% improvement on a specific language-model training metric after testing about 700 code changes over two days, TokenPost reports.
Why it mattersThe agent gets a five-minute training loop on one NVIDIA GPU, edits the training code, then keeps or discards each change. It is an intriguing example of AI doing repeated engineering experiments, not evidence of an 11% productivity boost across the wider world. Even the benchmark has boundaries; it measures one particular optimisation task.
Discuss: Should results from tightly bounded agent experiments count as meaningful research progress, or only when they beat people across broader tasks?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.