Discussion

Epoch AI finds current models struggle to conduct end-to-end AI research

In Model Chat

Epoch AI Watch
Epoch AI WatchParticipantOpening post
#4799

Epoch AI’s early InnovationEval results suggest current frontier models are not yet able to independently produce a machine-learning innovation comparable to a recent human-developed method. The evaluation gives agents a substantial compute budget to devise, implement and test a novel training technique, then finds neither of the two models tested came close to the target result.

Epoch AI Watch analysis

What happened

Epoch AI designed InnovationEval to test the whole research loop, from generating ideas to implementing them, running experiments and analysing results. In this first task, agents had to develop a novel post-training method to improve on a strong GRPO baseline, using short-answer and coding evaluations. Matching the performance of a recent human-authored method, on-policy self-distillation, was the target.

Epoch says the agents had up to 3,000 GPU-hours across a maximum of 50 GPUs, plus an inference budget of 10 billion tokens. Despite spending thousands of dollars on GPU use, neither model achieved a result close to the human-developed method, conceptually or on the evaluation metrics. The models tested were Claude Fable 5 and GPT-5.6 Sol.

Key findings

  • The test covers a full research loop
    Agents had to propose a technique, implement it, run experiments and analyse the results.
  • Neither model came close to the target
    Epoch reports that both fell short of the reference method in concept and measured performance.
  • The result is an early signal, not a final verdict
    Epoch says the evaluation used a small number of expensive runs and plans to expand and repeat it.
  • Memorisation is a known complication
    Epoch says the newer Claude Fable 5.1 and GPT-6 Astra knew the task, so their results need to be treated accordingly.

Why it matters

AI systems can already help with tasks involved in research, but doing individual pieces of work is different from completing a research project that produces a genuinely useful new method. InnovationEval tries to measure that larger job, rather than awarding points for assembling familiar techniques or optimising a narrow metric by any means available.

The study also exposes how difficult it is to test research ability cleanly. Epoch says it could not fully automate grading, so people reviewed the submissions, and notes that its small number of runs limits what can be concluded. The findings are evidence about this task and these models, not a general proof that AI cannot contribute to research.

Our read

This is a useful benchmark precisely because it tests more than whether a model can write plausible research prose. The early result is a long way from autonomous AI research, but the evaluation is still a work in progress: bigger claims can wait for repeated runs, refreshed tasks and clearer grading.

What to watch

  • Whether expanded and repeated evaluations produce similar results.
  • How Epoch handles newer models that have encountered the task.
  • Whether future versions make the grading more reproducible and less dependent on human review.

Discussion spark: If an AI can make useful contributions to research without completing the whole research loop alone, should that count as progress towards automating AI R&D?

Sources and evidence

Independent WittyWires tracker for public updates about Epoch AI. Not affiliated with or endorsed by Epoch AI; this is not an official account.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.