Six AI models were given real tasks from Epoch AI, and even the strongest struggled when the work called for open-ended judgement rather than a tidy, checkable answer. The report offers a more grounded test of workplace automation than another leaderboard score: can models produce work to the standard people actually use?
Epoch AI Watch analysis
What happened
Epoch AI says its new evaluation gave six models tasks drawn from its own work, including research design and graphic generation. The models received relevant context, used their highest available reasoning settings and attempted tasks without further intervention. People then graded the outputs against Epoch’s employee standards.
The initial task suite contains 11 tasks across five categories. Epoch says GPT-6 Astra and Claude Fable 5.1 broadly led, but neither could yet automate its work. The researchers found that models could suggest promising research directions while struggling to design informative experiments, and sometimes treated results from flawed setups as key findings. They also often converged on similar ideas, despite the wider space of possible approaches. Read Epoch AI’s report.
Key findings
- Strongest models still fall short of the job
GPT-6 Astra and Claude Fable 5.1 led overall, but did not reliably produce Epoch-quality work autonomously. - Open-ended research exposes a judgement gap
Models could suggest promising directions, yet struggled to design experiments that measured what they claimed. - Flawed setups could yield confident-looking findings
Epoch says models sometimes presented results driven by weaknesses in their own experimental design as key findings. - Open-weight models lagged on defined tasks too
They struggled on some well-specified work that frontier closed-weight models handled more reliably. - Models missed house standards
Even with reference material, they could miss implicit expectations such as visual style and the topics Epoch’s audience values.
Why it matters
Many familiar benchmarks reward answers that are straightforward to check. Epoch argues that real work also involves gathering context, using tools and making judgements that do not reduce neatly to a right-or-wrong test. Its evaluations could help expose gaps that polished benchmark scores leave out.
This is one organisation testing models against its own work, not a universal verdict on every job or workplace. Still, the finding is useful: capability on a well-defined task does not automatically translate into reliable end-to-end work.
Our read
This is the kind of test AI needs more of: less “can it ace the quiz?” and more “does the result hold up when the task gets messy?” Epoch’s employee-standard grading is a practical lens, though readers should keep in mind that the tasks and standards come from Epoch itself.
For now, the report makes a strong case for treating model output as a contribution to research work, not a substitute for checking whether the experiment actually supports the conclusion.
What to watch
- Whether Epoch expands the 11-task suite and publishes results as models change.
- How the models perform on the full range of tasks and categories.
- Whether other organisations test automation against their own standards and share comparable results.
Discussion spark: Should workplace automation be judged against an organisation’s own standards, or does that make the test too specific to compare across employers?
Sources and evidence
- Largest AI data center: doubling every 7 months | Epoch AI (Publication date not supplied)
Independent WittyWires tracker for public updates about Epoch AI. Not affiliated with or endorsed by Epoch AI; this is not an official account.