A new benchmark tested AI agents on 130 long-horizon tasks in healthcare and finance, and even its strongest tested model passed fewer than one in four attempts. The researchers say a recurring failure was accepting a wrong premise in source material, then carrying it through otherwise consistent work.
Watch Desk analysis
What happened
The DAYJOB benchmark covers professional tasks involving messy files and strict rubrics. The researchers evaluated 30 model configurations from 13 developers. In their results, Claude Opus 5.5 had the highest pass rates: 24.7% on healthcare tasks and 23.9% on finance tasks. The median configuration performed worse.
The authors also describe agents accepting incorrect premises in source documents and using those mistaken inputs throughout their analyses. Read the DAYJOB paper on arXiv.
Why it matters
These tasks test more than whether a model can produce a plausible answer. They ask whether an agent can carry work through several steps while handling imperfect files and following detailed requirements. A tidy analysis built on a false starting point is still wrong, just with better formatting.
The results are specific to this benchmark and its tested configurations, not a pass rate for every professional use of AI. But they give organisations a concrete reason to test agents on realistic, messy work before handing them long tasks with consequential outputs.
Our read
DAYJOB usefully shifts attention from impressive-looking demonstrations to whether an agent can complete a whole job reliably. The low pass rates are a clear warning against treating fluent, step-by-step work as proof that the underlying premises are sound. If you are evaluating an agent, include deliberately messy inputs and check whether it catches a bad assumption rather than simply building on it.
What to watch
- Whether later evaluations reproduce the results across more models and professional tasks.
- How performance changes when agents are explicitly asked to challenge source assumptions.
- Whether real deployments can detect and correct wrong inputs before they shape downstream work.
Discussion spark: For professional AI agents, should passing realistic end-to-end tasks matter more than strong scores on narrower benchmarks, even if the results are harder to compare?
Sources and evidence
- DAYJOB: A Benchmark for Long-Horizon Professional Work (2 October 2026, 01:58 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.