A study from Apple and EPFL found that, under equal time budgets, elaborate agent harnesses did not outperform a minimal coding-agent setup using the same frontier model. The result puts a useful question to AI developers: how much of an agent’s performance comes from its scaffolding, and how much from the model doing the work?
Watch Desk analysis
What happened
The researchers compared open-source, state-of-the-art harnesses with a minimal-harness coding-agent baseline for autonomous machine-learning engineering. They report no performance advantage for the more elaborate setups in that comparison, and argue that the underlying frontier model was the primary driver.
Their research summary describes systematic ablation studies. The finding is about the tested agents and task setting, not proof that harness design never matters.
Why it matters
Agent harnesses can add planning, tools and other layers around a model. If those additions do not improve results in a given task, they may bring complexity without a matching payoff. For teams building coding agents, the comparison makes a simple baseline worth keeping: otherwise, it is hard to tell whether another layer helped or merely made the diagram busier.
Our read
This is a useful challenge to the assumption that more agent machinery automatically means a better agent. The result does not settle which harness works best across other tasks, models or budgets, but it gives developers a reason to test each extra component rather than treating sophistication as a score.
What to watch
- Whether other teams reproduce the comparison with different models and coding tasks.
- Which harness components help in settings beyond the evaluated ML-engineering work.
- Whether longer or differently allocated time budgets change the result.
Discussion spark: When building coding agents, should teams start with the simplest possible harness and add complexity only when tests justify it, or are there important capabilities a minimal baseline is likely to miss?
Sources and evidence
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? (1 October 2026, 14:20 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.