FrontierSWE’s creators have released version 2 of their benchmark, expanding it to 34 technical challenges and introducing evaluations that run for up to 20 hours. They say Claude Fable 5.1 leads the results, ahead of GPT-5.6 and GLM-5.3, a ranking that puts long-horizon AI coding ability under fresh scrutiny.
Watch Desk analysis
What happened
The FrontierSWE v2 announcement adds 21 challenges to the benchmark, including tasks in scientific computing and AI research. Its creators say the new version uses a coding-agent harness called proximus to measure how far models progress during a 20-hour evaluation. The FrontierSWE v2 announcement gives the benchmark’s account of its changes and results.
The creators report that Claude Fable 5.1 performed best, ahead of GPT-5.6 and GLM-5.3. The supplied announcement does not give the individual scores or further detail on how the models were run, so the ranking is a reported result rather than a complete basis for comparing performance.
Why it matters
A longer evaluation can test work that does not fit neatly into a short coding prompt: models have more time to make progress on demanding, specialised tasks. Expanding the challenge set also gives evaluators more ground on which to compare systems. But a leaderboard is only as informative as its tasks, scoring and test conditions; the headline order does not settle which model is best for everyday software work.
Our read
The notable change is the attempt to measure sustained progress, not just whether a model can ace a brief coding test. FrontierSWE v2 offers a useful new point of comparison, and its reported winner is worth watching. For now, treat the ranking as an early signal. Scores, evaluation details and repeatable results would make it much easier to tell what the gap means in practice.
What to watch
- The scores and per-task results for Claude Fable 5.1, GPT-5.6 and GLM-5.3.
- More detail on the proximus harness, scoring and evaluation conditions.
- Whether other evaluations find a similar ordering on long-running technical work.
Discussion spark: Should AI coding benchmarks prioritise long, demanding tasks, or tests that better reflect the everyday work most developers actually do?
Sources and evidence
- FrontierSWE v2 (29 September 2026, 14:14 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.