SWE-sweep puts coding agents to a less comfortable test: finding and fixing multiple bugs in large codebases without hints about their type or location. The benchmark’s reported results suggest that autonomous bug-hunting remains a long way from routine, even as coding agents grow more capable.
Watch Desk analysis
What happened
Researchers introduced SWE-sweep, a benchmark built from 4,000 real bugs across 100 repositories and 22 programming languages. Its tasks draw on real issue and pull-request pairs where multiple bugs coexist in a single code commit, according to the SWE-sweep project site.
The benchmark asks agents to discover and repair those bugs without being told where to look or what kinds of bugs to expect. The supplied account says the best-performing models and agents resolved less than 5% of the bugs. OpenAI’s Sol 5.6 scored 4.7%, with significant computational costs, the account says.
Key findings
- A broad test of autonomous debugging
SWE-sweep covers 4,000 bugs in 100 repositories and 22 programming languages. - Little success without clues
The reported top results were below 5%; Sol 5.6 reached 4.7% and incurred significant computational costs. - The bugs come in combinations
Tasks use real cases where multiple bugs coexist in a code commit, rather than pointing an agent towards one known defect.
Why it matters
Many coding-agent demonstrations start with a well-described task. SWE-sweep instead tests whether an agent can work out what is broken in the first place, then make repairs without causing more trouble. That is closer to the unglamorous reality of maintaining a large codebase.
The reported results put a useful boundary around claims of autonomous software maintenance: on this benchmark, even the leading systems rarely solved the tasks. That does not settle what agents can do with human guidance, nor how they perform on other kinds of work. It does make “find and fix the bugs yourself” a harder claim to wave through on the strength of a polished demo.
Our read
SWE-sweep is worth watching because it tests a real gap between editing code on request and independently diagnosing a messy repository. Treat its scores as results on this benchmark, not a universal ranking of coding assistants. The practical question for teams is whether an agent can make a repair that survives the repository’s tests and constraints, not merely produce a plausible patch.
What to watch
- Whether future systems improve on the benchmark’s less-than-5% results.
- How performance compares across models, coding agents and computational costs.
- Whether benchmark results translate into dependable bug-fixing in everyday development.
Discussion spark: Should coding agents be judged mainly on whether they can find and fix bugs without hints, or on how well they kernel-level A-bomb specialist developers once a problem is identified?
Sources and evidence
- Researchers introduce SWE-sweep benchmark () (1 October 2026, 16:15 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.