MIT and Sakana AI researchers have developed SIFT, a framework that uses an AI judge to help select promising changes to coding agents before spending heavily on benchmark tests. In reported experiments, it improved results on several coding benchmarks while reducing some evaluation costs, offering a practical route to explore more agent designs without running every candidate through the full test suite.
Watch Desk analysis
What happened
SIFT, short for Recursive Self-Improvement via Fast Tree Search, has an agent propose changes to its own code or tools. It first runs a small check, then uses a language model to compare the modified agent with promising candidates in an archive. Those judgements help prioritise which versions receive more expensive evaluations. The stages can run in parallel, so the search need not wait for each full test to finish before exploring another branch.
VentureBeat reports that, on the Polyglot coding benchmark, a SIFT run using o3-mini reached 35.1% accuracy in under five hours, using 42 CPU hours and about $150 in API credits. The report says that result exceeded DGM’s 30.7% in the comparison. On TerminalBench, judge-selected candidates averaged 36.7% across repeated full-benchmark runs, compared with 28.1% for the candidate that scored highest on the smaller search set. The researchers also tested SIFT on a 60-task SWE-bench Verified subset.
Key findings
- An AI judge helps narrow the search
SIFT compares candidate implementations pairwise, then uses the rankings to direct further development and testing. - Small tests can mislead
On TerminalBench, the candidate with the best score on the search set did worse across repeated full-benchmark runs than the judge’s selection. - The reported gains vary by benchmark
Results on the SWE-bench subset favoured judge-guided selection, though less cleanly than some other comparisons. - Full testing still matters
The researchers used benchmark runs to confirm results; the judge helps choose what to test, rather than proving a change works.
Why it matters
Self-improving agents can generate many possible changes, but testing each one at full scale is slow and costly. SIFT’s useful idea is to spend evaluation effort selectively: use cheap checks to catch broken changes, let a judge rank candidates, then reserve heavier tests for the more promising versions. That could let teams explore more designs for coding agents without costs rising in lockstep with the number of ideas.
There is a catch worth keeping in view: a judge’s ranking is a shortcut to deciding what deserves testing, not a substitute for measuring whether the agent actually performs better. The reported results concern coding agents and particular benchmarks, so they do not establish that the approach will transfer to every agent or task.
Our read
SIFT is interesting less as a machine that improves itself by magic than as a more economical way to run experiments on agent designs. That is a useful distinction, and a refreshing one in a field where the word “self-improving” can do rather a lot of work in a headline. Teams considering the approach should keep full evaluations in the loop and check that their small tests reflect the failures they care about.
What to watch
- Whether other teams reproduce the results on coding benchmarks.
- How the judge performs when ranking candidates on different tasks and models.
- Whether the approach proves useful beyond coding agents.
- How often cheap checks and judge rankings agree with full evaluations.
Discussion spark: Should an AI judge be trusted to decide which agent changes get expensive testing, or should every candidate face the full benchmark?
Sources and evidence
- MIT's SIFT cuts coding agent eval costs – VentureBeat (2 October 2026, 22:50 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.