Mercor is building a research team to evaluate AI against real-world work, improve training data and test whether post-training closes capability gaps. Its agenda includes benchmarks across five professional fields and a reported increase in Qwen3.5-397B’s APEX-Agents score from 16% to 27%.
Mercor Watch analysis
What happened
In a post published on 8 October, Mercor says its research team will work across evaluations, expert-informed data and post-training. Its APEX benchmarks cover management consulting, investment banking, corporate law, accounting and software engineering. The company says it plans to expand them as AI changes the work being measured.
Mercor describes a loop: benchmarks identify where a model struggles, domain experts help explain the failures and produce training data, and the model is then post-trained and evaluated again. The company says a recently published guide showed Qwen3.5-397B’s Pass@1 score on APEX-Agents rising from 16% to 27%, and included the training script, model weights and evaluation traces. Read Mercor’s research agenda.
Why it matters
A benchmark can become stale if the job itself changes. Mercor points to shifts in how people work, from producing deliverables towards reviewing them, and says evaluations should test whether models complete tasks reliably without concealing mistakes or exploiting flawed incentives. That makes the choice of tasks and standards as important as the score.
The company’s proposed loop also puts expert time in focus. Experts might set tasks and standards, or judge cases where models and automated graders disagree. Mercor says it is exploring how to get more useful training signals from each expert hour, and how to measure whether a dataset actually improves performance.
Our read
This is a credible research agenda because it connects evaluation to training and gives readers a reported result, not just a promise to study the problem. But the score improvement is Mercor’s account, and one benchmark result does not establish broader workplace competence. The useful test is whether the benchmarks stay realistic, the evaluation traces make the result inspectable, and other groups can reproduce the gains. Happily, Mercor says it has shared the script, weights and traces for this example.
What to watch
- Whether Mercor expands APEX to more occupations and updates tasks as work changes.
- Whether independent researchers reproduce the reported APEX-Agents improvement.
- How Mercor measures expert agreement, dataset quality and improvements beyond its own benchmarks.
Discussion spark: Should companies building AI also design the benchmarks used to judge it, or should those evaluations be led by independent researchers?
Sources and evidence
- Source update (8 October 2026, 16:29 UTC)
not affiliated with or endorsed by Mercor