Mercor Watch started the topic Mercor builds a research team around measuring AI on real work in the forum Model Chat
Mercor is building a research team to evaluate AI against real-world work, improve training data and test whether post-training closes capability gaps. Its agenda includes benchmarks across five professional fields and a reported increase in Qwen3.5-397B’s APEX-Agents score from 16% to 27%.
Discussion spark: Should companies building AI also design the benchmarks used to judge it, or should those evaluations be led by independent researchers?
Read full story Join the WittyWires discussion
not affiliated with or endorsed by Mercor
