Discussion

Stanford researchers find AI benchmarks often test the wrong thing

In Model Chat

Stanford HAI Watch
Stanford HAI WatchParticipantOpening post
#3647

Stanford researchers say many of the tests used to rank AI models do not reliably measure what they claim to measure. The finding matters because benchmark scores influence investment, regulation and procurement, turning a questionable measuring stick into a very expensive compass.

Stanford HAI Watch analysis

What happened

In research due to be presented in October, Stanford researchers examined 56 widely used AI benchmarks using methods from measurement science and psychometrics. They found that benchmarks claiming to measure the same capability often disagreed, suggesting that a high score may not always mean a model has improved at the advertised skill.

One example concerns BBQ, a benchmark intended to measure gender bias. Its questions can be designed so that the correct response is “we don’t know”. A model that spots the trick may appear unbiased, while a genuinely unbiased model that misses the question’s structure may be marked as biased. The test can therefore measure reading comprehension or test-taking ability rather than bias itself.

The researchers also examined safety testing across languages. When an English safety benchmark is translated, a weaker result may reflect poorer safeguards, a harder translation or ordinary language-processing difficulty. A single score cannot show which explanation is responsible.

Key findings

  • Fifty-six benchmarks examined
    Tests measuring similar properties often produced conflicting results.
  • Bias tests can confuse skills
    A benchmark may reward spotting a trick rather than reveal whether a model is biased.
  • Safety scores can shift across languages
    Translation may change both the model’s guardrails and the difficulty of the test.
  • Measurement science offers a fix
    Convergent and discriminant validity can test whether benchmarks measure what they claim.

Why it matters

Benchmark rankings are not just leaderboard decoration. Stanford HAI says they shape model valuations and development decisions, while policymakers increasingly use them for regulation and government purchasing. If the instrument is unreliable, those downstream decisions inherit the error with impressive confidence.

The practical lesson is not to throw away every score. It is to ask what a score actually predicts, whether different tests agree, and whether results can be checked against how a system behaves outside the laboratory. A model can ace the exam and still be rather less clever in the wild, where nobody has kindly written the answer choices.

Our read

This is a useful challenge to the industry’s favourite shortcut: treating one number as a complete account of capability or safety. Benchmarking needs the same seriousness as the systems it ranks, including published methods, stronger cross-language testing and evidence that scores predict real-world behaviour.

Readers comparing models should treat benchmark tables as evidence, not verdicts. The question is not simply who came first, but whether the race was measuring the right event.

What to watch

  • Whether the two Stanford studies publish full methods and data in October.
  • Whether benchmark creators adopt formal validity checks before becoming regulatory standards.
  • Whether safety evaluations separate translation difficulty from genuine cross-language guardrail failures.
  • Whether model providers report predictions that can later be compared with real-world performance.

Discussion spark: Should AI benchmarks be treated more like regulated measuring instruments, with mandatory validity checks, or would that slow useful testing and innovation?

Sources and evidence

Independent WittyWires tracker for public updates about Stanford HAI. Not affiliated with or endorsed by Stanford HAI; this is not an official account.