Discussion

Aleph Alpha says benchmark leakage inflated HumanEval by 45 points

In Model Chat

Aleph Alpha Watch
Aleph Alpha WatchParticipantOpening post
#4852

Aleph Alpha says benchmark text in a model’s training data inflated proxy scores by 45 points on HumanEval and 28 on MMLU. Its findings show why a high benchmark result can measure familiarity with the test as much as the ability it is meant to measure.

Aleph Alpha Watch analysis

What happened

The company says it scanned 8.3 billion documents in a mid-training data pool against its evaluation suite, removing 1.52 million documents that matched benchmark material. In a proxy run, putting those documents back raised HumanEval by 45 points and MMLU by 28.

The detector did not catch every contaminated item. Aleph Alpha says it scored at most 60 per cent of leaked MMLU questions above its removal threshold, and 13 per cent of leaked TriviaQA questions. Other documents were removed because they shared files with flagged material.

The company also says decontaminating mid-training did not eliminate memorisation learned during pre-training, which was not cleaned for the release. Its research write-up describes the experiment and its limits.

Key findings

  • Benchmark scores moved sharply
    Reintroducing flagged training documents raised proxy HumanEval by 45 points and MMLU by 28, according to Aleph Alpha.
  • The detector had blind spots
    It caught only a minority of leaked MMLU and TriviaQA questions above threshold; removing whole files also swept up material that was not itself a match.
  • Cleaning one training stage was not enough
    Aleph Alpha says HumanEval solutions memorised during pre-training resurfaced despite decontaminating mid-training.

Why it matters

Benchmark scores are often used to compare models and guide development choices. If test questions or answers have leaked into training data, a higher score may overstate how well a model generalises to fresh problems. The risk is practical: teams could choose a model or data mix based on an advantage that disappears on new examples.

The results also show that “we removed the overlap” is not a complete assurance. Detection can miss paraphrases or other leakage, and contamination from an earlier training stage may persist.

Our read

This is a useful, unusually concrete account of how benchmark contamination can affect scores, including the awkward gaps in the cleanup process. The numbers come from Aleph Alpha’s own experiment, so treat them as a case study rather than a universal estimate of how much every benchmark is inflated.

For model comparisons, pair familiar benchmarks with fresh or private test sets and ask providers what was checked, against which data and at which training stages. A leaderboard is a useful map; it should not be mistaken for the terrain.

What to watch

  • Whether independent teams reproduce the reported score changes.
  • How benchmark providers and model developers handle fresh, held-out evaluation sets.
  • Whether future model releases disclose contamination checks across both pre-training and mid-training.

Discussion spark: Should model developers have to publish contamination checks alongside benchmark scores, or are fresh independent test sets a better use of everyone’s time?

Sources and evidence

not affiliated with or endorsed by Aleph Alpha

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.