The Instruments That Stopped Measuring
Researched by an agentic pipeline · reviewed and gated by the author
When MMLU cannot distinguish between frontier models and GSM8K scores 99 per cent across the board, citing these numbers in a procurement evaluation is the appearance of rigour without its substance.
The Measurement Problem No One Acknowledged
Every enterprise that selects, procures or governs an AI system relies, somewhere in the decision chain, on benchmark scores. Model A scores 92 per cent on MMLU; Model B scores 91 per cent. The procurement team picks A. The compliance function records the score. The board receives a report citing it.
The EvalEval Coalition — more than thirty researchers across ETH Zurich, Stanford, MIT, Cambridge and UC Berkeley — has now demonstrated, with an uncertainty-aware methodology across sixty standard benchmarks, that this entire chain of reasoning is structurally unsound [S1].
Their finding is not that benchmarks are imperfect. It is that nearly half of them have lost the ability to distinguish between leading models at all.
What the Coalition Measured
The study applied a saturation index to sixty text-based LLM benchmarks selected from 190 candidates identified across sixty-one technical reports published between January 2022 and November 2025. Twenty-three annotators with dataset curation expertise catalogued fourteen properties per benchmark. A Bayesian regression model (R² = 0.884 ± 0.012) identified which properties predict saturation [S1].
Twenty-nine of sixty benchmarks — 48.3 per cent — exhibit high or very high saturation, meaning the score differences between frontier models fall within measurement uncertainty. Fourteen benchmarks exceed 0.9 on the saturation index, where 1.0 represents complete inability to discriminate [S1][S2].
The individual numbers confirm what the aggregate suggests. Frontier models cluster between 88 and 93 per cent on MMLU, above 99 per cent on GSM8K, above 95 per cent on HellaSwag, and between 90 and 95 per cent on HumanEval [S5]. At these compression levels, a two-point score difference is statistical noise, not a capability signal.
Every Common Safeguard Has Failed
The study tested five hypotheses about what might protect a benchmark from saturation. The results eliminate every commonly cited defence.
Private test sets show no statistically meaningful difference in saturation compared with public benchmarks [S1]. Keeping the answers hidden does not prevent models from saturating the measurement.
Open-ended generation formats fare no better. No significant difference in saturation exists between closed-ended and open-ended benchmarks (p = 0.40) [S1]. Switching from multiple-choice to free-form answers does not restore discriminative power.
Multilingual coverage appeared protective until recency was controlled for. Multilingual benchmarks average 32.9 months old compared with 48.9 months for English-only tests. The apparent resistance is youth, not linguistic diversity [S1].
Expert curation partially delays saturation but does not prevent it. The effect does not hold against age [S1].
The single strongest predictor of saturation is benchmark age. This is a structural finding: saturation is not a design failure that better engineering can fix. It is an exposure dynamic inherent in the relationship between a static measurement instrument and a rapidly improving population of subjects.
The Consequences Already Measured
Three independent findings confirm the measurement failure is not theoretical.
A contamination-free variant of MMLU removed questions demonstrably present in training data and recorded significant score drops for frontier models [S1]. The standard MMLU score reflects, in part, what a model has memorised rather than what it can do.
The CLEAR framework documented a 37 per cent gap between lab benchmark scores and real-world deployment performance, with consistency dropping from 60 per cent on a single run to 25 per cent across eight consecutive runs [S6]. In operational conditions, benchmark scores are poor predictors of system behaviour.
NIST’s response is instructive. In July 2026, it launched the Artificial Intelligence Technology Evaluation programme, a sequestered testbed where datasets are kept entirely separate from any training process and models are evaluated blind [S4]. The programme’s architecture implicitly acknowledges that public benchmarks can no longer serve an assurance function.
What Remains Unproven
The study’s constraints matter for how far the conclusion travels.
The design is observational. The coalition demonstrates that saturated benchmarks no longer discriminate between models, but cannot determine in each case whether saturation reflects genuine task mastery or measurement degradation [S1]. A model scoring 99 per cent on GSM8K may have genuinely mastered grade-school mathematics, or it may have absorbed enough of the distribution to appear to have done so.
The saturation index involves defensible but not unique parameter choices — an α = 0.5 downweighting and a top-five model comparison [S1]. The headline rate could shift with different parameters, though the direction is robust.
The study covers text-only benchmarks. Multimodal evaluation, increasingly central to enterprise deployment, is unexamined [S1]. And the saturation tool itself has limited external adoption, meaning the methodology has not yet been independently stress-tested [S3].
What a Programme Should Conclude
Any procurement, compliance or governance process that cites standard benchmark scores as evidence of model capability is operating on instruments that cannot do what they claim to do. This is not a matter of imprecision. It is structural failure.
When MMLU cannot distinguish between frontier models and GSM8K scores 99 per cent across the board, citing these numbers in a procurement evaluation is the appearance of rigour without its substance.
The surviving benchmarks — SWE-bench Verified, GPQA Diamond, LiveCodeBench, ARC-AGI-2 — share a pattern: continuously refreshed tasks, expert curation, and contamination-resistant design [S5]. The principle transfers to enterprise evaluation: measure what the system actually needs to do, not a proxy chosen for convenience.
The cost of rebuilding is real but bounded. A documented 50× cost variation between evaluation approaches achieving comparable accuracy on the same tasks [S6] indicates that efficient alternatives exist. The expense lies not in building task-specific evaluation but in admitting that the standard instruments have failed.
Three conclusions follow from the evidence. First, audit every evaluation gate in the AI programme that references a standard benchmark. Where a saturated benchmark is the sole basis for a decision, the decision has no valid measurement behind it. Second, replace saturated benchmarks with evaluation built from production-representative data, following the design pattern the surviving benchmarks demonstrate. Third, treat model selection as a deployment trial against the enterprise’s own workload, not a comparison of published scores. The 37 per cent gap between benchmark and deployment performance [S6] means that selecting on benchmark scores is selecting on the wrong variable.
The benchmarks have not become wrong. They have become uninformative. The difference matters: there is nothing to fix. The instruments need replacing.
Sources
- Akhtar, Reuel et al. — When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation — February 2026 — https://arxiv.org/html/2602.16763v1
- EvalEval Coalition — When AI Benchmarks Stop Measuring Progress — June 2026 — https://evalevalai.com/2026/06/30/saturation-blog/
- EvalEval — benchmark-saturation GitHub repository — 2026 — https://github.com/evaleval/benchmark-saturation
- NIST — Announcing NIST’s Artificial Intelligence Technology Evaluation (AITE) — July 2026 — https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite
- SentryML — LLM Benchmarks in 2026: Which Still Discriminate, and How to Run — 2026 — https://sentryml.com/posts/llm-benchmarks-2/
- Kili Technology — AI Benchmarks 2026: Top Evaluations and Their Limits — 2026 — https://kili-technology.com/blog/ai-benchmarks-guide-the-top-evaluations-in-2026-and-why-theyre-not-enough