> ## Content Index
> Fetch the complete content index at: https://aiadoption.org/llms.txt
> Use this file to discover other available public pages before exploring further.

# The benchmark saturation problem
- URL: https://aiadoption.org/ai-analysis/the-benchmark-saturation-problem/
- Published: 2026-10-02T03:47:20.000Z
- Updated: 2026-10-02T03:47:20.000Z
- Description: Benchmarks built to last for years are saturating in months. Humanity's Last Exam gained 30 percentage points in 2025. SWE-bench Verified went from 60% to near-100% in eighteen months. The measurement infrastructure that AI procurement depends on is being outpaced by the thing it measures.
- Author: Jassie
- Tags: AI Analysis, AI Procurement, AI Research

The benchmarks built to last for years are saturating in months. That is not framing. It is what the data shows.

Humanity's Last Exam, a benchmark assembled in 2024 specifically to be hard for AI and favourable to human experts, gained 30 percentage points in 2025\. SWE-bench Verified, the real-world software engineering benchmark, moved from approximately 60% accuracy in early 2024 to near-100% by late 2025\. Multiple legacy benchmarks (ImageNet, SuperGLUE, MMLU) now operate above human baseline, which means they have stopped being useful for differentiating frontier models.

The pattern in the data: a benchmark is designed to be hard. It is published. Within months, top models have closed most of the gap to human performance. Within roughly a year, the benchmark is saturated. The measurement window, the period during which the benchmark produces useful signal about which frontier model is more capable on the dimension it measures, has compressed from years to months.

The structural cause has two components. The first is that frontier model capability is now climbing fast enough on most benchmark dimensions that any single benchmark has a short useful life. The second is that publication of a benchmark creates incentives: published evaluations are public training signal whether or not the underlying questions are in the training set, because labs can study what kinds of questions the benchmark contains and optimise for that style. Both components compress the measurement window in the same direction.

This is not a problem that better benchmarks solve. Benchmarks designed to be harder face the same compression. Humanity's Last Exam was assembled specifically to be hard for AI; it showed the same 30-point saturation curve in approximately one year that earlier benchmarks showed over five-to-ten year windows.

## Three observations follow.

**First: the methodological assumption that the AI capability evaluation literature has been operating under, that benchmarks are stable instruments measuring something durable, has been broken by the rate of capability gain.** The literature has not fully updated to this. Comparing models published twelve months apart on a benchmark that saturated six months ago is a defensible exercise that produces no useful signal about the comparative capability of the models.

**Second: the same compression is starting to affect benchmarks as instruments of communication.** A benchmark's score for a model is now best interpreted as "where this model sat relative to the frontier at the time it was evaluated", not as a stable property of the model. A model that scored 65% on a benchmark in mid-2025 may evaluate at 92% on the same benchmark today if rerun, not because the model improved but because the evaluation environment, prompting conventions, and reference baselines shifted.

**Third: composite benchmarks, including scaled multi-benchmark indices, carry more durable signal than individual benchmarks, because they aggregate across measurement instruments and reduce per-instrument noise.** But composite benchmarks have their own problem. They obscure which specific capabilities are saturating, which are still climbing, and which have plateaued. A composite that averages over benchmarks at different saturation stages reports a more stable trajectory but a less informative one.

The challenge for evaluation infrastructure design is straightforward to state and hard to solve. The infrastructure has to be designed for a rate of change that exceeds the design-and-validation cycle. Traditional benchmark publication cycles (design, publish, community adopts, results published, benchmark gets used for years) no longer match the underlying rate. The benchmarks that survive will either be ones deliberately impossible to saturate (because the underlying task space is open-ended in a way the benchmark captures) or ones continuously refreshed at a tempo matching the capability climb.

Early examples are visible. SWE-bench Verified was followed by SWE-bench Multimodal, SWE-bench Live, and SWE-bench Pro, each a refresh that introduces new tasks the prior version could not measure. ARC-AGI-1 was succeeded by ARC-AGI-2 once the original saturated. Terminal-Bench had a 2.0 release within months of 1.0 saturating. The benchmark-as-static-artefact model is giving way to a benchmark-as-living-platform model.

> There is a deeper methodological observation about what these refreshes can and cannot fix. A benchmark designed to be hard for current frontier models, refreshed every six months, can keep producing useful signal about frontier capability. It cannot produce useful signal about whether the model your enterprise plans to deploy in 2027 will be capable of the specific workload your enterprise needs. That requires evaluation infrastructure that lives inside the enterprise, designed around the enterprise's workload, scored by people who use the output. Public benchmarks, even the refreshed ones, are now best understood as one data point about which models are in the frontier tier. They are no longer useful as the evaluation of choice for downstream deployment decisions.

The methodological observation: the AI evaluation literature is mid-way through a paradigm shift from benchmark-as-yardstick to benchmark-as-snapshot-of-frontier-state. Practitioners using benchmarks as procurement criteria are operating under the older paradigm. The literature has begun the shift. Procurement has not. The gap between the two is where most benchmark-driven procurement decisions are now producing under-performing outcomes.

---

### Sources

- **Primary**: Stanford AI Index 2026, Chapter 2 (Technical Performance) 2.1 (and 2.5 for SWE-bench) — [hai.stanford.edu/ai-index/2026](http://hai.stanford.edu/ai-index/2026?ref=aiadoption.org)
- **Composite benchmark analysis**: AI Index 2026 own analysis, drawing on the Index's curated suite
- **Humanity's Last Exam**: [Center for AI Safety, lastexam.ai](https://lastexam.ai/?ref=aiadoption.org) — expert-level benchmark, 2024 release
- **SWE-bench data**: [SWE-bench Leaderboard](https://www.swebench.com/?ref=aiadoption.org), 2026 — software engineering benchmark trajectory