If you run vendor evaluations against public benchmarks, you are working with measurement instruments that have between 2% and 42% invalid questions. That is not a margin-of-error discussion. It is a structural finding about what enterprise AI evaluation infrastructure can support, and it is now visible in the data.

A 2025 audit (Truong et al. 2025) examined GSM8K, one of the most widely-cited mathematical reasoning benchmarks in enterprise AI procurement decks, and found roughly 42% of questions had integrity problems. Ambiguous phrasing, incorrect labels, or framings that admit multiple defensible answers. The same audit found 2% on MMLU Math. That spread is not noise. It tells you the underlying measurement quality varies by an order of magnitude across the benchmarks that procurement frameworks treat as comparable.

What does an invalid question look like in practice? Truong et al. flagged problems including questions where the stated answer does not follow from the stated reasoning, questions with multiple defensible interpretations that produce different numeric answers, arithmetic errors in the source-of-truth annotation, and word problems where the natural reading and the intended reading diverge. These are not edge cases catchable by casual review. They require structured audit work that benchmark authors did not perform when publishing the dataset, and that procurement teams have no incentive to perform when comparing vendor scores.

There is a second issue running in parallel. Arena, the human-vote leaderboard that has become the most-cited capability ranking in 2025–2026, is not measuring what most procurement frameworks assume it measures. Singh et al. found that high Arena standings partly reflect adaptation to the platform itself (what questions get asked, how responses are formatted, which model behaviours human voters reward) rather than general capability that would transfer to a different evaluation environment. A model can climb Arena rankings by getting better at performing well on Arena.

The procurement framing that worked when "vendor A scored 87% and vendor B scored 84% on benchmark X" treated the score gap as signal. That framing assumed the benchmark itself was a stable yardstick. The data now suggests the yardstick has variable accuracy, and the variance is not disclosed in vendor marketing materials that report the headline number.

This is not a critique of the labs running these benchmarks. Building evaluation infrastructure that keeps pace with frontier AI capability is genuinely hard. Benchmarks designed to last for years get saturated within months once they are public, because labs train against them. The integrity problem and the adaptation problem are two faces of the same structural issue: public benchmarks become artefacts of the public benchmark ecosystem rather than neutral instruments measuring something independent.

For practitioners running enterprise evaluations, three implications follow.

Benchmark scores from vendor marketing materials are downstream of measurement-instrument integrity that the vendor does not disclose. The 84-versus-87 gap on a benchmark with 42% invalid questions is not a reliable basis for procurement preference. Treating it as one introduces variance into vendor selection that compounds across hundreds of decisions enterprise-wide.

Internal evaluations on tasks specific to the enterprise's actual workload carry higher signal than public benchmarks. This is not new advice, but the benchmark integrity data makes it concrete. A smaller, well-constructed internal evaluation (50–100 representative tasks scored by people who actually need the model's output) carries more procurement signal than a 1,000-question public benchmark with 2–42% question integrity issues.

Benchmark scores should be reported alongside the benchmark's audit status, not just the raw number. "Vendor X scores 87% on GSM8K (Truong et al. audit: 42% invalid questions)" is a fundamentally more useful procurement input than "Vendor X scores 87% on GSM8K." The first lets the reader weight appropriately. The second invites overconfident decisions.

There is a longer-term issue too. The pace at which benchmarks saturate (Humanity's Last Exam gained 30 percentage points in 2025; SWE-bench Verified went from approximately 60% in early 2024 to near-100% by late 2025) is now faster than the cycle time to design and validate replacement benchmarks. The measurement infrastructure that AI procurement depends on is being outpaced by the thing it measures.

The benchmark ecosystem and the procurement evaluation ecosystem evolved in parallel without coordinating. Benchmarks were designed by research labs to measure progress against research questions: generalisation, reasoning, multi-step composition. Procurement frameworks adopted those benchmarks as proxies for "is this vendor better than that one for our use case." The first community values rapid iteration and is comfortable retiring saturated benchmarks. The second community values stability and comparability across vendor pitches. Those values are now in tension, and the integrity findings are a signal that the second community's needs are not being met by infrastructure designed for the first.

The single-number summary is more readable; it is also less true. The procurement frameworks that survive this period will be the ones that move evaluation closer to the enterprise's actual workloads, treat public-benchmark scores as one signal among several rather than as primary criteria, and disclose benchmark-integrity uncertainty alongside the headline score.


Sources

  • Primary: Stanford AI Index 2026, Chapter 2 (Technical Performance) 2.1 — hai.stanford.edu/ai-index/2026
  • Benchmark integrity study: Truong et al. (2025) — GSM8K and MMLU Math audit, invalid-question rate analysis
  • Arena adaptation study: Singh et al. (2025) — platform-adaptation effects in Chatbot Arena rankings
  • Saturation reference: Center for AI Safety, Humanity's Last Exam; SWE-bench Verified leaderboard