Gemini Deep Think won gold at the 2025 International Mathematical Olympiad. The same year, GPQA Diamond saturated and FrontierMath Tier 4 remained below 30%. Mathematical capability at the frontier is climbing fast on some axes and plateauing on others. Procurement frameworks anchored on "mathematical reasoning capability" as a single dimension are operating with insufficient resolution.

The IMO 2025 data point: Gemini Deep Think scored 35 points end-to-end in natural language within the 4.5-hour human time limit. A 30 percentage point year-on-year gain over the 2024 result. The benchmark is not multiple choice. It requires generating full mathematical proofs that human evaluators score against IMO marking criteria. The capability went from "competent silver-medal performance" in 2024 to "gold-medal-grade reasoning across all six problems" in 2025.

The GPQA Diamond data point: top model accuracy is now within a few percentage points of the human expert reference of approximately 81.2%. The benchmark, PhD-level science questions, was assembled in 2023 specifically to be hard for AI. By early 2026, multiple frontier models cluster near the human-expert ceiling. The benchmark is saturated relative to its design intent.

The FrontierMath Tier 4 data point: the most difficult known mathematical reasoning benchmark, research-grade problems that take human mathematicians hours to days to solve, remains below 30% across the leading models. Top model performance has improved year-on-year, but the curve is shallower than the equivalent curves for GPQA Diamond or IMO-style problems.

The pattern in the data: mathematical reasoning capability is not climbing at a uniform rate across difficulty tiers. The mid-range (PhD-level science, undergraduate-competition maths) has either saturated or is approaching saturation. The high end (research-grade mathematics, novel proof construction in unfamiliar territory) is climbing but on a different curve.

What this means for procurement frameworks evaluating mathematical reasoning capability.

The single-dimension framing ("how good is this model at maths?") produces a single number that obscures the structural variation. A model that scores 75% on GPQA Diamond may score 35 points on IMO problems and 22% on FrontierMath Tier 4. Each number is real. Each measures something different. The aggregate "maths capability" is a misleading abstraction.

For workloads in regulated domains where mathematical reasoning is part of the workflow (actuarial analysis, quantitative finance, scientific research), the specific tier of mathematical capability that matters depends on the workflow. Most enterprise workloads are at the mid-range tier, where multiple frontier models are now capable. A small minority are at the high end, where the capability gap between frontier models is wider and the procurement decision is more consequential.

Three observations about how the data should shape procurement frameworks.

The first: benchmark-driven procurement for mathematical reasoning should distinguish between the difficulty tier of the benchmark and the difficulty tier of the workload. A vendor that excels at GPQA Diamond is not necessarily the right vendor for workloads at the FrontierMath Tier 4 tier, and vice versa. The procurement question is "which vendor performs best at the difficulty tier of the actual workload", not "which vendor performs best on average across maths benchmarks."

The second: capability at the saturated tier (GPQA, AIME) is no longer a differentiator between frontier vendors. Procurement frameworks that rely on these benchmarks are evaluating on a dimension where every frontier vendor passes the threshold. The differentiator at this tier has moved to reliability, cost, and domain fit.

The third: capability at the high end (FrontierMath Tier 4) is genuinely differentiating between frontier vendors, with year-on-year improvement that suggests this tier will remain a differentiating dimension for procurement decisions through at least 2027.

The prescription: rebuild procurement evaluation for mathematical reasoning around explicit difficulty tier matching. The evaluation framework should specify the difficulty tier of the workload, identify benchmarks that test at that tier, evaluate vendor performance at that tier, and weight performance at adjacent tiers as a sanity check rather than as primary criteria.

Concretely, the framework should answer four questions in sequence.

What is the difficulty tier of the workload? Not "is the workload mathematical reasoning?" but "is the workload at the saturated-benchmark tier (GPQA-like), at the IMO-style problem tier, or at the research-frontier tier?"

Which benchmarks credibly measure capability at the identified tier? The list will be short. Most public benchmarks measure at one or two tiers, and the specific benchmark that matches the workload tier is usually identifiable.

Which vendors perform well at that tier? This question is answered by published benchmark scores at the relevant tier, supplemented by internal evaluation on workload-specific tasks.

Is the capability at that tier saturated, climbing, or differentiating? Saturated tiers are not procurement differentiators. Climbing tiers are. Differentiating tiers warrant deeper evaluation and shorter procurement cycles to capture the value of capability gains as they arrive.

The single-number "maths capability" framing was a useful procurement abstraction when most frontier models were operating at a similar tier of mathematical reasoning. The data now shows that abstraction is no longer matching the underlying capability landscape. Procurement frameworks that update accordingly will capture more value from the next 12–24 months of capability gains than frameworks that retain the single-dimension framing. Adjust the framework now while it's a planning exercise, or adjust it later when a procurement cycle exposes the mismatch.


Sources