SWE-bench Verified, the real-world software engineering benchmark that asks AI to resolve actual GitHub issues, moved from approximately 60% accuracy in early 2024 to near-100% on top models by late 2025. Terminal-Bench 2.0 climbed from 14.1% in February 2025 to 77.3% in early 2026. Coding capability has stopped being a frontier-model differentiator. Whether this is the actual plateau or a benchmark-saturation artefact is not yet settled in the data.
The trend over 18 months:
SWE-bench Verified accepted issues from real production repositories and tested whether AI agents could produce patches that pass the same tests human engineers would write. In February 2024, the top model was at approximately 60%. By late 2025, top models clustered in the low-to-mid 70s on the full benchmark and approached 100% on specific subsets. The benchmark is now operating in a regime where the top of the leaderboard is too clustered to identify a clear leader.
Terminal-Bench 2.0, which covers agentic software workflows that require navigating terminals, executing commands, and completing multi-step tasks, climbed from 14.1% (February 2025) to 77.3% (early 2026). Year-on-year capability gain of more than 60 percentage points.
The pattern: coding-specific benchmarks have saturated faster than any other capability dimension on record. The trajectory that took language understanding benchmarks five years to traverse, coding benchmarks have traversed in 18 months.
Two competing interpretations are circulating.
The first: coding is a genuinely solved problem at the level of well-bounded engineering tasks. Frontier models can now write code, debug code, integrate code into existing repositories, and reason about code architecture at a level that exceeds many junior-to-mid-level human engineers. The benchmark saturation is real, the capability is real, and the procurement implication is that coding capability is no longer a meaningful differentiator between frontier-tier vendors.
The second: the coding benchmarks are saturating faster than other benchmarks because coding is a particularly visible domain where model providers focus their post-training optimisation. The capability gain is real on the benchmark, but the benchmark covers a narrower slice of the actual coding task distribution than the headline suggests. Real-world coding workloads (large legacy codebases, undocumented systems, ambiguous requirements, multi-team coordination) are not well-represented in SWE-bench or Terminal-Bench. The benchmark plateau is not equivalent to a capability plateau in deployment contexts.
Both interpretations have evidence. The first is supported by the trajectory data and by GitHub developer surveys showing that AI coding tools are now used by 90%+ of professional developers. The second is supported by the relatively narrow scope of the public benchmarks compared to the long tail of real-world coding workloads, and by reports of capability variation in domain-specific coding tasks that the benchmarks do not cover: embedded systems, formal verification, niche language ecosystems.
For procurement frameworks evaluating coding-AI vendors, the contested interpretation creates uncertainty about which procurement approach captures the most value.
If interpretation one is correct, procurement should treat coding capability as table stakes. Every frontier vendor passes the capability threshold, and the procurement decision should be made on cost, integration quality, and developer-experience metrics rather than on capability benchmarks. The procurement framework should look more like a SaaS procurement framework than a capability evaluation framework.
If interpretation two is correct, procurement should still evaluate coding capability on the workload-specific dimensions that the public benchmarks do not cover: large codebase reasoning, edge-case language handling, domain-specific code patterns. The procurement decision should retain capability evaluation as primary criteria, with cost and integration as secondary.
The data does not currently settle which interpretation is correct. The benchmark trajectory is consistent with both. The available evidence about deployment performance in real workloads is fragmented and not yet systematic enough to validate either interpretation against the other.
Three observations follow regardless of which interpretation is correct.
The first: the capability differentiator between frontier-tier vendors has clearly moved off the public coding benchmarks. The benchmarks no longer separate vendors at the top of the leaderboard. Whatever the deployment-level capability picture turns out to be, the public benchmarks have stopped functioning as differentiators.
The second: the cost-per-task spread between coding-AI vendors is now wider than the capability spread. Vendor cost-per-coding-task varies by significant multiples across frontier vendors. Even under interpretation two, where capability still differentiates, the cost spread is consequential enough to matter for procurement decisions on most coding workloads.
The third: the speed of capability gain on coding benchmarks, 60-point gains within 18 months, suggests that whatever direction the next benchmark cycle takes, the gains will arrive quickly. Procurement frameworks built on assumptions about a stable coding capability landscape over multi-year procurement cycles will be repeatedly mis-aligned with the underlying capability state. Shorter procurement cycles capture more value than long cycles in this category.
The contested question: is the coding benchmark plateau real, or is it an artefact of the public benchmarks capturing only a slice of the coding task space? The data does not currently support a confident answer. Procurement frameworks that explicitly hold both interpretations open, and budget for coding capability uncertainty on a workload-by-workload basis, will out-perform frameworks that commit to either interpretation prematurely. The next 12–18 months of deployment data, particularly from large-codebase enterprise environments, will probably resolve which interpretation holds.
Sources
- Primary: Stanford AI Index 2026, Chapter 2 (Technical Performance) 2.5 — hai.stanford.edu/ai-index/2026
- SWE-bench data: SWE-bench Leaderboard, 2026 — real GitHub-issue resolution evaluation
- Terminal-Bench data: Terminal-Bench 2.0 Leaderboard, 2026 — agentic terminal-task evaluation
- Developer adoption context: GitHub Octoverse and Stack Overflow Developer Survey, 2025
Discussion