OSWorld, the benchmark that tests AI agents on real computer tasks across operating systems, rose from approximately 12% accuracy to 66.3% in one year. Within six percentage points of human performance. The capability that was not in your procurement evaluation framework last year now does most of the work.

GAIA, the general AI assistants benchmark, shows the same trajectory at a different point in time. Accuracy rose from roughly 20% in January 2025 to 74.5% in September 2025. Five-to-six times improvement on agent benchmarks in 2025. The discrete jumps in capability that the previous decade of AI development produced over years now happen over quarters.

This is a different kind of capability shift than the ones procurement teams have been tracking since 2023. Capability gains on language understanding (MMLU), maths (GSM8K), and code generation (HumanEval) followed an exponential curve but stayed within a recognisable task type. The model got better at answering questions, generating snippets, doing maths. Agents are a categorical shift: from answering questions to completing tasks.

Specifically, what an agent does that earlier model paradigms did not: read a screen, decide what to click, click it, observe the result, decide the next action, repeat across dozens of steps in pursuit of a goal that was specified at the start of the session. The model is not just emitting text: it is operating a computer. OSWorld measures this on real Linux, macOS, and Windows environments. The 66.3% score means the agent completed about two-thirds of multi-step computer tasks (file management, web research, application operation) end-to-end, without human intervention.

For procurement frameworks built around language model evaluation, the agent shift creates three problems.

The first is that the unit of evaluation changes. A language model's unit of evaluation is the response: given an input, how good is the output? An agent's unit of evaluation is the task: given a goal, did the agent complete it correctly? These require different evaluation infrastructure. Internal evaluations sized for language model output review (50–100 prompt-response pairs) do not translate to agent evaluations, which need scenarios, environments, and end-state checks.

The second is that the cost model changes. A language model query produces predictable token usage. An agent task produces variable token usage depending on how many steps the agent takes, how many tool calls it makes, how long it persists before deciding the task is complete or failed. Procurement cost forecasts built on per-query pricing under-estimate agent-task costs by significant multiples. Some workloads that are economically viable as language model queries become uneconomic as agent tasks, and the inverse is also true for workloads where agent automation eliminates significant human time.

The third is that the failure modes change. Language models fail by producing incorrect or unhelpful output, which is recoverable through prompt iteration or human review. Agents fail by taking incorrect actions that may have side effects (deleted files, sent emails, completed financial transactions) that are not recoverable through review after the fact. Agent procurement requires evaluation of failure modes specific to the deployment context, which most procurement frameworks have no fields for.

The capability is now real enough that these problems are practical procurement problems, not future-state hypotheticals. OSWorld at 66.3% means agents are useful for a meaningful share of multi-step desktop tasks today. The labs publishing these benchmarks are the same labs shipping the production agent products that enterprises are evaluating now.

The trajectory: agent capability is climbing on a curve that resembles the 2023–2024 language model capability curve, but with a one-year compression. OSWorld went from approximately 12% to 66.3% in one year. If the curve continues at that pace, agent benchmarks will be at human baseline within the next benchmark cycle, and saturated within two.

What this means for procurement frameworks: the next 12–18 months will see agent capability become a first-class procurement category, parallel to but distinct from language model procurement. The frameworks built now will be evaluating substantially different vendor capabilities than the language model evaluations that procurement teams have spent the past two years refining. Frameworks that try to fold agent evaluation into existing language model procurement processes will under-perform frameworks that build agent procurement as a separate workstream from the start.

The vendors will not wait. By the end of 2026, every frontier lab will be selling agentic products as first-class offerings, not as features bolted onto language model APIs. Procurement teams that have not built agent-specific evaluation by then will be evaluating agent vendors using language model evaluation infrastructure, and getting bad outcomes. The lead time to build evaluation infrastructure is months, not weeks. The window to start is closing.


Sources