If your AI deployment plan treats published productivity gains as a generalisable signal ("AI improves knowledge worker productivity by 26%, so we expect similar gains"), the 2025 study landscape reveals a structural problem with the framing. The reported gains range from -19% (METR study of experienced open-source developers) to +50% (marketing tasks). The mean masks a distribution where the gain depends substantially on task, worker, and measurement methodology.
[CHART fig_4427_2026 — AI productivity micro studies summary]
The study landscape is now substantial enough to compare. Brynjolfsson and colleagues (2025) studied AI assistance in customer support, finding +14–15% productivity gain. The gain was concentrated among less-experienced workers; experienced workers showed smaller gains or none. Cui and colleagues (2025) studied GitHub Copilot use among software developers and found +26% productivity gain. The gain was concentrated in the kinds of tasks where the AI suggestions were directly usable (boilerplate, common patterns). Ju and Aral (2025) studied AI assistance in marketing tasks and found +50% productivity gain. The gain was concentrated in creative output tasks where AI helped with ideation and first-draft generation.
The Becker et al. (2025) METR study showed -19% productivity for experienced open-source developers using AI assistance. The study has been the subject of replication attempts; the replication did not reproduce the negative effect. The interpretation is contested, but the study at minimum demonstrates that AI assistance can degrade productivity in some specific worker-task combinations rather than always improving it.
What does the distribution of gains tell us about deployment strategy?
The first pattern is that the gains concentrate at specific worker-task intersections rather than spreading evenly. Less-experienced workers gain more in tasks that benefit from AI suggestion. Experienced workers gain less or nothing in tasks they already do efficiently. The task structure matters too: tasks where AI output is directly usable produce larger gains than tasks where AI output requires substantial revision before use.
The second pattern is the variance is large enough that organisational averages are misleading. An organisation deploying AI assistance broadly should not expect a uniform 26% gain. The actual outcome depends on the mix of workers (experienced vs less experienced), tasks (AI-suitable vs not), and integration quality (does the AI assistance integrate into existing workflows?).
The third pattern is measurement methodology matters. The Brynjolfsson customer support study measured tickets resolved per hour, a direct production metric. The Cui Copilot study measured tasks completed in a specified time window, also direct. The Ju marketing study measured creative output quality and quantity, a composite metric. The METR study used a different measurement methodology that produced the divergent finding. Studies that measure different things produce different findings. The comparison "AI productivity gain is 26% per Cui et al." is not directly comparable to "AI productivity gain is 50% per Ju and Aral" because the measurement methodology and worker pool are different.
Three implications for deployment planning.
The first: target deployments at the worker-and-task combinations where the gain is most likely to materialise. The data shows the largest gains are at the intersection of less-experienced workers with AI-suitable tasks. Marketing tasks where AI assists ideation. Customer support tickets where AI suggests responses. Junior software development where AI generates boilerplate. Deployments that match this profile will see the largest measured productivity gains. Deployments that do not, experienced workers doing tasks they already do efficiently, will see smaller gains or none.
The second: build measurement into the deployment from the start, not retrospectively. The headline productivity numbers from the studies are useful as benchmarks but not as direct deployment predictors. Each enterprise deployment is its own data point. Productivity should be measured at the worker-task level (not the aggregate level) before and after AI assistance is introduced, with the measurement methodology defined in advance. Without this discipline, organisations end up with anecdotal evidence rather than actual deployment ROI data.
The third: the agent deployment data (single-digit scaled use across most functions per #77) suggests the productivity gains are not yet being realised at deployment scale. The gains in the studies are real, but the studies were typically of pilots or focused deployments. Scaling those gains to enterprise-wide deployment requires the agent layer that most organisations have not yet deployed. The productivity gains visible in the studies are leading indicators of what is possible; the deployment scale data is the lagging indicator of what is realised.
The contested question: are these productivity gains generalisable to broader knowledge work, or are they specific to the tasks studied? The case for generalisability rests on the consistency of the direction (most studies report positive gains) and the underlying mechanism (AI suggestion plus human curation produces faster output than human-only work). The case against generalisability rests on the study heterogeneity (different tasks, different measurement methods, different worker populations) and the negative findings (METR, despite the contested replication, demonstrated that the direction can flip in specific contexts).
The defensible interpretation: AI productivity gains are real for many worker-task combinations and not real for others. The gains depend on the specific combination in ways that aggregate numbers obscure. Strategic plans built on "AI improves knowledge worker productivity by [headline number]" will produce worse decisions than plans built on "AI improves productivity for this specific worker doing this specific task by this specific amount, based on our pilot measurement."
For executive teams setting AI investment direction, the planning anchor needs to shift from "we expect productivity gains of X%" to "we will measure productivity gains at the worker-task level for each deployment and decide expansion based on the measured gain." The discipline is harder. The data quality is much better, and the decisions that follow are correspondingly better.
Sources
- Primary: Stanford AI Index 2026, Chapter 4 (Economy) 4.4 — hai.stanford.edu/ai-index/2026
- Customer support study: Brynjolfsson et al., 2025 — GenAI assistance in customer support; +14–15% productivity gain concentrated in less-experienced workers
- Software development study: Cui et al., 2025 — GitHub Copilot productivity study; +26% gain concentrated in AI-suitable task patterns
- Marketing study: Ju & Aral, 2025 — marketing task productivity study; +50% gain in creative output tasks
- Negative-finding study: Becker et al. (METR), 2025 — experienced open-source developer study; -19% productivity finding (replication contested)
- Agent deployment context: McKinsey & Company "State of AI" Survey, 2025
- Macro productivity context: OECD / Filippucci; IMF; BIS, 2025
Discussion