The gap between top closed-weight and top open-weight AI models was 0.5% in August 2024. By March 2026 it had widened to 3.3%. The convergence narrative built into 2024-era multi-tier AI strategy plans needs updating, because the trend has reversed.

The data over three years: in May 2023, the leading closed-weight model (GPT-4-0314) outperformed the top open-weight model (Vicuna-13B) by 174 Elo points (15.2%). Stronger open-weight releases (Mixtral, WizardLM, Llama-3.1-405B) narrowed the gap to 7 Elo points (0.5%) by August 2024. The convergence direction looked clear. Open-weight was catching up. Strategic AI plans built in late 2024 incorporated that direction, treating open-weight frontier capability as a fast-arriving alternative to closed-weight model dependence.

The 2025 trajectory did not continue the convergence. Open-weight releases continued to ship (Llama 4, Mistral updates, Qwen 3), but closed-weight frontier capability advanced faster. By March 2026, six of the top ten models on the Arena Leaderboard are closed-weight. The gap to the top open-weight model is 3.3%. Not a wide gap, but moving in the opposite direction from 2024.

This is a meaningful reversal for strategic AI planning that treated the closed-vs-open dynamic as resolved. Three observations.

The first: the cause is not that open-weight development slowed. Llama 4, DeepSeek-V3, and Qwen 3 are objectively more capable than Llama 3.1, DeepSeek-V2, and Qwen 2.5. Open-weight capability is climbing on the same trajectory it climbed through 2024. The gap reopened because closed-weight capability climbed faster, partly because frontier closed-weight labs are operating at training compute scales that exceed what open-weight releases typically use, and partly because closed-weight labs are integrating reasoning-specific post-training techniques (extended reasoning, tool use, agent workflows) that open-weight releases adopt with months of lag.

The second: the gap dynamic is now asymmetric. When closed-weight frontier capability climbs by 100 Elo points, open-weight typically catches up within 6–12 months. When the closed-weight curve accelerates, as it did through 2025, the catch-up cycle does not keep pace. The result is a sawtooth pattern: open-weight closes the gap during periods when closed-weight is incrementally improving, then falls behind during periods when closed-weight steps up sharply. 2025 was a sharp-step period.

The third: frontier closed-weight labs have economic incentives that open-weight releases do not match. Closed-weight frontier development is a multi-billion-dollar investment that monetises through API access. Open-weight releases are typically published as research artefacts or as marketing investments. The economic model that funds closed-weight frontier development sustains a larger and faster R&D loop than the economic model that produces open-weight releases. That economic asymmetry is unlikely to resolve in either direction within 12–24 months.

For strategic AI planning, this reversal carries implications.

Plans built around "open-weight will reach parity with closed-weight by 2026" need updating. The data does not support the 2026 parity scenario. Plans that committed to open-weight-first deployment on the expectation of parity should evaluate whether the 3.3% capability gap matters for the workloads. For some workloads (well-bounded enterprise tasks, fine-tunable use cases), the gap is irrelevant and the cost advantages of open-weight deployment outweigh the capability difference. For other workloads (agentic, multi-step, frontier-reasoning), the 3.3% gap matters and may grow.

Multi-tier AI deployment strategies, open-weight for some workloads and closed-weight for others, remain viable but need clearer trigger criteria. The triggers should be workload-specific: "is the workload at the frontier of reasoning capability, requiring the most recent closed-weight model?" versus "is the workload bounded enough that open-weight at 3–5% lower capability is sufficient?" Frameworks that handle this trigger-criterion logic systematically will out-perform frameworks that try to commit to either tier as default.

The deeper structural point: the closed-vs-open dynamic is not a transient phase that resolves in either direction. It is a stable structural feature of the AI capability ecosystem in which closed-weight labs operate at the frontier and open-weight releases follow with predictable lag. The lag varies, sometimes closing (2024) and sometimes widening (2025), but the closed-weight frontier maintains its leading position because the economics of frontier development favour closed-weight monetisation.

The trajectory: the open-weight gap will continue to oscillate in the single-digit percentage range. Strategic plans built on either "open-weight will reach parity" or "the gap will widen indefinitely" are both anchored on extrapolations the data does not support. The defensible planning anchor is "single-digit gap, oscillating, with closed-weight maintaining a structural lead at the absolute frontier." Plans that build for that reality, with explicit trigger criteria for when each tier is appropriate, will continue to capture value in both tiers.

For executives setting AI direction: the open-vs-closed conversation should be framed as a portfolio question, not a phase question. Both tiers will be present in the procurement landscape for the foreseeable future. The strategic plan should specify which workloads sit in which tier and what triggers the move from one tier to the other. By 2028, the closed-weight frontier will sit further ahead on agentic and reasoning-specific capability; the open-weight tier will sit further ahead on cost, fine-tunability, and data-residency control. Plan for both axes.


Sources

  • Primary: Stanford AI Index 2026, Chapter 2 (Technical Performance) 2.1 — hai.stanford.edu/ai-index/2026
  • Leaderboard data: Chatbot Arena — historical Elo, Style Control On, top closed vs. top open-weight models
  • Open-weight release context: Llama 4, Mixtral, DeepSeek-V3, Qwen 3 technical releases and weight publication