If your assessment of clinical AI diagnostic reasoning is calibrated against early-generation LLM performance (where general-purpose models matched but did not substantially exceed physician performance on standardised tasks), the 2025 multi-agent and reasoning-model benchmark data shows the gap has changed structurally. Microsoft's AI Diagnostic Orchestrator (MAI-DxO), paired with OpenAI's o3 reasoning model, scored 85.5% accuracy on diagnostically challenging cases from the New England Journal of Medicine. The 21 practising physicians with 5-20 years of clinical experience working under comparable conditions scored approximately 20%. Multi-agent frameworks more broadly have shown diagnostic accuracy gains of 7% to over 60% over single-agent baselines. Two complementary evaluations confirm the magnitude (Brodeur et al. 2025; Microsoft MAI-DxO evaluation 2025; Gorenshtein et al. 2025; Zheng et al. 2025; Liu et al. 2025).

The Brodeur et al. 2025 evaluation provides the granular comparison data. OpenAI's o1-preview reasoning model was tested across multiple clinical reasoning task types:

  • NEJM clinicopathological conferences (n=143 cases): the model included the correct diagnosis in its differential 78% of the time, with 52% top-1 accuracy.
  • NEJM Healer cases (n=80 responses): the model achieved a perfect revised-IDEA score on 78 of 80 cases. By comparison: GPT-4 reached perfect score on 47 of 80; attending physicians 28 of 80; residents 16 of 80.
  • Management reasoning (Gray Matters cases): o1-preview's median score was 86%. GPT-4 only reached 42%. Physicians with access to GPT-4 reached 41%. Physicians using conventional resources reached 34%.
  • Real emergency department cases (n=76): o1 produced diagnoses rated "exact/very close" in 67–83% of cases across three diagnostic stages, surpassing two attending physicians at each stage.
The data observation: across multiple evaluation types (differential diagnosis, structured clinical reasoning, management decision-making, and real-world emergency cases), current reasoning-model AI systems substantially exceed unaided physician performance.

The multi-agent dimension amplifies the gap. The MAI-DxO architecture coordinates multiple specialised AI agents (diagnostician, history-taker, treatment planner) through a structured reasoning protocol. Paired with o3 (the reasoning model successor to o1), the system reached 85.5% on NEJM challenging cases against the 20% physician baseline. Multi-agent frameworks more broadly have shown 7% to over 60% diagnostic accuracy gains over single-agent baselines (Gorenshtein et al. 2025; Zheng et al. 2025; Liu et al. 2025).

The data observation: the diagnostic reasoning gap between current AI systems and unaided clinicians is now structurally meaningful. The gap is not a benchmark artefact; it shows up across multiple independent evaluations, multiple case types, and multiple model families.

Three structural observations follow.

The first observation: the evaluation conditions matter. Each benchmark cited (NEJM clinicopathological conferences, NEJM Healer, Gray Matters management cases, real ED cases with blinded scoring) tests specific cognitive components of clinical reasoning. The systematic AI advantage across all four suggests the capability is general across reasoning tasks, not narrowly task-specific. The physician baseline is also calibrated against expected performance: physicians with 5-20 years of experience working without their usual tools. This is a deliberate research design: the goal is to isolate cognitive reasoning capability, not test combined clinician + tools + colleagues + clinical record performance.

The second observation: real-world clinical integration is the open question. These results suggest that current LLMs have surpassed most existing clinical reasoning benchmarks, but they reflect isolated cognitive evaluations rather than real-world clinical integration. Whether AI-assisted reasoning translates to improved patient outcomes remains an open question requiring prospective trials. The data is rigorous on the cognitive capability question; it does not yet establish that deploying these systems in clinical workflows produces better patient outcomes.

The third observation: the evaluation gap matters for procurement and clinical deployment. The MedAgentBench evaluation (Jiang et al., 2025) of LLM agents in a virtual EHR environment across 300 clinically-derived tasks found the best performing model achieved 69.7% task success rate. The 85.5% diagnostic reasoning performance and the 69.7% EHR task performance reflect different capability dimensions. Procurement and clinical deployment decisions need both: diagnostic reasoning capability for decision-support use cases, and operational task capability for workflow integration.

The data observation about the contested area: while diagnostic reasoning benchmarks show large AI advantages, the NOHARM benchmark found that leading LLMs produced 11.8 to 14.6 severely harmful recommendations per 100 clinical cases, with 76.6% being errors of omission (e.g., failing to recommend a critical test). These findings apply to general-purpose LLMs evaluated on open-ended clinical reasoning tasks, not to narrower task-specific tools driving current adoption. The capability is real; so are the failure modes. Both must inform deployment strategy.

Three implications for organisations setting clinical AI deployment strategy in 2026.

The first implication: diagnostic reasoning capability is now sufficient to materially augment physician performance for hard cases. The NEJM challenging cases represent the right tail of diagnostic difficulty. The AI advantage on these cases is large enough that physician + AI combinations will substantially outperform physician-only diagnosis on hard cases. The clinical implication: hard-case consultation workflows that integrate AI diagnostic reasoning are likely to improve diagnostic accuracy in the cases where accuracy improvement matters most.

The second implication: deployment design needs to engage with the failure modes. The NOHARM benchmark findings about severely harmful recommendations are not a minor caveat. Clinical AI deployment that doesn't include physician review of AI recommendations, structured handling of low-confidence outputs, and systematic monitoring of error patterns will materialise the failure modes. The deployment frameworks that work (ambient AI scribes, article #121; sepsis prediction with clinician oversight, article #124) maintain physician control over clinical decisions. The deployment frameworks that try to automate clinical decisions without physician oversight will produce harm.

The third implication: prospective clinical trial evidence on diagnostic accuracy improvement is the missing layer. The benchmark data shows capability; clinical trials are needed to show outcome improvement. The ARISE Network's January 2026 State of Clinical AI Report found that nearly half of clinical AI studies used exam-style questions rather than real patient data; only 5% used real clinical data. The evidence gap between cognitive benchmarks and clinical outcomes is the methodological frontier. Health systems with strong research infrastructure have a structural opportunity (and arguably obligation) to contribute prospective trial evidence on diagnostic AI in real clinical settings.

The data observation: the 85.5% vs 20% finding is not a one-off benchmark; it reflects a structural capability shift documented across multiple independent evaluations. The implications for clinical AI deployment strategy in 2026 are substantial, and the deployment frameworks that engage with both the capability and the failure modes will be the operationally aligned approaches.