If your procurement framework treats hallucination as a single problem that one vendor solves better than another, the 2026 data demands you re-frame the question. Across 26 leading models on the AA-Omniscience benchmark, hallucination rates span 22–94%. The spread is real. What it means for procurement is genuinely contested.

AA-Omniscience: hallucination rate across 26 leading models:

The benchmark first. AA-Omniscience, published by Artificial Analysis in early 2026, tests factual reliability across 6,000 questions in six domains: law, health, software engineering, mathematics, science, and humanities. Its scoring rewards correct answers, penalises incorrect ones, and applies no penalty for refusing to answer. The design encourages models to acknowledge uncertainty rather than guess. Results: Grok 4.20 Beta 0305 had the lowest hallucination rate at 22%, Claude 4.5 Haiku at 26%, MiMo-V2-Pro at 30%. At the higher end, gpt-oss-20B (high) reached 94% and Gemini 3 Flash reached 92%.

The Hughes Hallucination Evaluation Model (HHEM) leaderboard tells a different story. HHEM measures how often models introduce hallucinations when summarising documents from the CNN/Daily Mail corpus. Top 15 models cluster between 1.8% and 5.4%, with most at 4–5%. A different scale entirely.

HHEM-2.3: hallucination rate on CNN/Daily Mail summarisation:

Both benchmarks are measuring hallucination. Both are real. Neither is wrong. They are measuring different specific failure modes. HHEM tests summarisation faithfulness: does the model add things that weren't in the source? AA-Omniscience tests open-ended knowledge: does the model state things it shouldn't be confident about? A model that scores 1.8% on HHEM might score 70% on AA-Omniscience. Both numbers are true.

For a procurement framework, this surfaces a specific question structure that practitioners often skip: which kind of hallucination does the workload care about?

For workloads built around summarising user-supplied documents (legal contract review, research paper synthesis, customer support ticket triage), HHEM-like benchmarks are the relevant signal. The procurement question is "how well does the model stay faithful to the source documents?" Sub-5% hallucination on HHEM means the model rarely adds material that isn't in the source. For these workloads, the 22–94% range on AA-Omniscience is largely irrelevant: the model isn't being asked open-ended knowledge questions.

For workloads built around model-supplied knowledge (open-ended assistants, RAG systems where retrieval failures hand off to model parametric knowledge, conversational interfaces without source-grounding), AA-Omniscience-like benchmarks are the relevant signal. The procurement question is "how well does the model acknowledge uncertainty when it doesn't know?" The 22% versus 94% range here is significant. The difference between a model that refuses to guess on three-quarters of its uncertain answers and one that guesses confidently on virtually every question.

Most enterprise workloads are mixed. A customer service AI summarises ticket context (HHEM-relevant) and then provides recommendations from product documentation that may or may not cover the specific situation (AA-Omniscience-relevant). A legal AI summarises case law (HHEM-relevant) and answers questions about precedent that aren't fully addressed in the source documents (AA-Omniscience-relevant). The procurement framework needs to evaluate against both dimensions, not pick one as primary.

There is a deeper contested question underneath the data.

One interpretation: hallucination is a single underlying phenomenon, and the benchmarks measure different surface manifestations of it. Better models will improve on both AA-Omniscience and HHEM simultaneously. The 22–94% spread on AA-Omniscience reflects the current state of model engineering, not the underlying limit. As post-training methods improve, the spread will compress and both benchmarks will saturate together.

The other interpretation: hallucination is not a single phenomenon. The summarisation-faithfulness problem (HHEM) and the open-ended knowledge problem (AA-Omniscience) involve different model behaviours. A model that has been trained to be conservative about adding material to summaries may still be aggressive about answering open-ended questions, because those are different decision points in the model's behaviour. Under this interpretation, hallucination will not converge to a single ceiling. Different benchmarks will continue to produce different rankings, and the procurement framework needs to evaluate against the specific failure mode the workload exposes.

The data does not yet distinguish between these interpretations. Both are consistent with the 2026 evidence. The contested question for procurement: which interpretation is correct affects whether hallucination evaluation is a one-benchmark question or a multi-benchmark question, and that distinction matters for evaluation budget, vendor comparison structure, and confidence intervals on procurement decisions.

For practitioners, the defensible procurement framework holds both interpretations open. Evaluate against benchmarks that match the workload's specific failure-mode exposure: HHEM if summarisation, AA-Omniscience if open-ended, both for mixed. Track benchmark scores at the model-version level, not at the vendor level. The model that excels at one workload may not excel at another, even within the same vendor's family. Build internal evaluations on workload-specific tasks where the failure modes appear in the actual deployment context. The vendor that wins your procurement on AA-Omniscience may not win it on your workload's specific hallucination profile, and the procurement framework that captures this distinction will produce better outcomes than the one that treats hallucination as a single number.


Sources