Look at any frontier lab's model launch page in 2026. The capability benchmarks are nearly always reported: MMLU, GPQA, AIME 2025, SWE-bench Verified, plus a half-dozen others. Look at the same launch page for responsible AI benchmarks. Most fields are empty.
That asymmetry is the structural feature of responsible AI evaluation in 2026, and it shapes what procurement teams can actually compare.
The data: across seven major frontier model families (GPT-5.2, Gemini 3, DeepSeek-V3.2, Llama 4 Maverick, Grok 4.1, Claude Opus 4.5, Mistral 3 Large), reporting on the major capability benchmarks is near-universal. MMLU, GPQA, AIME 2025, SWE-bench Verified, and similar all show systematic disclosure across the seven labs.
Reported general capability benchmarks for popular foundation models:
Reporting on the responsible AI benchmark equivalents is sparse to absent. BBQ (2021), measuring fairness. HarmBench (2024), Cybench (2024), StrongREJECT (2024), WMDP (2024), measuring security and dangerous-capability evaluation. SimpleQA (2024), measuring factuality. MakeMePay and MakeMeSay (2024), measuring autonomy and human agency. Across the same seven frontier labs, most cells in the equivalent reporting matrix are empty.
Reported safety and responsible AI benchmarks for popular foundation models:
Only Claude Opus 4.5 reports results on more than two of the responsible AI benchmarks. Only GPT-5.2 reports StrongREJECT. The pattern is not that some labs are RAI-focused and others are not. It is that the asymmetry between capability disclosure and responsible AI disclosure is the norm across the field.
This is not because frontier labs are ignoring responsible AI. All major frontier labs conduct internal evaluations, red-teaming, and alignment testing. The lack of disclosure does not mean lack of effort. It means lack of comparable, externally verifiable disclosure.
For practitioners doing vendor evaluation in 2026, the asymmetry creates a specific problem. The procurement framework for capability evaluation has been refined over the past three years: there are common benchmarks, public reporting, and third-party verification (Arena, Artificial Analysis, Epoch's Benchmarking Hub). The procurement framework for responsible AI evaluation has none of those infrastructure pieces in place yet. A vendor's claim about model safety, fairness, or factuality is mostly self-reported, often qualitative, and rarely tied to specific benchmarks the procurement team can cross-reference.
Three structural causes for the asymmetry are visible in the literature.
The first is measurement difficulty. Fairness and bias are highly context-dependent. A fairness metric that works for a hiring tool may not apply in a clinical diagnostic setting. The capability benchmarks (does the model answer the question correctly?) generalise across deployments in a way that responsible AI benchmarks do not. Building a universally-applicable benchmark for fairness is harder than building one for mathematical reasoning.
The second is incentive misalignment. Public capability benchmarks are part of the competitive landscape: reporting a high score is good marketing. Public responsible AI benchmarks have less marketing value, partly because the cultural expectation is that responsible AI is table-stakes rather than a differentiator, and partly because high responsible AI scores can be seen as exposing the existence of risks the model encountered during training.
The third is verification difficulty. Capability benchmarks are mostly objective: the maths problem has a correct answer. Responsible AI benchmarks include dimensions where the right answer depends on context, values, or evaluation methodology. Third-party verification of "this model is safer than that one" is harder than third-party verification of "this model scores higher on MMLU."
These structural causes are not equally easy to fix. Measurement difficulty will persist. Incentive alignment may shift if procurement frameworks systematically reward responsible AI disclosure. Verification difficulty may resolve gradually as evaluation methodology matures.
The methodological observation: the responsible AI literature is operating at an earlier evolutionary stage than the capability literature. Capability evaluation has shared benchmarks, public reporting, third-party verification, and competitive incentives that drive standardisation. Responsible AI evaluation has none of those pieces in place yet. The lack of disclosure infrastructure is now itself one of the structural features practitioners need to plan around. Treating responsible AI evaluation as a domain where benchmark scores are the primary signal will produce procurement under-allocation to internal evaluation and over-reliance on vendor-supplied claims that lack the verification scaffolding the capability benchmarks have.
This is not a recommendation to ignore the responsible AI benchmarks that do exist. BBQ, HarmBench, StrongREJECT, SimpleQA, and the rest provide meaningful signal where vendors report them. It is a recommendation to recognise that responsible AI procurement evaluation in 2026 requires more weight on internal evaluation, more weight on red-team reports from external organisations where available, and less weight on public benchmark scores than capability procurement evaluation requires.
The literature will continue to develop. New benchmarks will appear. Reporting will become more standardised over the next 24–36 months. But the reporting gap visible in the 2026 data is the starting point for any responsible AI procurement framework being designed today.
Sources
- Primary: Stanford AI Index 2026, Chapter 3 (Responsible AI) 3.2 — hai.stanford.edu/ai-index/2026
- Underlying analysis: AI Index 2026 own analysis — benchmark reporting matrices across seven frontier model families
- Referenced capability benchmarks: MMLU, GPQA, AIME 2025, SWE-bench Verified, MMMU, ARC-AGI-2, FrontierMath, τ-bench, HLE
- Referenced RAI benchmarks: BBQ, HarmBench, Cybench, SimpleQA, Toxic WildChat, StrongREJECT, WMDP, MakeMePay, MakeMeSay
Discussion