If your deployment plan for AI in multilingual markets treats "the model supports French/Spanish/Arabic/Mandarin" as a sufficient signal of capability in that language, the 2026 data shows a structural gap in the framing. GPT-5 scored 99.8% on Standard Slovenian commonsense reasoning. The same model scored 88.6% on the Cerkno dialect. Mistral Medium 3.1 fell from 90.0% on standard Slovenian to 53.2% on dialect. Llama 3.3 fell from 87% to 53.6%. The English-vs-other-languages framing obscures the much steeper cliff at the dialect level.
Slovene DIALECT-COPA: accuracy on standard vs dialect:
The benchmark methodology: Slovene DIALECT-COPA tests commonsense reasoning in both Standard Slovenian and the Cerkno dialect. The benchmark was developed because researchers noticed that frontier models perform reasonably on standard varieties of widely-spoken languages but degrade sharply when the same content is presented in regional dialects, sociolinguistic registers, or code-switched forms.
The drop pattern in the data is not uniform. GPT-5 lost approximately 11 percentage points moving from standard to dialect, the smallest drop in the top 12 models. Mistral Medium 3.1 lost approximately 37 percentage points, among the largest. The order of models on the standard variety is not the same as the order on the dialect. A model that ranks third on standard Slovenian may rank tenth on the dialect.
The parallel pattern shows up in other regional benchmarks. HELM Arabic, a regional benchmark developed with Arabic.ai, has Arabic.AI's regionally-developed LLM-X at the top with 0.86 mean score, ahead of Gemini 2.5 Flash (0.82) and GPT-5.1 (0.81). The leading global frontier models are not the leading models for Arabic. Indic LLM Arena, run by AI4Bharat at IIT Madras across more than 20 Indian languages, shows GPT-5.2 leading but with the gap to next-tier models narrower than on English-centric leaderboards.
HELM Arabic: mean score (regional vs global frontier models):
The structural pattern: the rankings that hold on English-centric evaluations do not necessarily hold when benchmarks reflect local usage, dialect, and cultural references. Models trained predominantly on English and a handful of other widely-spoken languages perform best on those varieties and degrade, sometimes sharply, on everything else.
For deployment teams in multilingual markets, this surfaces a specific evaluation gap. The procurement decision based on global benchmark performance (Arena, MMLU, GPQA) assumes the model will perform similarly across the languages the deployment will face. The dialect data shows that assumption is wrong by 10–40 percentage points for many language pairs.
What does this look like in practice? An organisation deploying a customer service AI in Slovenia, evaluating vendors on Arena, will choose the vendor that performs best on English. The deployed system will face customers writing in Cerkno or other regional dialects. The vendor that won the procurement on Arena may rank fifth or tenth on the language the deployment actually faces. The deployment failure mode is invisible during procurement and becomes visible only in production.
The methodological observation: model evaluation for multilingual deployment requires benchmarks that test the specific language varieties the deployment will encounter, not English-centric benchmarks that assume language performance generalises. The global benchmarking infrastructure was built around the assumption that linguistic performance is approximately uniform across the languages a model claims to support. The 2026 data shows this assumption is false at the dialect level, false at the regional-vocabulary level, and false at the cultural-context level.
Three implications for procurement frameworks in multilingual deployments.
The first: vendor evaluation should require dialect-level testing, not just standard-variety testing. A vendor's claim of "supports Spanish" needs to be evaluated against the specific Spanish varieties the deployment will face: Mexican vs. Peninsular vs. Caribbean vs. Andean Spanish are not interchangeable for a model evaluated only on Peninsular. Procurement frameworks that treat "supports Spanish" as a binary procurement criterion are accepting deployment variance the data shows is large.
The second: regionally-developed models often out-perform global frontier models in regional contexts. Arabic.AI's LLM-X leads HELM Arabic. AI4Bharat models are competitive on Indian languages. Spain's ALIA family handles regional Spanish and Catalan better than global frontier models. For deployments where the linguistic context is regional, the procurement question should include regionally-developed vendors alongside the global frontier candidates. The capability evaluation needs to be done at the deployment-context language level, not at the English baseline.
The third: deployment monitoring needs to include dialect-level performance tracking. A model that performs well on standard variety in pre-deployment evaluation may degrade in production as the user population includes more dialect speakers, code-switching, or regional variation than the evaluation captured. Monitoring infrastructure that aggregates across "Spanish" or "Arabic" without distinguishing varieties will miss the failure modes the data shows are common.
For procurement teams operating in multilingual markets, the planning anchor needs to shift. "Does the model support the language?" is not the procurement question. "How does the model perform on the specific varieties of the language the deployment will face?" is. The shift sounds modest. The implementation differences are substantial: evaluation budget, vendor candidate pool, monitoring infrastructure all need to be designed for the dialect-level cliff, not the language-level baseline.
The English-centric AI deployment narrative of 2023–2024 underrepresented the linguistic complexity of global AI deployment. The 2026 data brings the complexity into view. Frameworks that absorb it will produce better deployment outcomes than ones that retain the English-centric simplifying assumption.
Sources
- Primary: Stanford AI Index 2026, Chapter 3 (Responsible AI) 3.7 Fairness and Bias — hai.stanford.edu/ai-index/2026
- Slovene dialect benchmark: Slovene DIALECT-COPA leaderboard, 2026 — commonsense reasoning evaluation across Standard Slovenian and Cerkno dialect
- Arabic benchmark: HELM Arabic — Stanford CRFM with Arabic.ai, 2026
- Indic benchmark: Indic LLM Arena Leaderboard — AI4Bharat at IIT Madras, 20+ Indian languages
Discussion