> ## Content Index
> Fetch the complete content index at: https://aiadoption.org/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI loses half its accuracy in regional dialects — and that's a fairness problem
- URL: https://aiadoption.org/ai-analysis/ai-loses-half-its-accuracy-in-regional-dialects-and-thats-a-fairness-problem-3/
- Published: 2026-09-27T22:45:02.000Z
- Updated: 2026-09-27T22:45:02.000Z
- Description: Leading models lost close to half their accuracy when tested in a regional dialect rather than the standard language. The benchmark scores vendors quote are measured in standard, high-resource language — which not all your customers speak.
- Author: Jassie
- Tags: AI Analysis, AI Ethics, AI Governance, Large Language Models

If your organisation operates in more than one language, or serves customers who do not all speak the standard form of your national language, one 2026 fairness finding belongs in your AI evaluation before you deploy anything customer-facing. Several leading models lost close to half their accuracy when tested in a regional dialect rather than the standard language. The headline benchmark scores that vendors quote are measured in standard, high-resource language. Your customers do not all speak that.

The specific test is illustrative. On a Slovenian commonsense reasoning task, the gap between standard-language and dialect performance was large enough to roughly halve accuracy for several leading models. Slovenian is not an exotic edge case. It is an official EU language. The finding is that the dialect within a language, not just the jump between major languages, is enough to collapse model reliability. The model that reasons competently in textbook Slovenian becomes substantially less reliable in the Slovenian people actually speak.

## Why this is a fairness finding, not just a localisation one

The instinct is to file this under localisation, a translation-quality problem to be solved with more language coverage. It is better understood as fairness and bias, and the distinction matters. A localisation gap is uniform: the model is equally worse for everyone using the second language. A fairness gap is differential: the model is worse specifically for the speakers of the non-standard form, who are frequently the populations already least well served, the regional, rural, minority-dialect and lower-resource-language communities.

Model accuracy, standard Slovenian vs Cerkno dialect: 

This is why the finding sits within the lens of inclusiveness and the global language gap. Models perform best in the languages and dialects most represented in training data, overwhelmingly English and the standard forms of a handful of major languages. Performance degrades along the same axis as data representation, which means the people the model serves worst are systematically the people whose language is least represented online. The bias is structural, inherited from the data distribution, and it tracks existing patterns of digital marginalisation.

*\[CHART NEEDS REVIEW: performance by language resource level, high vs low-resource languages — no single matching source figure identified\]*

## Where this bites in mid-market deployment

The deployments most exposed are the customer-facing ones in multilingual or multi-dialect markets. A support chatbot for an Australian organisation serving recent-migrant communities. A government-services AI in a country with regional language variation. A retail assistant in a market where the standard written language and the spoken dialects diverge. In each, the benchmark score that justified the deployment was measured on a population that does not match the population being served.

The failure mode is quiet, which makes it dangerous. The model does not announce that it is less reliable in dialect. It produces fluent, confident answers that are simply wrong more often. A halving of accuracy that is invisible to a monolingual procurement team will be very visible to the dialect-speaking customer who gets wrong answers at twice the rate of the standard-language customer. The reputational and equity exposure lands on the organisation, not the vendor.

## The evaluation criterion this implies

The practitioner takeaway is a single, testable evaluation rule: **evaluate the model in the language and dialect of the population you will actually serve, not the standard form the benchmark used.**

> Concretely, that means three things. First, build your evaluation set from real inputs in the actual dialects and languages of your user base, not from translated standard-language test cases, because translation into the standard form hides exactly the gap you are trying to measure. Second, measure the accuracy gap between your highest-resource and lowest-resource user languages explicitly, and treat a large gap as a deployment blocker for the affected population rather than a minor localisation backlog item. Third, where the gap cannot be closed, scope the deployment honestly: deploy with confidence for the populations the model serves reliably, and use human-in-the-loop or alternative channels for the populations it does not, rather than deploying uniformly and letting the worst-served group absorb the failure rate.

The vendor will not surface this for you. The published benchmark is measured where the model is strongest. The only way to know how the model performs for your actual users is to test it on their actual language, and the 2026 evidence says that for non-standard dialects and lower-resource languages the result will frequently be far worse than the headline number suggests. For an organisation that takes fairness seriously, that test is not optional. It is the difference between a deployment that serves everyone and one that serves the standard-language majority while quietly failing everyone else.

---

### Sources

- **Primary**: Stanford AI Index 2026, Chapter 3 (Responsible AI), 3.7 (Fairness and Bias). [hai.stanford.edu/ai-index/2026](http://hai.stanford.edu/ai-index/2026?ref=aiadoption.org)
- **Dialect evaluation**: Slovene DIALECT-COPA benchmark (Slobench), Standard Slovenian vs Cerkno dialect, as reported in the AI Index Chapter 3 highlight on inclusiveness and the global language gap.
- **Cross-reference**: Stanford AI Index 2026, Chapter 2 2.2 (Language), multilingual and dialect performance benchmarks.