Short answer. Model capability tracks the volume of text a language had online when the model was trained, and that distribution is extremely uneven. On MMLU-ProX, a benchmark that poses the same 11,829 questions in 29 languages, strong models score above 70% in English and around 40% in Swahili — a gap of up to 24.3 points. The output does not look broken: fluency holds up even as accuracy drops, so weaker-language text still reads well while being wrong more often.
Key takeaways
- MMLU-ProX poses the same 11,829 questions in 29 languages across 36 evaluated models, so score differences reflect the model rather than question difficulty.
- Reported gaps between high- and low-resource languages reach 24.3 points, with strong models scoring above 70% in English and around 40% in Swahili.
- Fluency degrades more slowly than accuracy, so weaker-language output reads well while being wrong more often — a failure mode English-only QA cannot see.
- A 29-country CSA Research survey of 8,709 consumers found 76% prefer buying with information in their own language, and 40% will not buy from a site in another language at all.
- A tiered pipeline — measuring capability per language, then scaling review intensity to match — costs less than defaulting weak-model markets to English.
Where does the AI language gap come from?
A model learns the distribution of the text it was trained on, and languages appear in that text in proportion to how much of the language was written down and digitised — not in proportion to how many people speak it. A low-resource language is one with a comparatively small or low-quality digitised text corpus available for training, regardless of its speaker population.
The consequence compounds through the stack:
- Tokenisation. Tokenisers trained predominantly on high-resource text split other languages into more tokens per unit of meaning, which costs context window and money.
- Instruction tuning. Instruction data is overwhelmingly English or translated from English, so the model's sense of what a good answer looks like is calibrated on English conventions.
- Evaluation. Suites were English-first, so regressions elsewhere were not visible during development.
- Retrieval. Fewer and lower-quality in-language sources exist to ground an answer on.
Each layer is individually reasonable. The accumulation is a large capability gap. The operational detail that matters most: fluency and accuracy degrade at different rates. Models learn the shape of a language from relatively little data, so output stays grammatical well past the point where factual reliability has dropped — the fluency-accuracy gap that lets fluent, wrong text pass an English-speaking reviewer's read.
What do the benchmarks show?
Independent evaluations put a number on the gap rather than leaving it as an impression, and the number is large enough to change how a pipeline should be built.
MMLU-ProX extends a reasoning-focused English benchmark into 29 typologically diverse languages using a semi-automatic translation process with expert validation, and each language version contains the same 11,829 questions. Across 36 state-of-the-art models, the authors report disparities of up to 24.3 points between high- and low-resource languages; the work was published at EMNLP 2025. A separate strand of research addresses a related problem: for many widely spoken African languages the gap was simply unmeasured, because no comparable benchmark existed at all before recent work began building one.
| Pipeline layer | Symptom in a low-resource language | What an English-only team sees |
|---|---|---|
| Generation | Higher factual error rate at unchanged fluency | Nothing. The output reads well |
| Tokenisation | More tokens per sentence; truncated context; higher cost | Unexplained cost variance between markets |
| Instruction following | Drift toward English rhetorical conventions | Text that is correct and reads as translated |
| Local knowledge | Wrong or missing market-specific facts, names, regulations | Confident text a local reader immediately distrusts |
| Evaluation | Regressions invisible to the test suite | Green dashboards and complaints from the regional office |
| Retrieval | Fewer, weaker in-language sources available | Thin answers attributed to a thin brief |
Reading down the third column explains why these programmes fail quietly: every symptom is either invisible from headquarters or attributable to something other than model capability. Building multilingual evaluation sets in the target language, rather than translating an English one, is the only way most teams will see this table's left column instead of its right one.
Why does the gap matter commercially?
The markets where models perform worst are frequently the markets where publishing in the local language matters most for conversion, which is the opposite of what a cost-driven rollout assumes. CSA Research's 29-country survey of 8,709 consumers found 76% of online shoppers prefer to buy with information in their own language and 40% will not buy from a site in another language at all.
Preference for one's own language is not weaker in markets with lower English proficiency; it runs the other way. A programme that responds to weak model performance by defaulting those markets to English has chosen the worst-performing commercial option while appearing, on a dashboard, as a quality-conscious decision. The alternative is not publishing unreviewed output either — it is tiering: setting review intensity by measured model capability per language, funded from the savings the same model produces where it is strong. Lifewood's own work on the economics of multilingual data collection covers how that funding trade-off is typically structured.
How do you build a pipeline that accounts for the gap?
The goal is not one uniform process; it is a process whose intensity is set by measured capability per language, which means measuring first.
- Tier the language list by measured capability, not by revenue. Run a small in-language evaluation per target language before committing to a workflow. Three tiers is usually enough: strong, adequate with review, and weak enough to require human authorship or heavy rewriting.
- Build evaluation sets in-language, not translated. A translated test set measures translation and embeds English framing. Items authored by speakers, covering local conventions and regulations, are the only way to see the gap that matters — see how multilingual AI models are benchmarked.
- Ground generation in retrieved in-language sources. Retrieval reduces reliance on what the model memorised, which is precisely what is thin here. Where in-language sources are scarce, retrieve from a strong-language source and translate under explicit review.
- Constrain the output format more tightly where capability is weaker. Templates, controlled terminology and fixed structures reduce the space in which a model can be creatively wrong.
- Put the reviewer in-market, with a rubric. A fluent speaker abroad catches grammar; an in-market reviewer catches register, currency of usage and local factual error, scored against a defined error typology such as MQM.
- Sample inversely to capability. Applying the same 5% sample to English and to a weak-tier language accepts a much higher escape rate in the language you can least afford it in.
- Re-measure when models change. Capability in a given language can move substantially between versions, in either direction, so re-tier at each model change.
What do teams get wrong about this gap?
Three assumptions recur, and none survive contact with the measurements above. Translating from English does not remove the underlying scarcity — it relocates the same data problem into a different translation pair, and pivoting through English tends to flatten culture-specific content, so translation still needs in-market review rather than replacing it. The gap is also not simply closing on its own timeline: multilingual capability improves unevenly, since a language whose digital corpus is not growing benefits comparatively little from a larger training run, and reported gaps persist across frontier models. Most consequential, fluent text is not proof of correct text — fluency and accuracy decouple, and the decoupling is worst exactly where a team is least equipped to verify it, which is why a colleague's spot check is not a substitute for a rubric-scored review by a qualified in-market reviewer. Lifewood's guide on culturally relevant data covers the translation-pivot problem in more detail.
How does Lifewood address this gap?
Everything above is implementable in-house for a handful of languages. It stops scaling somewhere between ten and twenty, because the process needs a qualified in-market reviewer available for every language on every release, plus a maintained in-language evaluation set per language — a staffing and network problem more than a technical one.
Lifewood operates that network as its core business: 100+ languages with region-native reviewers, 40+ delivery centres across 30+ countries, 56,000+ registered contributors, and a dual-layer human-in-the-loop review process held to a 95%+ accuracy threshold. That network is the point rather than a differentiator to skim past: this problem is solved with people in the market, and any vendor answer that skips in-market reviewers is answering a different question. Lifewood's multilingual data collection service and broader AI data services are built around that same tiered, in-market review model.