Short answer. Not with a translated benchmark, which is how most multilingual claims are currently evidenced.
Translated tests carry artifacts that models exploit as shortcuts, and they align poorly with what speakers of the language actually judge as good: one comparison found translated benchmarks correlating with local human judgment at a Spearman coefficient of 0.47 against 0.68 for natively constructed ones. Real testing means benchmarks written by speakers in the language, generation tasks rather than multiple choice, per-language reporting, and held-out sets that the model has never seen.
Why isn't a multilingual benchmark score proof of multilingual capability?
Because most multilingual benchmarks were built by translating an English test, so a high score can reflect the translation rather than the language.
A useful framing from recent survey work identifies three core challenges in multilingual evaluation: coverage, meaning which languages are tested at all; representativeness, meaning whether the test reflects the language and culture rather than an English original; and trust, meaning whether the result is scientifically reliable. Translated benchmarks and UScentric framing were named as the two main representativeness deficiencies.
The practical consequence is that "supports 100 languages" and "was evaluated in 100 languages" and "performs well in 100 languages" are three different claims, and the market routinely presents the first as though it were the third.
A Microsoft-affiliated survey of multilingual evaluation found that in every world region a substantial share of languages are evaluated primarily through translated content, with Europe and East Asia the most translation-dependent. Interestingly, Sub-Saharan Africa showed the highest share of natively authored content, a direct result of language-specific benchmarks built from scratch by regional research communities.
What is wrong with translated benchmarks?
Three things: they leave detectable traces of the source language, models exploit those traces as shortcuts, and they disagree with the people who actually speak the language.
Translationese. Translation leaves artifacts, tokens and syntactic structures that let a reader identify the source language.
The phenomenon is well documented in linguistics, and in an evaluation context it is a contaminant rather than a curiosity.
Cue inheritance. Because all non-English items derive from a common English source, model performance may reflect identification of English cues preserved through translation rather than genuine understanding. Earlier work demonstrated exactly this mechanism, showing that translation introduces subtle artifacts models exploit as non-semantic shortcuts. A model can score well by recognising the shape of the original question.
Disagreement with speakers. This is the most decision-relevant finding. Translated benchmarks align far worse with local human judgments than natively constructed alternatives, reported as a Spearman correlation of 0.47 against 0.68. If your benchmark and your users disagree, the benchmark is the thing that is wrong.
Quality of translation varies more than most buyers assume. Global-MMLU, a substantial and well-intentioned effort covering 42 languages, combined machine translation with crowdsourced human verification, but only around 20% of machine-translated texts underwent manual correction, and questions and answers were translated separately, producing observable grammatical inconsistencies in some languages.
The contrast with native authorship is measurable. In one Sinhala benchmark, naturalness of natively written STEM items was rated at 97.3% against 71.1% for the translated equivalent.
Do natively written benchmarks solve the problem?
They fix representativeness and introduce a different problem: comparability. And most of them still test the wrong thing.
Native benchmarks are a genuine advance. Efforts that write questions from real curricula, government exams and local elearning platforms preserve cultural context, correct terminology and authentic framing, and they avoid translationese entirely. Dialect-specific versions extend this further.
Two limitations are worth stating plainly.
Cross-lingual comparison becomes harder. If a German set draws on driving theory and a French set on different topics, a score difference confounds language proficiency with task difficulty. Natively sourced benchmarks are internally valid and awkward to compare across languages, which is precisely what a buyer wants to do.
Multiple choice tests retrieval, not generation. Both translated and native benchmarks overwhelmingly use multiple-choice formats, which measure knowledge retrieval and are subject to selection bias. Real-world utility depends on coherent, fluent, semantically faithful generation, and multiple choice does not assess it. Researchers have been explicit that assessing multilingual generation quality remains an open problem.
The practical reading for a company is that no single public benchmark answers the question. Native benchmarks tell you whether the model handles the language and culture. Generation evaluation tells you whether output is usable. Neither substitutes for the other.
How big a problem is contamination?
Significant for older benchmarks, much smaller for newer ones, and impossible to rule out entirely.
Contamination means evaluation questions appearing in the model's pretraining data, inflating scores through memorisation rather than capability. Large web-crawled corpora make it likely for any benchmark that has been public for a while.
Recent measurement gives a usable picture. Newer benchmarks showed markedly lower contamination rates, reported as 0.0% for MILU, 1.0% for INCLUDE and 1.7% for Global MMLU, with the analysis concluding that benchmark age and prominence are strong predictors of contamination risk. The corollary is uncomfortable: the most cited benchmarks are the most likely to be compromised.
Detection is also unsettled. Different methods measure different things, with surface-form overlap and memorisation-driven performance gaps producing divergent results on the same benchmark, and the authors noting that no single detection approach suffices for multilingual contamination.
Two practical responses follow. First, prefer recent benchmarks and treat long-standing leaderboard scores with more caution than their prominence suggests. Second, and more importantly for a company, maintain a private held-out evaluation set that has never been published. It is the only reliable defence against contamination, and it is also the only test that measures your use case rather than a general one.
What does a serious multilingual evaluation programme look like?
Five layers, run per language, with the results reported separately rather than averaged.
Layer 1: Native knowledge and comprehension. Questions written by speakers in the language, ideally sourced from local curricula or examinations. This tests whether the model knows things in that language rather than in translation.
Layer 2: Generation quality. Open-ended tasks judged by speakers for fluency, coherence and semantic faithfulness.
This is where multiple-choice benchmarks are silent and where users form their impressions.
Layer 3: Task performance on your actual use case. A private set drawn from real work: your domain, your customers' phrasing, your document types. Uncontaminated by construction.
Layer 4: Tone, register and appropriateness. Judged by speakers, because it is not automatable. This is the layer that determines whether output reads as respectful or rude.
Layer 5: Safety and refusal behaviour, per language. Guardrails hold only where they were trained, so a safety evaluation conducted in English tells you almost nothing about exposure elsewhere.
Two disciplines wrap around all five. Report per language, never as an aggregate, because a mean is dominated by the strongest languages in the set. And define what a good answer looks like in each market before you measure, since expectations for directness, length and formality differ and scoring everything against English norms produces confident but wrong conclusions.
Who actually builds these evaluation sets?
People who speak the language, working to a specification. This is the constraint that determines whether a company can evaluate honestly.
Every layer above requires the same input: native speakers writing items, judging outputs, adjudicating disagreements and documenting decisions. There is no automated substitute, and the shortcut of using a model to judge output in a language it handles poorly reproduces the original problem in a new place.
That makes evaluation an operational capability rather than a research one. It needs recruitment, screening, guidelines in the target language, inter-rater agreement tracking and adjudication, in each language, on an ongoing basis, because models change and evaluation sets have to keep pace.
This is a large part of what Lifewood does across its 50+ language capability: producing in-language evaluation and preference data, including tone and appropriateness judgements and per-language red-teaming, through screened native speakers working under a human-in-the-loop model in delivery centres across more than 30 countries. The reason it is distributed is the same reason native benchmarks outperform translated ones. Judgements about a language have to come from people who use it.
The closing point is a commercial one. Any supplier can claim language coverage. A company that can produce per-language evaluation results, on sets authored by speakers, including generation and safety, is making a claim that can be checked. In a market where support lists are cheap, checkable evidence is the differentiator.
Key takeaways
- Multilingual evaluation has three core challenges: coverage, representativeness and trust. Translated benchmarks and US-centric framing are the main representativeness problems.
- "Supports 100 languages", "was evaluated in 100 languages" and "performs well in 100 languages" are three different claims.
- Translated benchmarks leave translationese artifacts that models can exploit as non-semantic shortcuts, scoring on recognition rather than understanding.
- Translated benchmarks correlate with local human judgment at a reported Spearman of 0.47 against 0.68 for natively constructed ones.
- In Global-MMLU, only around 20% of machine-translated text was manually corrected, and questions and answers were translated separately.
- Native item naturalness was rated 97.3% against 71.1% for the translated equivalent in one Sinhala benchmark.
- Native benchmarks fix representativeness but complicate cross-lingual comparison, since different topics per language confound proficiency with difficulty.
- Most benchmarks in both families use multiple choice, which tests retrieval rather than generation and is subject to selection bias.
- Contamination rates were reported at 0.0% for MILU, 1.0% for INCLUDE and 1.7% for Global MMLU, with age and prominence predicting risk.
- No single contamination detection method suffices, so a private held-out set is the only reliable defence.
- A serious programme runs five layers per language: native knowledge, generation quality, private task performance, tone and appropriateness, and safety.
- Every layer requires native speakers, which makes evaluation an operational capability rather than a research exercise.
Sources and further reading
- "Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss", arXiv, on translated benchmark limitations, INCLUDE and Global-MMLU, and the multiple-choice constraint
- "The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks", arXiv, on cue inheritance and the 0.47 against 0.68 correlation finding
- Microsoft Research, "The State and Fate of Multilingual, Contextual Evaluation", on contamination rates and regional translation dependence
- Emergent Mind, "Multilingual MMLU", on natively authored benchmarks and the Sinhala naturalness comparison
- "Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets", arXiv, on Global-MMLU construction and correction rates
- "Multilingual European Language Models: Benchmarking Approaches and Challenges", arXiv, on translationese in benchmark construction
- Lifewood, company overview and delivery network