Skip to main content
AI Data

Why AI Models Need Data From Multiple Languages

August 2026 · 9 min read · Updated September 2026

Short answer. A model performs reliably only in the languages it genuinely learned, and training mostly on English produces four measurable penalties everywhere else: lower accuracy, higher running costs from token inflation, weaker safety guardrails, and cultural errors no benchmark catches. Translating unsafe prompts into low-resource languages such as Zulu produced harmful responses from GPT-4 roughly 80% of the time in Brown University research. A guardrail that fails in any language is a guardrail anyone can route around with a free translation tool.

Key takeaways

  • Multilingual training data changes three things inside a model: tokenizer efficiency, the internal layer that converts stored knowledge into a specific language, and generalisation to languages the model saw only briefly.
  • Analysis comparing tokenizers found some languages, such as Arabic, needing far more tokens than English for the same text, which raises cost, shrinks usable context, and adds latency — a pattern researchers call the "token tax."
  • Brown University researchers bypassed GPT-4's safety refusals roughly 80% of the time by translating unsafe English prompts into low-resource languages such as Zulu, because safety alignment is trained per language rather than inherited across them.
  • Translation moves words but not register, local reference points, or the assumptions baked into the original text, so it works as a bridge to language coverage but not as a substitute for native-speaker data.
  • Multilingual training increasingly improves English performance too: reinforcement learning on non-English reasoning data has produced accuracy gains that transferred across languages in ways supervised fine-tuning did not.

Why does a model that scores brilliantly in English still fail abroad?

Benchmark scores measure the language the model was trained in, not the language customers use, and capability does not automatically cross a language boundary.

Consider a support assistant that handles refund disputes flawlessly in testing, then ships to customers writing in Indonesian. It answers a question about an instalment plan as though it were a loan default. It responds to polite formal phrasing with a bluntness that reads as rude. Occasionally it replies in English for no reason. Nothing in the launch checklist predicted any of it, because the checklist was written in English and passed in English.

Research on multilingual reasoning finds the same pattern repeatedly: models handle a task well when it is posed in English and degrade when the identical task is expressed in a lower-resource language. The knowledge is often present. What is missing is the ability to reach it reliably through a different language, because the pathways between concepts and words were built almost entirely on English examples. This is one reason buyers increasingly ask what a multilingual AI data collection service includes before signing a training-data contract.

What does multilingual data actually change inside a model?

Multilingual data changes three things at three different depths: how text is compressed into tokens, how stored knowledge is converted into a specific language, and how well the model generalises to languages it saw only briefly.

Tokenizer. The component that breaks text into tokens before a model can process it. Tokenizer efficiency. Text is broken into tokens by a tokenizer, which learns to compress whatever it saw most during training. Feed one a diet of English and it learns compact representations of English and treats everything else as unfamiliar fragments. Multilingual data changes what the tokenizer considers ordinary — and that decision is fixed for the life of the model.

The conversion layer. Studies of how models handle facts across languages suggest knowledge is stored in a largely language-agnostic internal space, and that errors often appear at the final step, when that internal representation is converted back into a specific language. Multilingual training data strengthens exactly that conversion, drawing on the same underlying process as managed multilingual data collection. It is the difference between a model that knows something and a model that can say it correctly in Vietnamese.

Generalisation beyond the training list. Because related languages share structure, exposure to many languages helps a model make better guesses in ones it has seen only briefly. Breadth matters even for languages that never appear in the plan.

Why is running an AI product in another language more expensive?

Models charge and think in tokens, and the same sentence becomes far more tokens in languages the tokenizer was not built around, which raises cost, shrinks context and adds latency with no change in the underlying idea.

Practitioners call this the token tax: the extra tokens, and therefore extra cost, latency and context consumption, required to express the same meaning in a language the tokenizer was not optimised for. Research on tokenizer fairness has found that the same text translated into different languages can produce drastically different token counts — differences of several times over between languages, for the same underlying content.

Consequence Why it happens
Cost rises APIs bill per token in both directions
Context shrinks A document that fits in English may overflow the same context window in Telugu or Japanese
Latency grows There is simply more to process

The uncomfortable part is who pays it. Research examining tokenization across many languages has found the penalty tends to fall hardest on languages spoken in lower-income regions, which means the markets least able to absorb the cost are charged the most for the same intelligence. Better multilingual training data is one of the few structural fixes, because a tokenizer trained on genuinely diverse text encodes that text more efficiently from the start. Teams scoping this trade-off often start from high-resource vs low-resource languages in AI training.

Why do safety guardrails weaken in other languages?

Safety alignment is trained, not inherited, so a guardrail exists only in the languages it was taught in and a model can be strict in English while permissive in a low-resource language at the same time.

Researchers at Brown University showed that taking unsafe English prompts, translating them into low-resource languages such as Zulu with a free translation tool, and sending them to GPT-4 produced harmful responses roughly 80% of the time — against a model that refused the same requests in English. No technical skill was required.

Follow-on research into multilingual jailbreaking finds the picture continuing to evolve: as labs patch the most direct translation attacks, prompts spread across several conversational turns still succeed at meaningfully higher rates in low-resource languages than in English. The strategic point is the one that gets missed. This is not only a problem for speakers of those languages. A guardrail that fails in any language is a guardrail anyone can route around with a translation tool and thirty seconds. Multilingual safety data is a security control, not a diversity gesture, and it is the same discipline behind how to build safety and jailbreak datasets for LLM red teaming.

Can translation just do the job instead?

Translation moves words. It does not move meaning, context, register or intent, and it inherits every assumption baked into the original English. It is genuinely useful as a bridge, and the mistake is treating it as a substitute.

Three things break. Register and politeness get flattened — many languages encode formality, seniority and social distance in grammar, so a translated reply can be technically correct and socially wrong, which in a customer-facing product reads as rudeness. Local reference points disappear — prices, legal terms, document names, payment methods, holidays and units all carry local meaning a translated English answer quietly gets wrong. And the original framing survives, so a model built on English assumptions about how banking, healthcare or family structure works keeps those assumptions and simply expresses them in a new language.

There is also a compounding effect. When translated text is used to train or evaluate, translation errors become training signal, and the model learns a slightly warped version of the language that then looks correct to anyone checking with the same translation tool. Native speakers break that loop. Nothing else does.

Does multilingual data make a model better in English too?

Increasingly yes, for two independent reasons: training in more than one language appears to push a model toward strategies that generalise, and the supply of unused non-English text is now larger than the supply of unused English text.

Cross-lingual generalisation. In a 2025 study, a model trained with reinforcement learning on non-English reasoning data improved not only in that language but substantially on evaluations in several other languages — well beyond what supervised fine-tuning on the same data achieved. Other work finds that reasoning skill, as distinct from factual recall, transfers between languages remarkably well.

Supply. Epoch AI's research on data scaling projects that, at current trends, frontier training runs will fully consume the available stock of quality public human text somewhere between 2026 and 2032, with English being the part mined fastest and therefore exhausted first. Meanwhile most of the world's linguistic output has never been digitised at all — it sits in speech, in local platforms, in messaging, in undocumented dialects.

Put those together and multilingual data stops looking like a cost centre. It is simultaneously the largest untapped reserve of human-generated training data and a route to better general capability.

What should a team building AI actually do?

A team should choose language coverage deliberately before training starts, evaluate every language on its own rather than in an aggregate score, and keep native speakers in the loop rather than relying on translation or automated checks alone.

  • Pick your languages deliberately, and early. Language coverage affects tokenizer design and data mix, both painful to change later. Retrofitting a language after launch costs far more than including it in the plan.
  • Evaluate per language, never in aggregate. A single averaged score hides exactly the failures that matter. Accuracy, refusal behaviour, tone and safety should each be measured language by language, against benchmarks written by speakers rather than translated into their language, an approach detailed in how to build multilingual evaluation sets for LLMs.
  • Red-team in every language you ship in. Given how guardrails degrade, a safety evaluation conducted only in English tells you almost nothing about your actual exposure.
  • Use translation as a bridge, not a foundation. It is a reasonable way to bootstrap coverage, and not a substitute for in-language data when accuracy, safety or brand tone are on the line.
  • Keep speakers of the language in the loop. Automated checks confirm format and completeness. They cannot tell you the model was subtly condescending, used the wrong register for an elder, or invented a legal term that does not exist in that country.

How does Lifewood approach multilingual data collection?

Lifewood's multilingual work sits at the stages where authorship cannot be substituted: collection, annotation, evaluation and red-teaming carried out by native speakers rather than translated in.

Human-in-the-loop. A quality process in which native-speaking people review and correct machine output rather than accepting it automatically. The practical consequence is that a model's behaviour in a given language is judged by someone who speaks it, rather than by a score averaged across a dozen languages. That is a delivery-network question rather than a tooling one: 100+ languages including underrepresented dialects, and 40+ delivery centres across 30+ countries, are what make per-language red-teaming and per-language evaluation practical at the point in a programme where they matter — before launch, not after a safety incident. Programmes of this size draw on 56,000+ registered contributors, and separately, Lifewood's Bangladesh workforce alone logged 414,120 training hours in 2025.

Quality is verified under a human-in-the-loop model against a customer-approved gold set, at a 95%+ accuracy SLA, with per-language reporting rather than an aggregate figure. Teams scoping a programme can review options in top 10 global multilingual AI data collection companies.

Frequently asked questions

There is no universal number; it depends on where the product ships and who uses it. What matters more than the count is that each supported language has real in-language training and evaluation data behind it, rather than being listed as supported on the strength of translation alone.

Not necessarily. Recent research points the other way, finding that multilingual and cross-lingual training can strengthen general reasoning, which then shows up in English performance as well. In one 2025 study, reinforcement learning on non-English reasoning data produced gains that carried over into other languages the model was evaluated on.

The extra tokens — and therefore extra cost, latency and context consumption — required to express the same meaning in a language the tokenizer was not optimised for. Research on tokenizer fairness has found some languages need several times as many tokens as English for equivalent text, depending on the model and language pair.

Because alignment is learned from safety training data, which has historically been overwhelmingly English. Where that data is thin, the guardrail is thin, regardless of how strict the model appears in English. Brown University researchers bypassed GPT-4's refusals roughly 80% of the time simply by translating prompts into low-resource languages.

It helps with volume but not with authenticity. Synthetic text generated by an English-centric model tends to reproduce that model's blind spots in the target language, so it needs native-speaker validation before it can be trusted as training signal.

From deliberate collection: recordings, writing and annotation produced by native speakers, gathered with consent and compensation, then verified in-language. For most languages beyond the top twenty, the material does not exist online in usable volume or quality and has to be created.

Sources and further reading

  1. Low-Resource Languages Jailbreak GPT-4 (Yong, Menghini, Bach, Brown University)
  2. Multilingual jailbreaking of LLMs using low-resource languages
  3. Language Model Tokenizers Introduce Unfairness Between Languages (Petrov et al.)
  4. Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
  5. Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data (Epoch AI)

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team