Short answer. Multilingual LLM training data is provided by three kinds of supplier — large general annotation vendors, specialist language-data companies, and managed multilingual providers with in-market delivery centres — and quality is ensured by four mechanisms rather than by any single check: coverage designed before collection (languages, dialects, domains, demographics, stratified deliberately), native-speaker production and review in-market, chance-corrected agreement measured per language rather than in aggregate, and provenance and consent recorded per item. Aggregate quality figures are the main way weak multilingual corpora hide their weakness: an average dominated by English says nothing about the language where your model will actually fail.
Foundation models are increasingly judged on their worst supported language, not their best. That is where the complaints come from, where regulators look, and where the gap between a model that works and a model that embarrasses its owner is widest. The data behind those languages is usually the thinnest part of the corpus and the least examined part of the procurement.
This guide covers what makes a multilingual corpus good, how the quality is actually verified, and what to require from a supplier.
Key takeaways
- Multilingual training data comes from large general annotation vendors, specialist language-data companies, and managed multilingual providers with in-market delivery centres.
- Five failure modes drive most bad multilingual corpora: translated English, dialect and register collapse, domain skew, script and tokenisation blind spots, and aggregate quality reporting.
- Quality has to be measured per language — Cohen's kappa or Krippendorff's alpha for judgement tasks, word error rate for transcription — because a global average hides the language that will actually fail.
- Coverage is a design decision made before collection starts, across four axes: languages, varieties within a language, domains, and speaker demographics.
- Provenance, consent, licensing and contamination control are procurement gates for LLM data, not optional extras.
Why do multilingual corpora fail?
Five failure modes account for most of it, and none is exotic — all are cheap to prevent and expensive to fix after training.
Translated English. A corpus built by machine-translating English data carries English discourse structure, English cultural assumptions and English-shaped questions into every language. Models trained on it answer the English question in another language. It is fast, cheap, and produces exactly the fluent-but-foreign quality that native speakers detect immediately.
Dialect and register collapse. "Arabic" is not one variety, and a corpus built entirely from Modern Standard Arabic will fail on the spoken forms most users actually write. The same applies to Chinese regional varieties, Spanish across the Americas, and any language with a wide formal/informal split.
Domain skew. Web-scraped multilingual data over-represents news, encyclopaedia and forum text. If your model serves healthcare, finance or industrial support, the vocabulary it needs is the vocabulary least present in the easily-scraped tail.
Script and tokenisation blind spots. Languages with rich morphology, non-Latin scripts, or no whitespace word boundaries consume more tokens per unit of meaning and are more sensitive to normalisation errors. A pipeline built and tested on English will silently mangle some of them.
Aggregate quality reporting. The failure mode that hides the other four. A corpus reporting 97% quality across 40 languages can contain a language at 60%, and nobody will see it until users do. Companion reading on vendor selection across modalities is in 9 Criteria for Choosing AI Annotation Services.
How should coverage be designed before collecting data?
Coverage is a design decision made before collection starts, not an outcome measured afterward. Four axes are decided explicitly.
| Axis | What to specify | Why it matters |
|---|---|---|
| Languages | Tiered by commercial priority, with a target volume per tier | Prevents the long tail being whatever was easy to source |
| Varieties within a language | Dialects, regional forms, formal and informal register | The most common gap, and invisible in a language-level plan |
| Domains | Distribution across the subject areas the model serves | Web-scraped data skews to news and encyclopaedia text |
| Speaker and author demographics | Age, gender, region, education mix appropriate to the use case | Determines who the model works badly for |
A useful check during collection, computed per language rather than globally:
Coverage ratio = Items collected in stratum ÷ Items targeted in stratum
Report the minimum coverage ratio across strata alongside the mean. The mean tells you the programme is on schedule; the minimum tells you which language or dialect is going to fail evaluation. Teams scoping this for the first time can work through how to build multilingual evaluation sets for LLMs alongside the coverage plan, since the two should be designed together.
How do the four quality mechanisms work?
The four mechanisms work together rather than substituting for each other: coverage decides what gets collected, in-market production decides who produces it, per-language agreement decides whether it is measured honestly, and provenance decides whether it can legally be used.
1. Native, in-market production and review
Two distinctions that suppliers routinely blur. A native speaker is someone who acquired the language from childhood and reliably catches register, idiom and the "no one here would say that" class of error — a fluent speaker, useful for many tasks, does not reliably catch these. In-market means the reviewer currently lives in and uses the language day to day, tracking current usage, regulation and cultural reference; a diaspora reviewer is excellent for many tasks but drifts on all three over time.
Ask for headcount per language, with location — not a supported-language count. For speech and dialogue data, ask about coverage of varieties inside each language, because a Vietnamese capability sourced entirely from one city is not general Vietnamese coverage.
2. Agreement measured per language
The quality number that matters is chance-corrected agreement between independent annotators, computed per language:
Cohen's kappa = (Observed agreement − Expected agreement) ÷ (1 − Expected agreement)
Cohen's kappa is a chance-corrected agreement score between two raters that discounts the agreement you would expect from guessing alone. Raw agreement is misleading on unbalanced tasks — two annotators can agree 95% of the time while distinguishing almost nothing — and a global average hides exactly the languages you need to see. Require a per-language table, and treat any language reported only inside an aggregate as unmeasured; a fuller treatment of what the numbers mean is in inter-annotator agreement: Cohen's kappa, Krippendorff's alpha, and what the numbers actually mean.
For transcription work, the equivalent per-language figure is word error rate, with the convention set explicitly: what counts as an error for disfluencies, numerals, code-switching and proper nouns differs between suppliers, and a WER figure without a stated convention is not comparable to anything.
3. Gold sets, built per language
A gold set is a reference collection of correctly labelled items, built natively in the target language, used to check ongoing production against a known-correct answer. A gold set built in English and translated is not a gold set for the target language. Each one needs its own reference items, built by native speakers, refreshed periodically, and injected into live work at a known rate so quality is measured continuously rather than at delivery.
Ask three questions: who builds the gold set, how often it refreshes, and what share of production work is gold-injected. A supplier without a gold-set protocol is inspecting output rather than measuring it.
4. Provenance, consent and licensing per item
For LLM data specifically, this is a procurement gate rather than a nicety:
- Source and licence per item, with the right to use it for model training explicitly established rather than assumed.
- Consent for collected speech, image and text contributions, documented, including for onward use.
- PII handling — detection, redaction where required, and a deletion path that can be executed and confirmed.
- Contamination control — deduplication within the corpus and, where possible, screening against public evaluation sets, so your benchmarks measure capability rather than memorisation.
- Synthetic data, labelled as such. Model-generated training data has legitimate uses and different risks, covered in more depth in is it safe to train AI models on AI-generated data? — a corpus that mixes it in without labelling makes those risks impossible to manage.
How do you evaluate a corpus before training on it?
Five checks run this evaluation, and all of them are runnable on a sample before committing volume.
- Per-language sample read by a native speaker who was not involved in production. Ask for a plain judgement: would a competent local writer produce this?
- Translationese detection. Sample items and ask reviewers to guess the source language. If they can, the corpus is translated rather than native.
- Dialect and register distribution against the design targets, not against total volume.
- Domain distribution against the model's intended use, not against what was easiest to collect.
- Duplicate and near-duplicate rate, per language. High duplication is common in low-resource languages, where the available source material is small and the same text circulates widely.
Run these on a paid pilot before committing volume. Several thousand items per priority language, including your hardest, tells you more than any proposal.
What should you require from a supplier?
A supplier should be able to produce evidence against every mechanism above, not just describe the mechanism in a sales deck.
| Requirement | Evidence to request |
|---|---|
| Native, in-market production | Headcount per language, with location |
| Per-language quality reporting | Kappa or WER table by language, last quarter |
| Gold-set protocol | Who builds it, refresh cadence, injection rate |
| Coverage design | Stratification plan with minimum coverage ratio per stratum |
| Low-resource sourcing method | How they recruit and validate speakers in a language they do not yet cover, and how long it takes |
| Provenance and consent | Per-item record; licence position for model training |
| Contamination control | Deduplication method; evaluation-set screening |
| Security and residency | Where data is stored and processed; named sub-processors |
Red flags: a single aggregate quality figure; language coverage counted in supported languages; translated gold sets; no answer on how a new low-resource language is sourced; "we can support any language" without a sourcing method behind it. Buyers comparing multiple vendors on these points can start from the top multilingual AI training data companies and the top LLM training data companies, then request the evidence above from whichever shortlist survives.
How does Lifewood approach multilingual training data quality?
Lifewood supplies multilingual training data through a managed workforce in owned delivery centres rather than an open crowd, which is the model that makes per-language accountability possible.
The same reviewers stay with a language long enough for a gold set and an agreement figure to mean something. The coverage position is structural: 100+ languages, 40+ delivery centres across 30+ countries, and 56,000+ registered contributors, with region-native annotators rather than remote approximations, and a specialism in low-resource languages and regional dialects — the part of a corpus most likely to be sourced by translation elsewhere. Scope spans LLM work including RLHF, SFT, data distillation and response evaluation across 50+ languages, multilingual speech transcription and phonetic labelling, and bespoke field collection across geographic and demographic segments. The company was founded in 2004, giving the delivery model over two decades to mature, and engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement. More on the underlying services is at enterprise LLM training data and multilingual data collection.