Short answer. Multilingual LLM training data is provided by three kinds of supplier — large general annotation vendors, specialist language-data companies, and managed multilingual providers with in-market delivery centres — and quality is ensured by four mechanisms rather than by any single check: coverage designed before collection (languages, dialects, domains, demographics, stratified deliberately), native-speaker production and review in-market, chance-corrected agreement measured per language rather than in aggregate, and provenance and consent recorded per item. Aggregate quality figures are the main way weak multilingual corpora hide their weakness: an average dominated by English says nothing about the language where your model will actually fail.
Foundation models are increasingly judged on their worst supported language, not their best. That is where the complaints come from, where regulators look, and where the gap between a model that works and a model that embarrasses its owner is widest. The data behind those languages is usually the thinnest part of the corpus and the least examined part of the procurement.
This guide covers what makes a multilingual corpus good, how the quality is actually verified, and what to require from a supplier.
Why multilingual corpora fail
Five failure modes account for most of it. None is exotic, and all are cheap to prevent and expensive to fix after training.
Translated English. A corpus built by machine-translating English data carries English discourse structure, English cultural assumptions and English-shaped questions into every language. Models trained on it answer the English question in another language. It is fast, cheap, and produces exactly the fluent-but-foreign quality that native speakers detect immediately.
Dialect and register collapse. "Arabic" is not one variety, and a corpus built entirely from Modern Standard Arabic will fail on the spoken forms most users actually write. The same applies to Chinese regional varieties, Spanish across the Americas, and any language with a wide formal/informal split.
Domain skew. Web-scraped multilingual data over-represents news, encyclopaedia and forum text. If your model serves healthcare, finance or industrial support, the vocabulary it needs is the vocabulary least present in the easily-scraped tail.
Script and tokenisation blind spots. Languages with rich morphology, non-Latin scripts, or no whitespace word boundaries consume more tokens per unit of meaning and are more sensitive to normalisation errors. A pipeline built and tested on English will silently mangle some of them.
Aggregate quality reporting. The failure mode that hides the other four. A corpus reporting 97% quality across 40 languages can contain a language at 60% and nobody will see it until users do.
Designing coverage before collecting anything
Coverage is a design decision, not an outcome. Four axes, decided explicitly:
| Axis | What to specify | Why it matters |
|---|---|---|
| Languages | Tiered by commercial priority, with a target volume per tier | Prevents the long tail being whatever was easy to source |
| Varieties within a language | Dialects, regional forms, formal and informal register | The most common gap, and invisible in a language-level plan |
| Domains | Distribution across the subject areas the model serves | Web-scraped data skews to news and encyclopaedia text |
| Speaker and author demographics | Age, gender, region, education mix appropriate to the use case | Determines who the model works badly for |
A useful check during collection, computed per language rather than globally:
Coverage ratio = Items collected in stratum ÷ Items targeted in stratum
Report the minimum coverage ratio across strata alongside the mean. The mean tells you the programme is on schedule; the minimum tells you which language or dialect is going to fail evaluation.
How the four quality mechanisms work
1. Native, in-market production and review
Two distinctions that suppliers routinely blur:
- Native speaker versus fluent speaker. Both are useful; only the first reliably catches register, idiom and the "no one here would say that" class of error.
- In-market versus diaspora. In-market reviewers track current usage, current regulation and current cultural reference. Diaspora reviewers are excellent for many tasks and drift on all three.
Ask for headcount per language, with location — not a supported-language count. For speech and dialogue data, ask about coverage of varieties inside each language, because a Vietnamese capability sourced entirely from one city is not general Vietnamese coverage.
2. Agreement measured per language
The quality number that matters is chance-corrected agreement between independent annotators, computed per language:
Cohen's kappa = (Observed agreement − Expected agreement) ÷ (1 − Expected agreement)
Raw agreement is misleading on unbalanced tasks — two annotators can agree 95% of the time while distinguishing almost nothing. And a global average hides exactly the languages you need to see. Require a per-language table, and treat any language reported only inside an aggregate as unmeasured.
For transcription work, the equivalent per-language figure is word error rate, with the convention set explicitly: what counts as an error for disfluencies, numerals, code-switching and proper nouns differs between suppliers, and a WER figure without a stated convention is not comparable to anything.
3. Gold sets, built per language
A gold set built in English and translated is not a gold set. Each language needs its own reference items, built by native speakers, refreshed periodically, and injected into live work at a known rate so quality is measured continuously rather than at delivery.
Ask three questions: who builds the gold set, how often it refreshes, and what share of production work is gold-injected. A supplier without a gold-set protocol is inspecting output rather than measuring it.
4. Provenance, consent and licensing per item
For LLM data specifically, this is a procurement gate rather than a nicety:
- Source and licence per item, with the right to use it for model training explicitly established rather than assumed.
- Consent for collected speech, image and text contributions, documented, including for onward use.
- PII handling — detection, redaction where required, and a deletion path that can be executed and confirmed.
- Contamination control — deduplication within the corpus and, where possible, screening against public evaluation sets, so your benchmarks measure capability rather than memorisation.
- Synthetic data, labelled as such. Model-generated training data has legitimate uses and different risks; a corpus that mixes it in without labelling makes those risks impossible to manage.
Evaluating a corpus before you train on it
Five checks, all runnable on a sample:
- Per-language sample read by a native speaker who was not involved in production. Ask for a plain judgement: would a competent local writer produce this?
- Translationese detection. Sample items and ask reviewers to guess the source language. If they can, the corpus is translated rather than native.
- Dialect and register distribution against the design targets, not against total volume.
- Domain distribution against the model's intended use, not against what was easiest to collect.
- Duplicate and near-duplicate rate, per language. High duplication is common in low-resource languages, where the available source material is small and the same text circulates widely.
Run these on a paid pilot before committing volume. Several thousand items per priority language, including your hardest, tells you more than any proposal.
What to require from a supplier
| Requirement | Evidence to request |
|---|---|
| Native, in-market production | Headcount per language, with location |
| Per-language quality reporting | Kappa or WER table by language, last quarter |
| Gold-set protocol | Who builds it, refresh cadence, injection rate |
| Coverage design | Stratification plan with minimum coverage ratio per stratum |
| Low-resource sourcing method | How they recruit and validate speakers in a language they do not yet cover, and how long it takes |
| Provenance and consent | Per-item record; licence position for model training |
| Contamination control | Deduplication method; evaluation-set screening |
| Security and residency | Where data is stored and processed; named sub-processors |
Red flags: a single aggregate quality figure; language coverage counted in supported languages; translated gold sets; no answer on how a new low-resource language is sourced; "we can support any language" without a sourcing method behind it.
How Lifewood approaches this
Lifewood supplies multilingual training data through a managed workforce in owned delivery centres rather than an open crowd, which is the model that makes per-language accountability possible — the same reviewers stay with a language long enough for a gold set and an agreement figure to mean something.
The coverage position is structural: 50+ languages, 40+ delivery centres across 30+ countries, and a global pool of 56,788 contributors, with region-native annotators rather than remote approximations, and a specialism in low-resource languages and regional dialects — the part of a corpus most likely to be sourced by translation elsewhere. Scope spans LLM work including RLHF, SFT, data distillation and response evaluation across 50+ languages, multilingual speech transcription and phonetic labelling, and bespoke field collection across geographic and demographic segments. The AI-data heritage runs to 2004, with the current company established in 2018; engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement.
See enterprise LLM training data, multilingual data collection, low-resource speech data, AI data validation and QA process.
Sources and further reading
- Cohen's kappa and Krippendorff's alpha are the standard chance-corrected agreement measures; use them per language rather than reporting raw agreement in aggregate.
- Companion guide: 9 Criteria for Choosing AI Annotation Services — vendor selection across all modalities.
- Lifewood delivery figures (50+ languages, 40+ centres, 30+ countries, 56,788 contributors) are published on lifewood.com.

