Short answer. Translation converts words while leaving the underlying knowledge and assumptions unchanged, so a model can answer fluently in a language and still be wrong about the food eaten at a birthday there, the appropriate level of formality, or how a request should be phrased. In the BLEnD benchmark (Myung et al., NeurIPS 2024), models averaged 79.22% on United States everyday knowledge asked in English and 12.18% on Ethiopian everyday knowledge asked in Amharic. That knowledge is rarely written down, so it has to be collected from the people who live it.
Key takeaways
- The BLEnD benchmark measured 52,600 hand-crafted question-and-answer pairs across 16 countries and 13 languages, finding models averaged 79.22% on United States everyday knowledge in English versus 12.18% on Ethiopian everyday knowledge in Amharic.
- Model performance on cultural knowledge tracks how well represented a culture is online, not how difficult its language is to process.
- For low-resource languages, models scored better answering in English than in the local language; for mid-to-high-resource languages the pattern reverses.
- The knowledge that is missing is ordinary and unwritten — everyday food, social norms and local reference points that nobody thought needed recording online.
- Closing the gap requires native speakers authoring and evaluating content in-language, not a better translation pipeline.
Why isn't translation enough?
Translation is a language operation applied to content that was already shaped by another culture, so the words change but the worldview does not. Three categories travel badly: everyday practice, social norms, and reference points.
Ask a model what to bring to a colleague's house for dinner. In English it answers sensibly for a Western context. Translate that into Bengali and you get fluent Bengali advice about wine. Nothing was mistranslated — the response was built on knowledge from somewhere else, and the translation step faithfully carried the assumption across. This is the failure mode that survives every quality check a non-local team can run: the grammar is correct, the terminology is consistent, the register is plausible, and the content is wrong in a way only a reader from that market can see.
| What travels badly | Examples | What a translated answer produces |
|---|---|---|
| Everyday practice | What people eat, wear, play and celebrate | Fluent advice about the wrong food, the wrong gift, the wrong occasion |
| Social norms | Formality, directness, who is addressed how, what counts as a polite refusal | Correct information delivered in a register that reads as rude or absurd |
| Reference points | Legal terms, institutions, payment methods, holidays, units | Confident references to things that do not exist in that market |
Above all three sits a harder layer: what makes an answer helpful, appropriate or rude differs by society, and a model aligned on one society's preferences applies those preferences everywhere. Cultural competence — a model's ability to reflect the knowledge and norms of a specific culture rather than the culture best represented in its training data — is not a knowledge gap that more facts would close; it is a calibration inherited from the alignment data.
How large is the cultural gap, measurably?
The gap is large enough to be the dominant factor in some markets, and it is measurable rather than anecdotal. The BLEnD benchmark — 52,600 question-and-answer pairs across 16 countries and 13 languages, including Amharic, Hausa and Sundanese, hand-crafted by native speakers rather than scraped — was built specifically to test this. Building it that way matters, because a benchmark assembled from web text would only measure the same online sources the models already learned from.
| Measurement | Figure |
|---|---|
| Average score, United States everyday knowledge asked in English | 79.22% |
| Average score, Ethiopian everyday knowledge asked in Amharic | 12.18% |
| Spread across cultures for the best model tested | Up to 57.34 percentage points |
| Coverage | 52,600 pairs, 16 countries, 13 languages |
Two findings underneath the headline are more useful than the headline itself. Performance tracked how well represented a culture is online, not how difficult its language is — the ranking of cultures follows their digital footprint. And for low-resource languages, models answered better in English than in the local language, while for mid-to-high-resource languages the reverse held: models did better when asked in the local language. The implication is precise: for those cultures, the model knows more about the culture through English than through the language that culture actually speaks, meaning whatever it absorbed came from outside descriptions rather than from the community itself.
The cultural coverage gap — the score on a reference culture in its own language minus the score on a target culture in its own language — is the more honest metric to track, since an overall multilingual average is dominated by well-represented cultures and can report a model as broadly capable while it fails completely in a specific market.
What kind of knowledge is actually missing?
The missing knowledge is ordinary: the things everyone in a place knows and nobody writes down. The BLEnD authors point to what people eat at birthday celebrations, the spices they cook with, the instruments young people play, and the sports played at school — common knowledge locally, uncommon in the online sources models learn from.
You cannot fix that by scraping harder or translating more, because nobody thought it needed recording; it exists only in people. This distinguishes cultural competence from two things it is often confused with. It is not language coverage — a model can be fluent in a language and ignorant of the culture that speaks it, which is exactly what the Amharic result shows. And it is not localisation in the production sense of adapting formats, currencies and dates; those are surface conversions applied to content whose substance was decided elsewhere.
Where does the gap show up in a product?
The gap surfaces anywhere a system generates, ranks or judges content against assumptions from a single reference market, and it is hardest to see in the layer meant to catch it.
| Surface | What the cultural gap looks like |
|---|---|
| Assistants and chat | Advice that is fluent, confident and inapplicable — the wine-in-Bengali failure |
| Search and recommendation | Results ranked against assumptions from a different market |
| Content generation | Copy that reads as translated even when the grammar is flawless |
| Evaluation | Green dashboards, because the test set was translated from the reference market |
| Safety and moderation | Norms enforced from one society applied to another, over- or under-blocking |
The evaluation row is the one that keeps the rest hidden: a translated benchmark carries the source culture across with it and can rank models wrongly for that market, so the measurement layer reproduces exactly the error it was installed to catch.
What closes the gap?
Closing the gap takes native speakers producing and judging content in their own language, then verifying each other's work — there is no shortcut, because the input is lived knowledge. Four practices do most of the work.
- Author in-language; do not translate in. Prompts, answers and examples written by people from the culture, not converted from English originals. Translation can bootstrap coverage, but it cannot supply knowledge that was never in the source, a discipline covered further in the guide to multilingual LLM training data quality.
- Collect the mundane deliberately. Everyday practice is the material that is missing, so it has to be asked for explicitly — contributors will not volunteer what they assume everyone knows, which is why programmes built around collecting conversational data across cultures design instruments that go looking for it.
- Evaluate with locally written test sets. Items authored by people from that culture, covering everyday knowledge, tone and appropriateness. A translated benchmark measures translation.
- Keep humans in the loop after launch. Cultural errors read as fluent and correct to anyone who is not from that culture, including automated checks and model-based graders that share the assumptions that produced the error.
What to specify when commissioning this work
- Which cultures, named separately from which languages, since they are not the same list and one language may span several.
- Whether contributors live in the market now, and for how long, since diaspora knowledge drifts, particularly on everyday practice — the concern behind guidance on recruiting native contributors for African language data.
- How everyday-knowledge topics are elicited, and who chose the topic list.
- Whether the evaluation set is authored locally or translated, and who wrote it.
- How disagreement between local reviewers is resolved, since two people from the same market can legitimately differ.
- Consent and fair compensation for contributors, both because it is right and because contributor networks in rare languages cannot be rebuilt once lost.
How does Lifewood address this gap?
Lifewood builds delivery around native-speaker collection and review rather than translation, on the reasoning that cultural knowledge cannot be gathered from outside the community that holds it.
Collection, annotation and evaluation run across 100+ languages including underrepresented dialects, produced by screened native speakers through 40+ delivery centres across 30+ countries, with human review layered over automated checks at a 95%+ accuracy threshold. The network is distributed because cultural knowledge is not portable — it has to be gathered where it lives, and a reviewer working from a translated guideline in another country cannot supply it. See the overview of multilingual data collection and the companion guide on low-resource language speech data collection, and the discussion of how contributors should be consented and paid fairly for this kind of work.