Short answer. A high-resource language is one with large volumes of digitised, labelled and unlabelled data available for training AI. A low-resource language lacks that data regardless of how many people speak it. The distinction is about the supply of usable data, not speaker population, which is why languages with tens of millions of speakers sit near the bottom of the scale while smaller ones sit near the top. The standard reference is the six-class framework proposed by Joshi and colleagues (ACL 2020), which places seven languages in the top class and around 2,191 in the bottom one — and the gap between them widens on its own, because almost every mechanism in the system rewards languages that already have data.
Researchers at Microsoft Research India opened their 2020 paper with a puzzle. Two languages, each the official language of a country, with comparable native-speaker populations — around 29 million and around 18 million. One has roughly 2 million Wikipedia articles; the other about 5,500. One appears in 69 items across the two major linguistic data catalogues; the other in 2. One is served by some of the best machine translation systems available; the other by very few, and poorly. The languages were Dutch and Somali.
Nothing about Somali makes it harder to learn or less expressive. What separates the two is a century of publishing, institutional investment, internet infrastructure and research attention accumulating on one side and not the other. "High-resource" describes that accumulated stock of digitised text, audio, dictionaries, parallel corpora and labelled datasets — an economic and historical fact about a language's environment, not a property of the language. The correction worth making early is that low-resource does not mean minor. Bhojpuri, Javanese, Sylheti, Hausa and Amharic each have tens of millions of speakers and sit well down the scale.
How are languages classified?
Joshi, Santy, Budhiraja, Bali and Choudhury, "The State and Fate of Linguistic Diversity and Inclusion in the NLP World" (ACL 2020), sorts the world's languages into six classes by how much labelled and unlabelled data exists for each. It is the common reference point in multilingual AI research, and the classes are usually given memorable names.
| Class | Name | Languages | Examples | Data position |
|---|---|---|---|---|
| 5 | The Winners | 7 | English, Spanish, German, Japanese, French | Dominant online presence and sustained investment; first to benefit from every advance |
| 4 | The Underdogs | ~18 | Russian, Vietnamese, Korean, Dutch | Plenty of unlabelled data, less labelled data, active research communities |
| 3 | The Rising Stars | ~28 | Indonesian, Ukrainian, Hebrew, Afrikaans | Strong web presence, but underserved by labelled dataset collection |
| 2 | The Hopefuls | ~19 | Zulu, Lao, Maltese, Irish | Small sets of labelled data, usually built by determined communities |
| 1 | The Scraping-Bys | ~222 | Bhojpuri, Cherokee, Fijian, Greenlandic | Some unlabelled text, almost no labelled data |
| 0 | The Left-Behinds | ~2,191 | The long tail | Essentially no digital resources at all |
The numbers underneath make the picture stark. Class 5's seven languages cover around 2.5 billion speakers; Class 0's 2,191 languages — about 88% of all those studied — cover around 1 billion. The bottom class holds the overwhelming majority of the world's languages and roughly 15% of its speakers, with close to nothing to train on.
Read the table downward and one boundary stands out. The distinction between Class 3 and Class 4 is not really web presence — Class 3 languages often have thriving online cultures. It is labelled data. The same is true a rung lower: what separates a Hopeful from a Rising Star is whether anyone has systematically produced annotated, verified datasets in that language.
Why does the gap widen instead of closing?
Because left alone, the distribution concentrates. Four reinforcing loops do the work.
Pretraining amplifies what already exists. Modern models learn largely from large unlabelled corpora. That was supposed to democratise things, and for the middle classes it partly has. But a language with almost no unlabelled text online gains almost nothing from a technique that consumes it. As the original researchers put it, unsupervised pretraining risks making the poor poorer.
Research attention follows resources. Analysis of publications across the main computational linguistics conferences found Class 5 languages ranking consistently in the top two or three, while Class 0 languages ranked on average somewhere between 600th and 1000th. Fewer papers means fewer benchmarks, fewer tools and fewer trained researchers — which means fewer papers.
Commercial incentives compound it. Investment flows toward markets that can pay, and those markets mostly speak Class 4 and Class 5 languages. The languages with the weakest business case are frequently the ones with the greatest need.
Speakers migrate to the served language. The subtlest loop and the most consequential. When technology works in one language and not another, people switch to the one that works. Every switch reduces the digital output of the smaller language, weakening its position further. The tools do not merely reflect language inequality; they accelerate it.
What do the measurements actually show?
Three independent bodies of evidence point the same way.
The typological gap. Comparing the structural features found in Classes 0 to 2 against those in Classes 3 to 5, the Microsoft study identified 549 feature categories out of 1,139 that exist in the under-resourced group and do not appear in the well-resourced one. Models trained on the top classes have never encountered nearly half the structural variety of human language. The consequence is specific: Amharic, the second most spoken Semitic language after Arabic, has nine typological features in that ignored group, while Arabic has none. On a cross-lingual similarity task, English into Arabic produced an error rate of 7.8; English into Amharic produced 60.71. The gap is not explained by difficulty but by what the model was never shown.
The benchmark gap on identical content. MMLU-ProX translates the same 11,829 questions into 29 typologically diverse languages, so a score difference between languages is a difference in the model rather than in the questions. Evaluating 36 state-of-the-art models, including reasoning-enhanced and multilingual-optimised ones, the authors report performance disparities of up to 24.3 points between high- and low-resource languages, with the strongest models scoring above 70% in English and falling to around 40% in Swahili. The work was published at EMNLP 2025.
The measurement gap itself. For many languages the size of the gap was unknown rather than small, because no benchmark existed. A 2024 study created roughly one million human-translated words of new benchmark data across eight low-resource African languages — Amharic, Bambara, Igbo, Sepedi, Shona, Sesotho, Setswana and Tsonga — covering more than 160 million speakers, precisely because standard benchmarks did not exist there.
Coverage gap = Score in your strongest language − Score in the target language
Compute that per language rather than reporting a multilingual average. An aggregate score is dominated by the high-resource languages in the set, and a model can post a strong mean while failing in the language a market actually speaks.
One further finding matters for planning: fluency and accuracy degrade at different rates. Models learn the shape of a language from relatively little data, so output stays grammatical well past the point where its factual reliability has dropped — and a reviewer who does not speak the language sees fluent text and concludes it is fine.
Does the commercial case follow the data?
It runs the other way, which is what makes the gap expensive rather than merely unfair. CSA Research's 29-country survey of 8,709 consumers found that 76% of online shoppers prefer to buy products with information in their own language, and 40% will not buy from a website in another language at all.
So the markets where models are least reliable are frequently the markets where the local language matters most commercially. A programme that answers weak model performance by defaulting those markets to English has chosen the option that performs worst commercially, while appearing on the dashboard as a quality-conscious decision.
What actually moves a language up a class?
Labelled data produced by native speakers. That is the binding constraint for almost every language below Class 4, and it is the one thing scraping cannot supply.
Resource class responds to investment, and there is a recent demonstration. Meta's Omnilingual ASR, released in late 2025, brought speech recognition to more than 1,600 languages, including over 500 never previously served by any such system. A significant part of the training corpus was commissioned specifically, gathered through fieldwork with local organisations that recruited and compensated native speakers, using open-ended prompts so people spoke naturally. Those languages did not rise because the internet changed. They rose because someone paid to go and collect the data.
Research on the classification found the same pattern from the other direction: some of the most neglected languages have small, focused communities working hard on them, while others with millions of speakers — Javanese and Igbo among them — have almost no such support.
Two planning conclusions follow, alongside the per-language measurement above.
- Check the class before you promise the market. Speaker count tells you the size of the opportunity; resource class tells you how much work reaching it takes. Confusing the two is how launch dates slip.
- Budget for data creation, not data acquisition. For Class 4 and 5 languages, data can often be sourced. Below that, it usually has to be produced — a different activity with different timelines and costs.
The operational side of that work is covered elsewhere: collecting speech data in low-resource languages for the recording and transcription pipeline, and multilingual LLM training data quality for how a multilingual corpus is specified and verified.
How Lifewood approaches this
Lifewood's multilingual work is built for the constraint this guide describes: producing labelled, in-language data where the web has not supplied it. That means 50+ languages including underrepresented dialects, collected and verified through 40+ delivery centres across 30+ countries, with native speakers screened before a project begins and human review layered over automated checks at a 95%+ accuracy threshold. The network is distributed because moving a language up a class means having people in the places where it is spoken.
See multilingual data collection, global AI data and low-resource speech data.
Sources and further reading
- Joshi, Santy, Budhiraja, Bali and Choudhury, "The State and Fate of Linguistic Diversity and Inclusion in the NLP World", ACL 2020 — the six-class taxonomy, the Dutch/Somali comparison and the typological analysis.
- MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, EMNLP 2025 (arXiv 2503.10497) — 29 languages, 11,829 items each, 36 models.
- "Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages", arXiv 2412.12417, December 2024 — the eight-language benchmark build.
- Meta AI, "Omnilingual ASR: Advancing Automatic Speech Recognition", 2025.
- CSA Research, "Can't Read, Won't Buy" — 8,709 consumers surveyed across 29 countries, 2020.

