Short answer. Africa holds over 2,000 languages, close to a third of the world's total, yet 88% of them are severely underrepresented or ignored in computational linguistics under the Joshi et al. classification. A 2025 PRISMA review found just 74 published ASR datasets covering 111 African languages — roughly 11,206 hours across five years, fewer than 15% of them reproducible. The gap is not about speaker numbers (Hausa ~80M, Amharic ~60M, Swahili well over 100M): the top 100 NLP languages cover about 96% of world GDP but under 60% of world population, and commercial attention followed the GDP. What changes it is capacity on the continent — 60% of the population is under 25, with about 12 million entering the workforce each year.
There is a statistic about African languages and AI that I keep coming back to, because it reframes the problem every time.
Africa is home to over 2,000 languages, close to a third of all the languages spoken on Earth. And according to the Joshi et al. classification that has become the standard reference in this field, 88% of African languages are either severely underrepresented or completely ignored in computational linguistics.
Not underserved. Not lagging. Ignored.
That gap is not an abstract research problem. It determines whether a farmer in northern Nigeria can use a voice assistant to check crop prices, whether a patient in rural Tanzania can get health information in the language they actually think in, whether a student in Senegal can use the same educational tools available to a student in Lyon. And it is a gap that will not close through better model architecture. It closes when someone goes and collects the data.
Lifewood brought delivery centres online in Africa as part of its recent expansion, alongside AV programme growth across Malaysia and Indonesia and the establishment of a US site. This piece is about why that decision matters operationally, what the continent's actual language landscape looks like, and what changes when the people doing the annotation live where the languages are spoken.
The scale of what is missing
Let me put some numbers around the gap, because the abstraction hides how severe it is.
A systematic literature review published in 2025, following PRISMA methodology and covering research from January 2020 to July 2025, catalogued the entire published landscape of automatic speech recognition for African languages. Across 71 studies that met the inclusion criteria, the researchers found 74 datasets covering 111 languages, totalling approximately 11,206 hours of speech.
Eleven thousand hours. For 111 languages. Across five years of published research.
For comparison, a single well-resourced commercial ASR programme for English might use tens of thousands of hours for one language. The entire published African language speech corpus, spanning more than a hundred languages, amounts to less than what a serious English-language project consumes.
The same review found that fewer than 15% of the studies provided reproducible materials, and that dataset licensing was frequently unclear. So even the small amount of data that exists is often not practically usable by the next team who needs it.
Meanwhile, roughly 814 African languages are classified as endangered. Nigeria alone has 171 languages facing the most severe threat levels, Cameroon has 75, Ivory Coast has 65. These are languages where the window for documentation is closing while the AI systems that could support them are being built without them.
Why the gap persists, and it is not what most people assume
The obvious explanation is that these are small languages and nobody has got around to them yet. That explanation is wrong on both counts.
They are not small. Hausa has roughly 80 million speakers. Yoruba and Igbo each have tens of millions. Amharic has around 60 million. Swahili is a lingua franca across East Africa with well over 100 million speakers including second-language users. These are not obscure languages spoken by isolated communities; they are the primary languages of major economies.
And it is not a queue. The AI4D African Language Program made this point directly: the top 100 languages in NLP cover about 96% of world GDP but fewer than 60% of the world's population. The languages that get attention are the ones attached to commercial return, not the ones with the most speakers. As the programme's authors put it, the low economic interest African languages represent for the companies driving NLP means that the work will have to be taken up by African researchers.
That is the structural reality, and it explains something important about the shape of what has been built so far. The most significant African language resources have come from grassroots and academic initiatives rather than commercial ones:
Masakhane, a grassroots organisation of African technologists creating datasets and models since 2019; the AI4D African Language Program's crowd-sourcing challenges and research fellowships; the African Languages Lab; benchmarks like IrokoBench and AfroLID built by African researchers.
The work has been done by people close to the languages, largely without commercial funding. What has been missing is production capacity: the ability to take a client requirement for 500 hours of verified Wolof speech, or a Yoruba instruction-tuning dataset, or Amharic red-teaming data, and deliver it at commercial scale, on a timeline, with documented quality.
What the continent actually has: the talent argument
Here is the part that gets underweighted in most discussions of African AI data work, which tend to focus on cost.
Africa has 60% of its population under the age of 25, making it the youngest continent by a wide margin, with roughly 12 million young people entering the workforce every year. Mobile penetration sits around 495 million subscribers and rising, with smartphone use climbing steadily. Kenya's technology ecosystem has earned the "Silicon Savannah" label; Nigeria's population of over 220 million supports a substantial technology sector; South Africa has sophisticated financial and professional services markets.
What that means practically for annotation work is a large, young, digitally fluent workforce that is often multilingual by default. A Nigerian annotator may speak Yoruba at home, English professionally, Nigerian Pidgin socially, and understand Hausa or Igbo from regional exposure. That combination is genuinely difficult to find outside the continent, and it is exactly what multilingual annotation programmes need.
It also means something for quality that cost-focused discussions miss entirely. Annotating Yoruba correctly requires knowing Yoruba, but annotating Yoruba well requires knowing how Yoruba is actually spoken in Lagos versus Ibadan, which tonal distinctions carry meaning in which contexts, and when a code-switch into English is natural speech rather than a specification violation. That knowledge does not transfer through a training document.
The labour question, honestly
Any piece about African AI annotation that skips the labour conditions debate is not worth reading, so let me address it directly.
The data annotation industry in East Africa has been the subject of serious criticism, and much of it is warranted. Kenya's draft Artificial Intelligence and Other Emerging Technologies Policy, published for consultation in 2026, addresses data annotation, content moderation and AI quality evaluation roles specifically. It proposes a fair-pay reference framework benchmarked against international rates rather than domestic minimums, and the gap it addresses is stark: reported earnings of roughly $1.46 to $3.74 an hour against $21 to $27 for equivalent United States roles.
The draft also proposes mandatory psychosocial support, written contracts, transparent pay reporting and grievance mechanisms. A 2026 CHI paper drawing on interviews with Kenyan data workers describes a "regime of entrapment" produced by precarious contracts and global labour arbitrage.
This is the context any company operating in Africa is entering, and pretending otherwise would be dishonest. What it means in practice is that the terms on which the work is done are becoming a procurement question rather than a values statement. Buyers subject to EU documentation obligations will increasingly be asked how contributors were treated, not only what was delivered.
There is also a straightforward operational argument alongside the ethical one. In languages where the qualified annotator pool is genuinely small, trained contributors are the scarce asset in the entire supply chain. Fair terms, stable work and progression paths are how a delivery operation retains the ability to deliver in Wolof or Tigrinya next quarter. Attrition in a rare-language programme is not an HR metric; it is a capacity loss that can take months to rebuild.
What having centres on the continent actually changes
This is where I want to be specific, because "we have centres in Africa" can mean anything from a sales office to a functioning delivery operation.
Recruitment reach. For a language like Wolof or Oromo, contributors cannot be sourced through a global job board. They are reached through local networks, community organisations, universities and word of mouth in the regions where the language is spoken. That requires people on the ground who know those networks. It is the difference between a project that fills its contributor quota and one that stalls at 40%.
Dialect and variety coverage. A language is not one thing. Swahili in Nairobi is not Swahili in Dar es Salaam. Hausa spoken in Kano differs from Hausa spoken in Niger. A programme that recruits from one location produces a dataset that teaches a model one variety and treats the others as errors. Distributed recruitment across the actual geography of a language is the only way to get representative coverage, and it requires physical presence.
Field conditions that match deployment. If a voice product will be used on inexpensive phones in noisy environments with intermittent connectivity, the training data should be recorded on inexpensive phones in noisy environments. Collection run remotely through a central studio produces clean audio and a model that fails in the field.
Data residency and consent. Several African jurisdictions have data protection frameworks with cross-border transfer restrictions, and voice data that identifies a speaker frequently qualifies as sensitive personal data. In-region collection and processing is not just operationally convenient; for some client programmes it is a compliance requirement.
Turnaround. A regional hub with reliable power and connectivity supporting field collection in the surrounding area converts an infrastructure problem into a logistics one. Recordings collected offline can be physically transported to a hub with bandwidth rather than pushed over a weak connection at the contributor's expense.
Lifewood's published work in this area includes a long-running supply relationship with a global voice AI company covering multilingual speech and LLM services, specifically expanding voice-AI coverage into low-resource Asian and African languages. That is the shape of the work: not a one-off dataset purchase, but sustained capacity to extend a client's language coverage into places where the data has to be created rather than sourced.
What this looks like over the next few years
The direction is reasonably clear even if the pace is not.
Language coverage is becoming a policy objective rather than only a commercial calculation. African governments and research institutions are funding language technology work directly, which creates a new class of buyer with different requirements: openness, documentation and explicit coverage mandates.
The grassroots research community has produced benchmarks that make African language work measurable in ways it was not five years ago. IrokoBench, AfroLID, AfriWOZ and the Masakhane corpus family mean a client can now evaluate whether a model actually works in Yoruba rather than assuming it does because the language appears on a support list.
And the commercial case is improving as the addressable market grows. Africa's mobile-first economy means voice and text interfaces in local languages carry real commercial return, not just a moral argument.
None of that closes the gap on its own. What closes it is the unglamorous work: finding speakers, obtaining consent, recording, transcribing, verifying, adjudicating and documenting, thousands of times over, in languages where none of this has been done before. That work has to happen where the languages are spoken. Which is, in the end, the whole argument for being there.
Key takeaways
- Africa is home to over 2,000 languages, close to a third of the world's total, yet 88% are severely underrepresented or completely ignored in computational linguistics per the Joshi et al. classification.
- A 2025 PRISMA systematic review found 74 published ASR datasets covering 111 African languages totalling roughly 11,206 hours of speech across five years of research, with fewer than 15% providing reproducible materials.
- Around 814 African languages are classified as endangered, with Nigeria facing threats to 171, Cameroon 75 and Ivory Coast 65.
- The gap is not about speaker numbers: Hausa has around 80 million speakers, Amharic around 60 million, Swahili well over 100 million including second-language users.
- The top 100 languages in NLP cover about 96% of world GDP but fewer than 60% of the world's population, which explains why commercial attention has not followed speaker counts.
- Most significant African language resources came from grassroots and academic work including Masakhane, the AI4D African Language Program, IrokoBench and AfroLID.
- Africa has 60% of its population under 25 and roughly 12 million young people entering the workforce annually, producing a large, often multilingual workforce.
- Kenya's 2026 draft AI policy proposes fair-pay benchmarking against international rates, addressing reported earnings of $1.46 to $3.74 an hour against $21 to $27 for equivalent US roles.
- In-region centres change five things: recruitment reach, dialect coverage across a language's geography, field conditions matching deployment, data residency compliance, and hub-based turnaround.
- Lifewood brought Africa centres online alongside AV expansion in Malaysia and Indonesia and a US site, with published work extending voice-AI coverage into low-resource African languages.
Sources and further reading
- Joshi et al., "The State and Fate of Linguistic Diversity and Inclusion in the NLP World", cited in the African Languages Lab paper for the 88% figure and the 2,000+ language count
- "Automatic Speech Recognition for African Low-Resource Languages: A Systematic Literature Review" (arXiv, 2025), PRISMA review covering 2020 to July 2025
- AI4D African Language Program (arXiv, 2021), on the 96% of world GDP versus 60% of population figure
- NaijaNLP, "A Survey of Nigerian Low-Resource Languages" (arXiv, 2025), on IrokoBench and AfroLID
- Princeton CDH, "African_UD", on Masakhane's founding and grassroots African NLP work
- Introl, "Africa's AI Data Center Boom", on demographic and mobile penetration figures
- ITWeb Africa, "Kenya sets standards for AI workers", on the draft fair-pay framework
- Lifewood, company timeline and case studies