Short answer. "We cover Swahili" is not a coverage statement. AfriVoices-KE scoped Kikuyu across five dialects — Kiambu, Murang'a, Nyeri and Kirinyaga, with Kirinyaga subdivided — and Kalenjin across two groupings, Nandi and Kipsigis, together accounting for an estimated 60 to 70% of speakers. That is a documented, bounded claim rather than a language name. Dialect and accent are also different things: dialectal variation is lexical, phonological, prosodic and morphosyntactic, while regional accent is primarily phonetic, and labels derived from geography or ISO codes conflate the two.
Locale Level?
A client asks for Kikuyu speech data. Kikuyu has roughly 8.15 million speakers, an ISO code, and a Wikipedia page. It looks like one thing to procure.
The AfriVoices-KE team, collecting exactly this, adopted a five-dialect framework covering Kiambu, Murang'a, Nyeri and Kirinyaga, with Kirinyaga itself subdivided into Kĩ-Ndia and Gĩ-Gĩchũgũ, in order to capture regional phonological and sociolinguistic variation.
Five varieties, from one line item on a scope document.
And for Kalenjin, a Southern Nilotic cluster with about 6.3 million speakers in the Rift Valley, the same team made the opposite decision. They focused on the two most widely spoken groupings, Nandi and Kipsigis, which together account for an estimated 60 to 70% of all Kalenjin speakers.
That is coverage scoping done well: an explicit decision, with a stated rationale, and a documented percentage of the speaker population reached. Most projects make the same decision implicitly, by recruiting wherever recruitment was easiest, and never write down what they ended up with.
The distinction that determines everything downstream
Before you can scope coverage you need a category system, and there is a distinction in the literature that most procurement conversations skip.
Dialectal variation involves differences in lexical choice, phonology, prosody and morphosyntax. Different words, different grammar.
Regional accent is better characterised as primarily phonetic differences within a shared dialectal lexicon and grammar. Same words, said differently.
Why this matters, in the words of the researchers who made the point: labels derived from geography or ISO codes can conflate lexical and grammatical differences with pronunciation differences, affecting both pooling decisions and the interpretation of model errors.
Read that twice, because it explains a category of project failure. If your dataset labels two varieties as separate because they sound different, but they share a lexicon, you have split data that should have been pooled. If it labels two as the same because they share a region, but they differ grammatically, you have pooled data that should have been split. Either way, when the model underperforms you cannot tell whether the problem is acoustic or linguistic.
A useful third axis sits alongside these. Speech corpora also vary by register, meaning language conditioned by activity or setting, and by affect, meaning emotion expressed through speech. A dataset can have excellent geographic coverage and cover only one register, which produces a model that handles formal speech across every region and conversational speech nowhere.
Why ISO codes are not a coverage plan
The labelling problem in published corpora is well documented and it makes existing data harder to reuse than it appears.
Dialect labels across datasets are variously defined by geopolitical boundaries, by coarse regional groupings, or by ISO-style codes such as apc or ary. The consequence, as the researchers put it, is that this hinders principled pooling and selection of data for low-resource varieties.
The recurring bottlenecks catalogued for Arabic speech technology apply broadly: limited open-access standardisation, uneven coverage of underrepresented varieties, inconsistent sub-dialect labelling, and insufficient documentation of data quality and collection conditions.
There is a further finding that runs against the instinct to maximise data volume. Pooling is not universally beneficial: variation in mutual intelligibility across the Arabic continuum can make indiscriminate mixing less effective than targeted cross-dialect transfer.
More data from adjacent varieties can make a model worse for a target variety. Which means coverage scoping is not "collect as many varieties as budget allows" but "decide which varieties belong together and which do not."
The dialect continuum problem
The hardest scoping case is a continuum, and it is more common than discrete-dialect models suggest.
In a dialect continuum, contiguous settlements speak very similar, mutually intelligible varieties, but more distant settlements of the same continuum speak varieties that are not mutually intelligible. The examples given in the literature include the West Romance, East Slavic and Scandinavian continua, with the Caucasus noted as a particularly complex case.
There is no natural boundary to draw. Any line you draw between "variety A" and "variety B" is a decision you made, not a fact you discovered.
The methodological consequence is stated plainly: to account for variation in a geographic region it is necessary to collect data from a significant number of different locations. And the historical failure mode is equally plain, since homogeneous distributions have traditionally been assumed in most studies.
For a data programme this means the sampling frame is locations, not languages. Three recording sites in a continuum will produce a dataset that represents three points and interpolates nothing.
How coverage actually gets mapped
Four approaches, with different costs and different reliability.
Traditional dialect surveys. Structured elicitation of specific variables across regions. Authoritative and extremely slow: only a handful of comprehensive surveys have been completed in the UK and the US in over a century of research.
App-based large-scale collection. The most impressive recent example collected data on 26 alternations from over 47,000 speakers across more than 4,900 localities in the UK via a mobile phone app. That is a scale traditional survey methods cannot approach.
Corpus-based mapping. Deriving variation from large geolocated text corpora. The critical question is whether it generalises, and there is now evidence that it does: a study comparing 139 lexical dialect maps built from a 1.8 billion word geolocated UK Twitter corpus against the BBC Voices dialect survey found broad alignment between the two sources. That validation licenses corpus-based mapping for general inquiry into regional variation, which matters because it is orders of magnitude cheaper than survey work.
Existing reference works. Ethnologue and comparable references give dialect inventories and speaker estimates.
AfriVoices-KE cited these for its variety frameworks and speaker numbers. They are a starting point rather than a scoping plan, because inventories vary between sources and speaker figures are estimates.
The practical approach for a commercial programme is usually a combination: reference works to enumerate candidate varieties, corpus or app data to check which distinctions are actually live, and local expertise to decide which matter for the deployment.
Coverage decisions have a time dimension
Two processes from dialectology are worth knowing because they mean a coverage map has a shelf life.
Geographical diffusion is the process by which a linguistic feature spreads gradually from one place to another, usually through face-to-face contact between speakers of different dialects.
Dialect levelling is the erosion of differences between local dialects, usually toward a standard variety.
A UK study comparing new survey data against data from the 1950s found that some dialect variables had changed and others had stayed the same across more than sixty years. That is the useful nuance: variation does not simply disappear, but it does move, unevenly.
For a data programme this means that a coverage framework built from a reference work published two decades ago may be describing distinctions that have levelled, and missing ones that have emerged. Older speakers and younger speakers in the same locality may need separate treatment.
What a documented coverage scope looks like
Drawing the practice together, here is what should exist on paper before collection starts.
An enumeration of candidate varieties, with the source for that inventory named.
A stated selection, with the rationale. The AfriVoices-KE Kalenjin decision is the model: two varieties chosen, with the estimated share of speakers covered stated as 60 to 70%.
Speaker population estimates per variety, with the source, so the coverage claim is checkable.
Sampling locations rather than just varieties, particularly in continuum situations.
Explicit pooling rules. Which varieties will be treated as one label in the dataset and which will not, decided on linguistic grounds rather than administrative convenience.
A distinction between dialect and accent labels in the metadata schema, so downstream users can pool or split appropriately.
Register and speaker-factor coverage, not just geography.
A statement of what is out of scope, which is the part that gets omitted and the part a client most needs.
One technical detail worth borrowing from the Voxlect benchmark: when building dialect classification data they excluded audio clips shorter than three seconds as insufficient for robust dialect classification, and discarded samples labelled simply as "British" for lacking specificity on regional varieties such as Scottish. Both are quality decisions that only make sense if you know what granularity you are aiming for, which is another reason to decide the framework first.
Where our own work sits
Declaring the interest: Lifewood's language capability is stated as 50-plus languages and dialects, and that second word carries most of the operational weight.
Two observations from doing this work.
The first is that the variety framework has to be built with people from the region, not selected from a reference work in a capital city. Reference inventories are a starting point and they are frequently coarser than the distinctions speakers themselves make, or occasionally finer, preserving distinctions that have levelled. The people who can tell you which is which are the people who live there, and that is a recruitment question before it is a linguistics question.
The second is that coverage scoping is where most multilingual projects quietly go wrong, because it happens early, it looks administrative, and it is usually settled by whoever wrote the statement of work. A project that specifies "Kikuyu" and recruits in Nairobi will deliver something, and what it delivers will be a particular variety spoken by people who moved to the city, labelled as the language as a whole. Nothing in the delivery statistics will reveal that. It surfaces later as a model that works well in one district and poorly in four.
The remedy is unglamorous: decide the framework explicitly, recruit against locations rather than against a language name, and report coverage per variety rather than in aggregate.
Key takeaways
- AfriVoices-KE adopted a five-dialect framework for Kikuyu covering Kiambu, Murang'a, Nyeri and Kirinyaga, with Kirinyaga subdivided into Kĩ-Ndia and Gĩ-Gĩchũgũ.
- For Kalenjin the same team covered two groupings, Nandi and Kipsigis, accounting for an estimated 60 to 70% of speakers. That is a documented, defensible scoping decision.
- Dialectal variation involves lexical, phonological, prosodic and morphosyntactic differences. Regional accent is primarily phonetic difference within a shared lexicon and grammar.
- Labels derived from geography or ISO codes conflate the two, affecting pooling decisions and making model errors harder to interpret.
- Register and affect are additional axes: a dataset can have full geographic coverage and only one register.
- Dialect labels across published corpora are variously defined by geopolitical boundaries, coarse regional groupings or ISO codes, which hinders principled pooling for low-resource varieties.
- Documented bottlenecks include limited standardisation, uneven coverage, inconsistent sub-dialect labelling and insufficient documentation of collection conditions.
- Pooling is not universally beneficial: variation in mutual intelligibility can make indiscriminate mixing less effective than targeted cross-dialect transfer.
- In a dialect continuum, contiguous settlements are mutually intelligible while distant ones are not, so any boundary is a decision rather than a discovery. Sampling frames should be locations, not languages.
- Traditional dialect surveys are authoritative and slow, with only a handful completed in the UK and US in over a century.
- App-based collection reached over 47,000 speakers across more than 4,900 UK localities on 26 alternations.
- Corpus-based mapping generalises: 139 lexical maps from a 1.8 billion word geolocated Twitter corpus showed broad alignment with the BBC Voices survey.
- Geographical diffusion spreads features through contact and dialect levelling erodes local differences toward a standard, so coverage frameworks have a shelf life.
- A documented scope needs variety enumeration with sources, stated selection with rationale and speaker share, sampling locations, explicit pooling rules, dialect versus accent metadata, register coverage and a statement of what is out of scope.
- Voxlect excluded clips under three seconds as insufficient for dialect classification and discarded samples labelled only "British" for lacking regional specificity.
Sources and further reading
- "AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages", arXiv, on the five-dialect Kikuyu framework, the two-grouping Kalenjin selection covering 60 to 70% of speakers, and speaker population figures
- "Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology", arXiv, on the dialect versus regional accent distinction, ISO and geopolitical labelling problems, register and affect axes, and the finding that pooling is not universally beneficial
- "Detecting linguistic variation with geographic sampling", Journal of Linguistic Geography, on dialect continua, mutual intelligibility across distance and the need to sample many locations
- "Mapping Lexical Dialect Variation in British English Using Twitter", Frontiers in Artificial Intelligence, on the 139 map comparison against BBC Voices, the Leemann et al. app study covering 47,000 speakers in 4,900 localities, and the scarcity of comprehensive traditional surveys
- "Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe", arXiv, on minimum clip duration for dialect classification and the exclusion of insufficiently specific labels
- York English Language Toolkit, "Mapping dialects", on geographical diffusion, dialect levelling and the sixty-year UK comparison finding mixed change across variables
- Lifewood, multilingual and dialect data collection