Skip to main content
AI Data

How to Scope Language Coverage at Locale Level

July 2026 · 10 min read · Updated September 2026

Short answer. "We cover Swahili" is not a coverage statement. AfriVoices-KE scoped Kikuyu across five dialects — Kiambu, Murang'a, Nyeri and Kirinyaga, with Kirinyaga subdivided — and Kalenjin across two groupings, Nandi and Kipsigis, together accounting for an estimated 60 to 70% of speakers. That is a documented, bounded claim rather than a language name. Dialect and accent are also different things, and labels derived from geography or ISO codes conflate the two.

Key takeaways

  • AfriVoices-KE used a five-dialect framework for Kikuyu (Kiambu, Murang'a, Nyeri, Kirinyaga split into Kĩ-Ndia and Gĩ-Gĩchũgũ) and a two-grouping framework for Kalenjin (Nandi and Kipsigis) covering an estimated 60 to 70% of speakers.
  • Dialectal variation involves lexical, phonological, prosodic and morphosyntactic differences; regional accent is primarily phonetic difference within a shared lexicon and grammar.
  • Labels derived from geography or ISO codes conflate dialect and accent, which distorts pooling decisions and makes model errors harder to diagnose.
  • In a dialect continuum, contiguous settlements are mutually intelligible while distant ones are not, so any boundary drawn between varieties is a decision rather than a discovery.
  • Coverage frameworks have a shelf life: geographical diffusion spreads features and dialect levelling erodes local differences over time.

Why isn't a language name enough to define coverage?

A single language name can hide several distinct varieties, so a coverage statement has to name which varieties were reached and roughly what share of speakers they represent, not just the language.

A client asks for Kikuyu speech data. Kikuyu has roughly 8.15 million speakers, an ISO code, and a Wikipedia page. It looks like one thing to procure.

The AfriVoices-KE team, collecting exactly this, adopted a five-dialect framework covering Kiambu, Murang'a, Nyeri and Kirinyaga, with Kirinyaga itself subdivided into Kĩ-Ndia and Gĩ-Gĩchũgũ, in order to capture regional phonological and sociolinguistic variation.

Five varieties, from one line item on a scope document.

And for Kalenjin, a Southern Nilotic cluster with about 6.3 million speakers in the Rift Valley, the same team made the opposite decision. They focused on the two most widely spoken groupings, Nandi and Kipsigis, which together account for an estimated 60 to 70% of all Kalenjin speakers.

That is coverage scoping done well: an explicit decision, with a stated rationale, and a documented percentage of the speaker population reached. Most projects make the same decision implicitly, by recruiting wherever recruitment was easiest, and never write down what they ended up with.

What's the difference between a dialect and a regional accent?

Dialectal variation covers differences in vocabulary, grammar, phonology and prosody — different words and different grammar. Regional accent is primarily a phonetic difference within a shared lexicon and grammar — the same words, said differently. Treating the two as interchangeable is the mistake that undoes a coverage plan later.

Why this matters, in the words of the researchers who made the point: labels derived from geography or ISO codes can conflate lexical and grammatical differences with pronunciation differences, affecting both pooling decisions and the interpretation of model errors.

Read that twice, because it explains a category of project failure. If your dataset labels two varieties as separate because they sound different, but they share a lexicon, you have split data that should have been pooled. If it labels two as the same because they share a region, but they differ grammatically, you have pooled data that should have been split. Either way, when the model underperforms you cannot tell whether the problem is acoustic or linguistic.

A useful third axis sits alongside these. Speech corpora also vary by register, meaning language conditioned by activity or setting, and by affect, meaning emotion expressed through speech. A dataset can have excellent geographic coverage and cover only one register, which produces a model that handles formal speech across every region and conversational speech nowhere.

Why can't ISO codes just define the dialect boundaries?

ISO codes were built to identify languages, not to capture sub-dialect distinctions, so treating an ISO code as a coverage plan hides exactly the variation a coverage plan needs to describe.

The labelling problem in published corpora is well documented and it makes existing data harder to reuse than it appears. Dialect labels across datasets are variously defined by geopolitical boundaries, by coarse regional groupings, or by ISO-style codes such as apc or ary. The consequence, as the researchers put it, is that this hinders principled pooling and selection of data for low-resource varieties.

The recurring bottlenecks catalogued for Arabic speech technology apply broadly: limited open-access standardisation, uneven coverage of underrepresented varieties, inconsistent sub-dialect labelling, and insufficient documentation of data quality and collection conditions.

There is a further finding that runs against the instinct to maximise data volume. Pooling is not universally beneficial: variation in mutual intelligibility across the Arabic continuum can make indiscriminate mixing less effective than targeted cross-dialect transfer, a distinction covered further in our guide to collecting code-switched speech and text.

More data from adjacent varieties can make a model worse for a target variety. Which means coverage scoping is not "collect as many varieties as budget allows" but "decide which varieties belong together and which do not."

How do you scope coverage in a dialect continuum?

A dialect continuum has no natural dividing line, so the sampling frame has to be geographic locations rather than named varieties, with enough sites to interpolate the variation between them.

A dialect continuum is a chain of contiguous settlements speaking very similar, mutually intelligible varieties, where more distant settlements of the same continuum speak varieties that are not mutually intelligible with each other. The examples given in the literature include the West Romance, East Slavic and Scandinavian continua, with the Caucasus noted as a particularly complex case.

There is no natural boundary to draw. Any line drawn between "variety A" and "variety B" is a decision made, not a fact discovered.

The methodological consequence is stated plainly: to account for variation in a geographic region it is necessary to collect data from a significant number of different locations. And the historical failure mode is equally plain, since homogeneous distributions have traditionally been assumed in most studies.

For a data programme this means the sampling frame is locations, not languages. Three recording sites in a continuum will produce a dataset that represents three points and interpolates nothing — a problem our overview of low-resource language speech data collection covers in more depth.

How is dialect coverage actually mapped?

Coverage gets mapped through a mix of traditional surveys, large-scale app or corpus data, and existing reference works, each trading cost against reliability differently.

Traditional dialect surveys use structured elicitation of specific variables across regions. They are authoritative and extremely slow: only a handful of comprehensive surveys have been completed in the UK and the US in over a century of research.

App-based large-scale collection can move much faster. The most impressive recent example collected data on 26 alternations from over 47,000 speakers across more than 4,900 localities in the UK via a mobile phone app — a scale traditional survey methods cannot approach.

Corpus-based mapping derives variation from large geolocated text corpora. The critical question is whether it generalises, and there is now evidence that it does: a study comparing 139 lexical dialect maps built from a 1.8 billion word geolocated UK Twitter corpus against the BBC Voices dialect survey found broad alignment between the two sources. That validation licenses corpus-based mapping for general inquiry into regional variation, which matters because it is orders of magnitude cheaper than survey work.

Existing reference works such as Ethnologue and comparable references give dialect inventories and speaker estimates. AfriVoices-KE cited these for its variety frameworks and speaker numbers. They are a starting point rather than a scoping plan, because inventories vary between sources and speaker figures are estimates.

The practical approach for a commercial programme is usually a combination: reference works to enumerate candidate varieties, corpus or app data to check which distinctions are actually live, and local expertise, of the kind discussed in our guide to recruiting native contributors for African language data, to decide which matter for the deployment.

Does a coverage framework expire?

Yes. Dialect boundaries shift over time as features spread between areas and local differences erode toward a standard variety, so a coverage map built from an older reference source can already be out of date.

Geographical diffusion is the process by which a linguistic feature spreads gradually from one place to another, usually through face-to-face contact between speakers of different dialects. Dialect levelling is the erosion of differences between local dialects, usually toward a standard variety.

A UK study comparing new survey data against data from the 1950s found that some dialect variables had changed and others had stayed the same across more than sixty years. That is the useful nuance: variation does not simply disappear, but it does move, unevenly.

For a data programme this means that a coverage framework built from a reference work published two decades ago may be describing distinctions that have levelled, and missing ones that have emerged. Older speakers and younger speakers in the same locality may need separate treatment, a point that also applies to work on collecting accented and non-native speech.

What should a documented coverage scope include?

A documented scope should exist on paper before collection starts, and it should name the varieties covered, the sources used to identify them, and what was deliberately left out.

An enumeration of candidate varieties, with the source for that inventory named, comes first. A stated selection follows, with the rationale: the AfriVoices-KE Kalenjin decision is the model, two varieties chosen, with the estimated share of speakers covered stated as 60 to 70%.

Speaker population estimates per variety, with the source, keep the coverage claim checkable. Sampling locations rather than just varieties matter particularly in continuum situations, and explicit pooling rules should state which varieties will be treated as one label in the dataset and which will not, decided on linguistic grounds rather than administrative convenience.

A distinction between dialect and accent labels in the metadata schema lets downstream users pool or split appropriately. Register and speaker-factor coverage should be recorded, not just geography. And a statement of what is out of scope is the part that gets omitted most often, and the part a client most needs.

One technical detail worth borrowing from the Voxlect benchmark: when building dialect classification data the researchers excluded audio clips shorter than three seconds as insufficient for robust dialect classification, and discarded samples labelled simply as "British" for lacking specificity on regional varieties such as Scottish. Both are quality decisions that only make sense once the target granularity is decided first.

How does Lifewood approach locale-level coverage?

Lifewood treats coverage scoping as a linguistics decision made with people from the region, then documents it the way AfriVoices-KE documented Kikuyu and Kalenjin: named varieties, a stated rationale, and a checkable share of speakers reached.

Lifewood's language capability spans 50+ languages, and the dialect and accent distinctions inside each one carry most of the operational weight. Two observations from doing this work follow.

The first is that the variety framework has to be built with people from the region, not selected from a reference work in a capital city. Reference inventories are a starting point and they are frequently coarser than the distinctions speakers themselves make, or occasionally finer, preserving distinctions that have levelled. The people who can tell you which is which are the people who live there, and that is a recruitment question before it is a linguistics question, one our piece on Africa's role in AI annotation explores further.

The second is that coverage scoping is where most multilingual projects quietly go wrong, because it happens early, it looks administrative, and it is usually settled by whoever wrote the statement of work. A project that specifies "Kikuyu" and recruits in Nairobi will deliver something, and what it delivers will be a particular variety spoken by people who moved to the city, labelled as the language as a whole. Nothing in the delivery statistics will reveal that. It surfaces later as a model that works well in one district and poorly in four. Buyers comparing multilingual data collection companies should ask each vendor to show this kind of documented scope rather than a language list.

The remedy is unglamorous: decide the framework explicitly, recruit against locations rather than against a language name, and report coverage per variety rather than in aggregate. Programmes that need this depth across many languages at once are usually run as part of a broader multilingual data collection engagement rather than as a one-off.

Frequently asked questions

Dialectal variation involves differences in vocabulary, grammar, phonology and prosody. Regional accent is primarily phonetic difference within a shared lexicon and grammar. Conflating them causes data to be pooled or split incorrectly.

Because they were not designed to capture sub-dialect distinctions, and published corpora label varieties inconsistently by geopolitical boundary, coarse region or ISO code. This makes principled pooling and selection difficult, particularly for low-resource varieties.

Not necessarily. Pooling is not universally beneficial, and mutual intelligibility varies. The defensible approach is an explicit selection with a stated rationale and speaker share covered, as AfriVoices-KE did with two Kalenjin varieties covering 60 to 70% of speakers.

By sampling locations rather than named varieties, since contiguous settlements are mutually intelligible and distant ones are not, and any boundary drawn is a decision rather than a discovery.

Increasingly yes. A comparison of 139 lexical maps derived from a 1.8 billion word geolocated Twitter corpus against the BBC Voices survey found broad alignment, which supports corpus-based mapping for general inquiry.

Effectively yes. Geographical diffusion spreads features and dialect levelling erodes local differences, so a framework built from an older reference work may describe distinctions that have levelled and miss ones that have emerged.

Sources and further reading

  1. "AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages", arXiv, on the five-dialect Kikuyu framework, the two-grouping Kalenjin selection covering 60 to 70% of speakers, and speaker population figures
  2. "Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology", arXiv, on the dialect versus regional accent distinction, ISO and geopolitical labelling problems, register and affect axes, and the finding that pooling is not universally beneficial
  3. "Detecting linguistic variation with geographic sampling", Journal of Linguistic Geography, on dialect continua, mutual intelligibility across distance and the need to sample many locations
  4. "Mapping Lexical Dialect Variation in British English Using Twitter", Frontiers in Artificial Intelligence, on the 139 map comparison against BBC Voices, the Leemann et al. app study covering 47,000 speakers in 4,900 localities, and the scarcity of comprehensive traditional surveys
  5. "Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe", arXiv, on minimum clip duration for dialect classification and the exclusion of insufficiently specific labels
  6. York English Language Toolkit, "Mapping dialects", on geographical diffusion, dialect levelling and the sixty-year UK comparison finding mixed change across variables

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team