Skip to main content
AI Data

How Speech Data Is Collected for Low-Resource Languages

July 2026 · 9 min read · Updated September 2026

Short answer. Speech data for low-resource languages is collected rather than found, because no large public corpus exists to scrape. The work is field operations: recruit and verify native speakers stratified by dialect, age, gender and region; design prompts that elicit read and spontaneous speech; record across real device and acoustic conditions; transcribe against a written convention agreed in advance; and verify with native reviewers measuring error rate and agreement per dialect. Consent and fair pay are part of the method, not an afterthought.

Key takeaways

  • Low-resource speech corpora cannot be scraped from the web; every usable hour is produced deliberately through field recruitment and recording.
  • Coverage should be planned across five axes — dialect, speaker demographics, speech type, acoustic condition and device — before any recording starts.
  • Word error rate and inter-transcriber agreement should be reported per dialect and per acoustic condition, not as a single language-level average.
  • A transcription convention covering orthography, disfluencies, code-switching and uncertainty marking must be agreed in writing before transcription begins.
  • Informed consent, a traceable withdrawal path, and PII redaction are required for spontaneous speech, which routinely contains names and personal detail speakers did not intend to make permanent.

Why is low-resource speech data collection different from high-resource languages?

There is little or nothing to scrape, dialect variation is often larger, and orthography may still be unsettled — so every corpus has to be produced deliberately rather than assembled from existing media.

High-resource languages have decades of broadcast, subtitled media and public corpora to draw on. Low-resource languages frequently have none of that in usable quantity. A low-resource language is one without enough existing digitised text, audio or annotated data for a model to learn it well from public sources alone, which means collection has to substitute for scraping entirely.

Orthography can also be unstable: some languages have competing spelling conventions, recent standardisation, or predominantly oral use, and transcription cannot begin until a convention is chosen and written down — a decision that affects every downstream quality metric. Dialect distance is often larger than in high-resource languages, so varieties a language-level plan treats as one thing may differ enough that a model trained on one performs poorly on another. Code-switching, where speakers mix languages within a single sentence, is normal rather than exceptional in many multilingual regions, and a corpus that excludes it excludes how people actually talk. Underlying all of this is a recruiting problem: finding, verifying, retaining and fairly paying speakers in a language a team does not currently cover is field work, and it is the single largest determinant of whether a programme delivers, a challenge covered in more depth when recruiting native contributors for African language data.

How is coverage designed before recording begins?

Coverage is planned in advance across five stratification axes, each with a numeric target, rather than left to whoever happens to be recorded.

Axis Specify Why
Dialect and regional variety Named varieties with a target share each The most common gap; invisible in a language-level plan
Speaker demographics Age bands, gender, urban/rural, education mix Determines who the model works badly for
Speech type Read, scripted, spontaneous, conversational Read-only corpora produce models that fail on natural speech
Acoustic condition Quiet, ambient noise, outdoor, in-vehicle, telephony Must match deployment conditions
Device Phone models common in the market, headset, far-field Microphone characteristics shape the signal materially

Coverage ratio is hours collected in a stratum divided by hours targeted in that stratum. The minimum across strata should be tracked and reported alongside the total, since a programme at 90% of total hours can still sit at 20% on a dialect that carries a third of the market. The comparison between high-resource and low-resource languages in AI training sets out why this planning step matters more the further a language sits from the high-resource end of that spectrum.

What prompts and scripts are used to collect speech?

A corpus generally needs three material types, each contributing something the others cannot: read speech for clean phonetic coverage, elicited spontaneous speech for naturalness, and conversational speech for real dialogue dynamics.

Read speech, where speakers read prepared sentences, is the cheapest and most controllable option and gives clean phonetic coverage, but it does not sound like conversation. Elicited spontaneous speech — a prompt or scenario spoken in the speaker's own words — is more expensive to transcribe and far closer to deployment reality. Conversational speech, with two or more speakers, overlaps, interruptions and repairs, is hardest to transcribe and essential for anything handling real dialogue. Prompts should be written in the target language culturally rather than translated from English, since translated prompts produce sentences native speakers would never say and are audible as awkward reading in the resulting data. The prompt set should also be checked explicitly for phonetic coverage — including sounds that are rare but distinctive — since in a language without existing corpora nothing else will surface that gap.

How are native speakers recruited and verified?

Speakers are recruited in-market and screened by a native speaker to confirm they use the specific regional variety a project needs, not just the language in general.

Verifying the variety rather than only the language matters because a speaker who says they speak the language may use a different regional variety than the one targeted, and skipping that short screening corrupts the stratification silently. Diaspora speakers are valuable contributors in general but their speech drifts over time — vocabulary, borrowings and register shift away from current in-market usage — so recruiting in-market is preferred where the deployment is also in-market. Retention across multiple sessions needs planning where a design calls for repeat recordings per speaker. Fair, transparent pay at rates appropriate to the local market, with terms the speaker can read in their own language, is both an ethical requirement and a practical one: underpaid collection produces rushed sessions, drop-out, and a speaker pool that will not return for the next programme, a point developed further in how data contributors should be consented and paid.

What recording conditions should a low-resource speech corpus include?

Recordings need to match the noise, devices and audio bandwidth of the deployment environment, not the quiet room where recording is easiest.

A model trained entirely on quiet-room recordings degrades in the environments where it will actually run, so a corpus should include the noise conditions of the use case — street, market, vehicle, home with background sound — and the devices people actually use in that market, which may differ from the devices a project team uses. Telephony-band audio should be included if the model will meet it, since bandwidth reduction is not something a wideband-trained model handles by default. Metadata should be recorded per session — device, environment, speaker stratum, session conditions — because missing metadata is the reason a corpus cannot be re-balanced later. Related considerations for speech captured outside studio conditions are covered in collecting accented and non-native speech for robust voice AI.

What transcription conventions need to be agreed before work starts?

A written transcription convention — covering orthography, numerals, disfluencies, code-switching, non-speech events, speaker labels and uncertainty marking — has to be fixed before the first file is transcribed, because it governs every quality figure produced afterwards.

The convention needs to state which spelling standard applies and what happens to words with no standard spelling, whether numerals and dates are written as spoken or as digits, and how filled pauses, false starts and repetitions are handled — transcribed, tagged, or dropped. It should specify how mixed-language spans are marked, how laughter, noise and overlapping speech are recorded, and how speakers are labelled in multi-party audio. The convention also needs a defined way for a transcriber to mark unintelligible audio rather than guess, because a convention with no way to say "I could not hear this" produces confident wrong transcriptions — more damaging than gaps, since they train the model on invented text. Diarization and timestamping conventions for multi-speaker audio are covered in more detail in speech and audio annotation: transcription, diarization and timestamping.

How is transcription quality verified?

Quality is measured with word error rate (WER) — the proportion of substituted, inserted and deleted words relative to a reference transcript — reported per dialect and acoustic condition, alongside inter-transcriber agreement on a double-transcribed subset.

WER is calculated as substitutions plus insertions plus deletions, divided by the number of reference words. Reporting it per dialect and per acoustic condition matters because an aggregate, language-level figure is dominated by the easiest stratum and hides weak performance elsewhere. Measuring inter-transcriber agreement on a subset that two transcribers each work independently gives the earliest and cheapest signal that a convention is ambiguous and needs revision. A gold set built natively in the language, refreshed periodically and injected into live work at a defined rate, keeps quality measured continuously rather than only at delivery, and native reviewers should listen to audio rather than only reading transcripts, since a wrong dialect, an unnatural reading, or a non-native speaker is inaudible on the page and obvious in the recording. The programme mechanics behind this verification loop are set out in how a speech data collection programme actually runs.

How does Lifewood approach low-resource language speech collection?

Lifewood treats low-resource language speech as a specialism delivered through in-market field operations rather than as an extension of general speech work, combining bespoke recording with multilingual transcription and phonetic labelling.

The operating model depends on a physical footprint of 40+ delivery centres across 30+ countries, spanning China, the Philippines, Malaysia, India and Bangladesh alongside Africa, Europe and North America, supporting 100+ languages with 56,000+ registered contributors. That footprint is what makes in-market recruitment and verification possible rather than nominal, which is the step that decides whether dialect stratification in a plan is real once recording starts. Lifewood has run multilingual data operations since its founding in 2004, with speech and language engagements spanning voice-AI developers; buyers evaluating a partner for this specifically should also review multilingual data collection and broader AI data services scope alongside the speech specialism itself.

Frequently asked questions

Vendors that do this well run field operations, not data marketplaces: in-market recruitment, dialect-level stratification, and a written transcription convention agreed before work starts. Verify a specific vendor's in-market presence in the target language and dialect, ask for per-dialect word error rate rather than a language-level average, and check how consent and PII redaction are handled before signing.

Because it largely does not exist there in usable quantity. High-resource languages have decades of broadcast, subtitled media and public corpora; many low-resource languages have little recorded material, unstable orthography, or predominantly oral use. Every usable hour has to be produced deliberately through recruitment and recording rather than found.

It depends on the task, acoustic conditions, and whether a multilingual base model already has related-language exposure. Coverage matters more than raw hours: are the dialects, demographics, speech types, devices and noise conditions of the deployment all represented, and in what proportion? A smaller, well-stratified corpus regularly outperforms a larger, skewed one.

Collecting read speech in quiet conditions only. It is the cheapest and cleanest data to produce, and it trains a model that fails on natural, noisy, device-mediated speech — which is all the speech it will actually meet once deployed.

Word error rate against a natively-built gold set, reported per dialect and per acoustic condition rather than as a language-level average, plus inter-transcriber agreement on a double-transcribed subset. Both figures are only comparable if the transcription convention was fixed in writing beforehand.

Informed consent in the speaker's own language, covering intended use including commercial and onward use, retention period, and a withdrawal path traceable to specific files. Spontaneous speech also needs PII detection and redaction, since speakers say names and personal details they did not intend to make permanent.

Sources and further reading

  1. Lifewood — low-resource speech data
  2. Lifewood — multilingual data collection
  3. Lifewood — AI data services

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team