Skip to main content
AI Data

How a Speech Data Collection Programme Actually Runs

Short answer. As a managed pipeline, not a crowdsourced funnel: we design prompts and scripts for the target model, recruit and verify native speakers through our delivery centres, record…

Mumu D. · July 2026 · 8 min read

Download PDF

Short answer. As a managed pipeline, not a crowdsourced funnel: we design prompts and scripts for the target model, recruit and verify native speakers through our delivery centres, record under controlled conditions with session-level checks, and pass every batch through our dual-layer human review before delivery. The field's own audits explain why: loosely controlled community collections show serious quality problems in exactly the low-resource languages where the data matters most.


What does the market now expect from a speech corpus?

Scale with verification. The reference projects of the past two years pair thousands of hours with named speakers, demographic balance and documented review.

The bar moved fast. AfriVoices-KE, one of 2026's landmark collections, targeted roughly 3,000 hours across five Kenyan languages — 750 hours of scripted and 2,250 hours of spontaneous speech — recorded from 4,777 native speakers deliberately spread across regions and demographics. NaijaVoices delivered 1,800 hours of authentic Igbo, Hausa and Yorùbá speech; Mali's Bambara went from a single 30-hour corpus to a 612-hour dataset — a 20x jump — recorded, notably, in a controlled environment with consistent quality control before transcription. At the other end of the spectrum sits found data: MLCommons' Unsupervised People's Speech assembled over a million hours across 89+ languages by mining the web, of which about 736,000 hours survived validation.

The gap between those two ends is the reason managed programs exist. A 2025 audit of widely used multilingual speech datasets found serious quality issues concentrated precisely in low-resource, lessinstitutionalised languages, and noted that even the flagship community-driven platform lacks welldocumented quality control for its source texts and recordings — making reliability "largely unknown".

Volume is now cheap; verified hours, balanced speakers and documented review are what a training team is actually buying. That is the specification our program is built against.

What a reference-grade collection looks like now AfriVoices-KE: spontaneous speech collected 2,250 hrs NaijaVoices: Igbo, Hausa, Yorùbá 1,800 hrs AfriVoices-KE: scripted speech 750 hrs Bambara, 2022–2025: 30 hrs → 612 hrs 20x And why managed beats mined 4,777 named, demographically balanced native speakers behind AfriVoices-KE's ~3,000 hours ~26% of the web-mined million-hour People's Speech corpus was lost at validation (736K hours survived)

"Unknown" how audits describe the reliability of community collections without documented QC — worst in lowresource languages Figures from the AfriVoices-KE, Bambara and MLCommons publications and the 2025 multilingual speech-quality audit, as cited below.


How is a collection designed before anyone records?

The specification does most of the quality work: what the model needs decides the prompts, the speaker mix and the environments — before a microphone is switched on.

Start from the model's diet, not the language's dictionary. A wake-word model, a call-centre ASR system and an expressive voice-synthesis model need different speech: scripted prompts for phonetic coverage, spontaneous conversation for real-world ASR, expressive reads for synthesis. The reference projects encode this in their ratios — AfriVoices-KE's 750 scripted against 2,250 spontaneous hours; India's national TTS framework splitting each speaker's ten hours into nine of neutral read speech and one of expressive storytelling. Our design stage fixes the same decisions per project: prompt sets and scenarios, dialect coverage (the Kenyan project explicitly covered both Nandi and Kipsigis within Kalenjin — a dialect decision, made upfront), demographic quotas, and recording environments, from quiet-room capture to deliberately natural settings when the model must survive background noise.

Write the review criteria with the prompts. Because our dual-layer review will later judge every recording, the design stage also defines what "pass" means — audible criteria like clipping, truncation, mispronunciation of the prompt, wrong dialect or code-switching where none was asked for — so reviewers apply a written standard rather than taste. The cultural layer is designed here too: our cultural voice synthesis work across 30 languages taught us that pacing, register and idiom are language-specific review criteria, not universal ones, and the guideline for each language is drafted with the region-native team that will apply it.


What happens in the recording and review stages?

Controlled capture with session-level checks, then a two-pass human review with authority to reject — the same discipline the strongest public projects describe.

Recording is run like a production, because it is one. The professional playbook is visible in the bestdocumented programs: India's 22-language TTS effort checks equipment before every session, audits audio after recording with re-records where needed, and mandates a fifteen-minute break for every forty-five minutes of recording — fatigue is an audio-quality variable. Our sessions run the same way through the delivery-centre network: verified native speakers, session-level technical checks (levels, noise floor, sample integrity), and immediate flagging of failed takes while the speaker is still available, because a re-record on the day costs minutes and a re-record after delivery costs a recruitment cycle.

Review is dual-layer, and rejection is real. Every batch then passes the same two-pass human-in-theloop review we apply across our data work: a first pass checks each clip against the written criteria — audio quality, prompt fidelity, speaker eligibility — and an independent second pass audits the result, with authority to reject and with decisions recorded. Transcription and metadata go through the same gate: a beautiful recording with a wrong transcript is a defect, and the field's audits show transcript-audio mismatch is exactly where uncontrolled collections decay. The Bambara project's phrasing — controlled environment, consistent quality control prior to transcription — is the pattern; we simply run it as standing infrastructure rather than per-project scaffolding.


What ships at delivery — and what did running this teach us?

Audio plus its paper trail: transcripts, speaker metadata, consent records and review provenance — and three lessons the program keeps re-teaching.

The deliverable is a documented dataset, not a folder of audio. What leaves the program is the recording set with aligned transcripts, per-clip metadata (speaker demographics, dialect, environment, device), the consent record behind every voice, and the review provenance — what was checked, by whom, and what was rejected. That package is what lets a client's ML team trust the corpus without re-auditing it, and it mirrors what the strongest public datasets now publish about themselves.

Three lessons from running it. First, the specification is the quality system — almost every defect we reject at review traces back to something a sharper prompt set or guideline would have prevented. Second, speaker supply is the schedule — finding qualified voices, especially for smaller languages and for bilingual requirements (the Indian TTS team notes how hard it is to find talent fluent in both the native language and English), takes longer than recording them, which is why our standing contributor network across 30+ countries is the program's real asset. Third, spontaneous speech is where programs earn their keep: scripted audio is easy to check against its prompt, but natural conversation needs native-speaker judgment on every clip — and that judgment, at scale, in 50+ languages, is precisely what a managed program exists to supply.

A caution on the numbers. Corpus sizes and project details above are as published by the cited papers and organisations; our own program details are first-party descriptions of how we work, stated at the level we publish them. Dataset scales in this field move quickly — verify current figures against the original sources before quoting them.

The Lifewood speech collection pipeline 1 2 3 4 SPECIFY RECRUIT & VERIFY RECORD & CHECK REVIEW & DELIVER Prompts, speaker quotas, dialects, environments and written pass criteria — designed from the model's needs Native speakers sourced and vetted through delivery centres in 30+ countries, with consent on record Controlled sessions with technical checks and same-day re-records — fatigue and noise managed as variables Dual-layer human review with authority to reject; audio ships with transcripts, metadata and provenance The same dual-layer discipline we apply to annotation and AIGC, pointed at audio — because a voice dataset is a dataset first.


Key takeaways

    • The market's bar is scale with verification: reference corpora like AfriVoices-KE (~3,000 hours, 4,777 balanced native speakers) and NaijaVoices (1,800 hours) publish their speaker mixes and review processes, while audits call the reliability of loosely controlled collections "largely unknown".
    • Web-mined volume is cheap and lossy — roughly a quarter of the million-hour People's Speech corpus fell out at validation — which is why managed collection remains the route to training-grade audio.
    • Design does most of the quality work: prompt sets and scripted/spontaneous ratios chosen for the target model, dialect and demographic quotas fixed upfront, and written pass criteria drafted with region-native teams before recording begins.
    • Recording runs like a production: equipment checks per session, audio audited with same-day re-records, and speaker fatigue managed as an audio-quality variable — the discipline documented in India's 22language TTS framework.
    • Every batch passes our dual-layer review — independent second pass, authority to reject, decisions recorded — covering audio, transcripts and metadata alike.
    • Delivery is a documented dataset: aligned transcripts, per-clip speaker and environment metadata, consent records, and review provenance.
    • The lessons: the specification is the quality system, speaker supply sets the schedule, and spontaneous speech is where native-speaker judgment at scale earns its keep.
    • Corpus figures are as published and move quickly — verify at source.

Sources and further reading

Frequently asked questions

Usually both, in a ratio set by the model: scripted prompts buy phonetic coverage and easy verification; spontaneous conversation buys the disfluencies, code-switching and pacing real users produce. The reference projects run roughly one part scripted to three parts spontaneous for ASR-oriented corpora.

Enough to represent the population the model will hear — which is a diversity question before a volume question. Thousands of hours from a handful of voices trains a model on those voices; the landmark projects spread comparable hours across thousands of demographically balanced speakers.

They are excellent for pretraining and benchmarking, but audits show quality is uneven exactly in lowresource languages, and public sets rarely match a product's domain, dialects or acoustic conditions. Most programs blend public foundations with custom, verified collection for what the model actually faces.

The same way the audio is: written criteria, a first transcription pass, and an independent native-speaker review against the recording — because transcript-audio mismatch is the classic decay mode of uncontrolled collections.

Documented, informed consent covering the recording itself, the uses of the data, and the metadata kept about them — recorded per speaker and delivered as part of the dataset's provenance. Voice is personal data everywhere and biometric-adjacent in several jurisdictions; the consent record is part of the product.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team