Skip to main content
AI Data

How a Speech Data Collection Programme Actually Runs

July 2026 · 8 min read · Updated September 2026

Short answer. As a managed pipeline, not a crowdsourced funnel: we design prompts and scripts for the target model, recruit and verify native speakers through our delivery centres, record under controlled conditions with session-level checks, and pass every batch through our dual-layer human review before delivery. The field's own audits explain why: loosely controlled community collections show serious quality problems in exactly the low-resource languages where the data matters most.

Key takeaways

  • Reference corpora like AfriVoices-KE (roughly 3,000 hours, 4,777 balanced native speakers) and NaijaVoices (1,800 hours) publish their speaker mixes and review processes, while audits call the reliability of loosely controlled community collections "largely unknown."
  • Web-mined volume is cheap and lossy: roughly a quarter of the million-hour People's Speech corpus fell out at validation, which is why managed collection remains the route to training-grade audio.
  • Design does most of the quality work — prompt sets, scripted-to-spontaneous ratios, dialect and demographic quotas, and written pass criteria are all fixed before recording begins.
  • Every batch passes a dual-layer human review with independent second-pass authority to reject, and delivery includes aligned transcripts, metadata, consent records and review provenance.
  • The specification is the quality system, speaker supply sets the schedule, and spontaneous speech is where native-speaker judgment at scale earns its keep.

What does the market now expect from a speech corpus?

Scale with verification. The reference projects of the past two years pair thousands of hours with named speakers, demographic balance and documented review.

The bar moved fast. AfriVoices-KE, one of 2026's landmark collections, targeted roughly 3,000 hours across five Kenyan languages — 750 hours of scripted and 2,250 hours of spontaneous speech — recorded from 4,777 native speakers deliberately spread across regions and demographics. NaijaVoices delivered 1,800 hours of authentic Igbo, Hausa and Yorùbá speech; Mali's Bambara corpus went from a single 30-hour collection to 612 hours — a 20x jump — recorded in a controlled environment with consistent quality control before transcription. At the other end of the spectrum sits found data: MLCommons' Unsupervised People's Speech assembled over a million hours across 89+ languages by mining the web, of which about 736,000 hours survived validation.

Scripted speech is read aloud from a fixed prompt, built for phonetic coverage and easy verification; spontaneous speech is unscripted conversation, built to capture the disfluencies and pacing real users produce. The gap between managed and mined data is the reason programmes like ours exist: a 2025 audit of widely used multilingual speech datasets found serious quality issues concentrated precisely in low-resource, less-institutionalised languages, and noted that even the flagship community-driven platform lacks well-documented quality control for its source texts and recordings — making reliability "largely unknown." Anyone scoping a low-resource language collection is buying against that gap.

What a reference-grade collection looks like now: AfriVoices-KE contributed 2,250 hours of spontaneous speech and 750 hours of scripted speech from 4,777 named, demographically balanced native speakers; NaijaVoices added 1,800 hours across Igbo, Hausa and Yorùbá; Bambara scaled 30 hours to 612 hours between 2022 and 2025. By contrast, roughly 26% of the web-mined, million-hour People's Speech corpus was lost at validation, and audits describe the reliability of undocumented community collections as simply "unknown" — worst in low-resource languages. Figures are from the AfriVoices-KE, Bambara and MLCommons publications and the 2025 multilingual speech-quality audit, cited below.

Volume is now cheap; verified hours, balanced speakers and documented review are what a training team is actually buying. That is the specification our programme is built against.

How is a collection designed before anyone records?

The specification does most of the quality work: what the model needs decides the prompts, the speaker mix and the environments — before a microphone is switched on.

Start from the model's diet, not the language's dictionary. A wake-word model, a call-centre ASR system and an expressive voice-synthesis model need different speech: scripted prompts for phonetic coverage, spontaneous conversation for real-world ASR, expressive reads for synthesis. The reference projects encode this in their ratios — AfriVoices-KE's 750 scripted against 2,250 spontaneous hours; India's national TTS framework splitting each speaker's ten hours into nine of neutral read speech and one of expressive storytelling. Our design stage fixes the same decisions per project: prompt sets and scenarios, dialect coverage (a recent Kenyan project explicitly covered both Nandi and Kipsigis within Kalenjin — a dialect decision made upfront), demographic quotas, and recording environments, from quiet-room capture to deliberately natural settings when the model must survive background noise. This is the same design discipline behind our multilingual data collection programmes more broadly.

Write the review criteria with the prompts. Because our dual-layer review will later judge every recording, the design stage also defines what "pass" means — audible criteria like clipping, truncation, mispronunciation of the prompt, wrong dialect, or code-switching where none was asked for — so reviewers apply a written standard rather than taste. The cultural layer is designed here too: our cultural voice synthesis work across 30 languages taught us that pacing, register and idiom are language-specific review criteria, not universal ones, and the guideline for each language is drafted with the region-native team that will apply it, the same approach we use when we recruit native contributors for language data more generally.

What happens in the recording and review stages?

Controlled capture with session-level checks, then a two-pass human review with authority to reject — the same discipline the strongest public projects describe.

Recording is run like a production, because it is one. The professional playbook is visible in the best-documented programmes: India's 22-language TTS effort checks equipment before every session, audits audio after recording with re-records where needed, and mandates a fifteen-minute break for every forty-five minutes of recording — fatigue is treated as an audio-quality variable. Our sessions run the same way through the delivery-centre network: verified native speakers, session-level technical checks (levels, noise floor, sample integrity), and immediate flagging of failed takes while the speaker is still available, because a re-record on the day costs minutes and a re-record after delivery costs a recruitment cycle. Consent is captured at this stage too, following the same principles covered in how data contributors should be consented and paid.

Dual-layer human review means every unit of data passes through two independent checks — a first pass against written criteria, then a second, separate pass with authority to reject — rather than a single reviewer's judgement. Every batch passes the same two-pass review we apply across our data work: a first pass checks each clip against the written criteria — audio quality, prompt fidelity, speaker eligibility — and an independent second pass audits the result, with authority to reject and with decisions recorded. Transcription and metadata go through the same gate: a clean recording with a wrong transcript is a defect, and the field's audits show transcript-audio mismatch is exactly where uncontrolled collections decay — a failure mode we also cover for diarization and timestamping work. The Bambara project's own phrasing — controlled environment, consistent quality control prior to transcription — describes the same pattern; we run it as standing infrastructure rather than per-project scaffolding.

What ships at delivery, and what did running this teach us?

Audio plus its paper trail: transcripts, speaker metadata, consent records and review provenance — and three lessons the programme keeps re-teaching.

The deliverable is a documented dataset, not a folder of audio. What leaves the programme is the recording set with aligned transcripts, per-clip metadata (speaker demographics, dialect, environment, device), the consent record behind every voice — the documented, informed agreement covering the recording, its uses, and the metadata kept about the speaker — and the review provenance: what was checked, by whom, and what was rejected. That package is what lets a client's ML team trust the corpus without re-auditing it, and it mirrors what the strongest public datasets now publish about themselves. The same discipline applies whether the target is a wake-word model or studio-cast TTS voice data.

Three lessons from running it. First, the specification is the quality system: almost every defect rejected at review traces back to something a sharper prompt set or guideline would have prevented. Second, speaker supply is the schedule: finding qualified voices, especially for smaller languages and bilingual requirements, takes longer than recording them, which is why a standing contributor network across many countries is the programme's real asset — the same network behind our global multilingual speech data collection services. Third, spontaneous speech is where programmes earn their keep: scripted audio is easy to check against its prompt, but natural conversation needs native-speaker judgment on every clip, at scale, across many languages — precisely what a managed programme exists to supply.

The pipeline runs in four stages: specify (prompts, speaker quotas, dialects, environments and written pass criteria, designed from the model's needs); recruit and verify (native speakers sourced and vetted through delivery centres, with consent on record); record and check (controlled sessions with technical checks and same-day re-records, with fatigue and noise managed as variables); and review and deliver (dual-layer human review with authority to reject, with audio shipped alongside transcripts, metadata and provenance). It is the same dual-layer discipline applied to annotation and AIGC work, pointed at audio, because a voice dataset is a dataset first.

A caution on the numbers: corpus sizes and project details above are as published by the cited papers and organisations; our own programme details are first-party descriptions of how we work, stated at the level we publish them. Dataset scales in this field move quickly — verify current figures against the original sources before quoting them.

Frequently asked questions

Usually both, in a ratio set by the model: scripted prompts buy phonetic coverage and easy verification; spontaneous conversation buys the disfluencies, code-switching and pacing real users produce. The reference projects run roughly one part scripted to three parts spontaneous for ASR-oriented corpora.

Enough to represent the population the model will hear — a diversity question before a volume question. Thousands of hours from a handful of voices trains a model on those voices; landmark projects spread comparable hours across thousands of demographically balanced speakers instead.

They are excellent for pretraining and benchmarking, but audits show quality is uneven exactly in low-resource languages, and public sets rarely match a product's domain, dialects or acoustic conditions. Most programmes blend public foundations with custom, verified collection for what the model actually faces.

The same way the audio is: written criteria, a first transcription pass, and an independent native-speaker review against the recording — because transcript-audio mismatch is the classic decay mode of uncontrolled collections.

Documented, informed consent covering the recording itself, the uses of the data, and the metadata kept about them — recorded per speaker and delivered as part of the dataset's provenance. Voice is personal data everywhere and biometric-adjacent in several jurisdictions, so the consent record is part of the product.

Programmes built for this work combine delivery-centre recruitment of native speakers, model-specific prompt design, and dual-layer human review — the discipline reference projects like AfriVoices-KE and NaijaVoices document publicly. Lifewood runs speech and audio collection this way, including for low-resource languages, with consent and review provenance included in every delivery.

Sources and further reading

  1. "AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages" (arXiv), on the ~3,000-hour, 4,777-speaker collection and its scripted/spontaneous split
  2. "Data Quality Issues in Multilingual Speech Datasets" (arXiv), the audit finding serious quality problems in low-resource languages and undocumented QC in community collections
  3. "A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages" (arXiv), on session checks, re-recording, speaker breaks and read/expressive splits
  4. "Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara" (arXiv), on the 30-to-612-hour scale-up and controlled collection with QC before transcription
  5. Factored / MLCommons, "Unsupervised People's Speech," on the million-hour web-mined corpus and its 736,000 validated hours
  6. Lifewood, on speech and audio data collection and cultural voice synthesis across 30 languages

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team