Short answer. To collect multilingual speech data for conversational AI, define the intended conversations first, recruit consented native or proficient speakers across the needed language varieties, record realistic speech in controlled and natural conditions, and use human-in-the-loop transcription and review. Document contributors and limits, then test results separately by language, accent, noise condition, and task.
Key takeaways
- Multilingual conversational AI needs data that reflects real interactions such as commands, questions, turn-taking, repairs, and code-switching, not only read sentences.
- Languages should be specified at the locale or variety level, such as es-MX and es-ES, rather than treating Spanish as one uniform category.
- Voice recordings are sensitive personal data, so collection needs informed consent, minimal metadata, restricted access, and defined retention and deletion rules.
- Quality and model performance should be measured by language, accent, and condition, because one overall accuracy score can hide poor results for an important market.
What speech data does a conversational AI system need?
A conversational AI system needs audio paired with enough context to be useful: a transcript, a language or locale label, recording conditions, and appropriately collected consent and speaker metadata. Depending on the product, it may also need intent labels, dialogue turns, entity tags, interruption markers, and outcome labels.
Speech data is audio recorded from speakers and paired with transcripts and metadata so that a model can learn to recognise, interpret, or generate spoken language.
The required mix depends on the product:
| Product capability | Data to prioritise |
|---|---|
| Speech-to-text | Natural speech, accurate transcripts, accents, background noise, device variation |
| Voice assistant | Commands, questions, follow-up requests, interruptions, wake-word conditions |
| Customer-service bot | Domain vocabulary, numbers, account references, repair phrases, emotional and fast speech |
| Speech translation | Aligned source and target speech and text, plus locale-specific terminology |
| Text-to-speech | Consent-cleared voice recordings, pronunciation coverage, style and prosody labels |
Public datasets can be a useful baseline, but they rarely match an enterprise's domain, acoustic environment, or consent requirements. Mozilla's Common Voice 18 release, for example, contained 20,789 community-validated hours across 129 languages. That is valuable evidence of multilingual scale, but not a substitute for data designed around a particular product and user journey (Mozilla Foundation). For the managed-service side, see global multilingual speech data collection services.
How should you plan the collection before recruiting speakers?
Start with a data specification rather than a recruitment target such as 10,000 hours. The specification should settle which languages and varieties matter, what situations users will be in, which language behaviours the model must understand, what the quality bar is, and what the data will and will not be used for.
A data specification is the written contract for a collection that fixes languages, conditions, behaviours, acceptance rules, and permitted uses before any speaker is recruited.
Which languages and varieties matter?
List the language, country or region, script where relevant, expected accents, and common mixing patterns. A Malaysian English customer-support assistant, for example, may need English mixed with Malay, Mandarin, Tamil, product names, and local place names. Treating all of this as generic English creates predictable blind spots. The guide to collecting code-switched speech and text covers this in more depth.
What situations will users be in?
Define microphones, devices, noise levels, network conditions, speaking distance, and whether people will be walking, driving, or using a speakerphone. If production audio is noisy but training data is studio-clean, performance estimates will be too optimistic.
What language behaviours must the model understand?
Include the patterns your product expects: short commands, complete questions, follow-ups, self-corrections, interruptions, numbers, dates, spelling, names, and code-switching. Read speech is easier to collect, but it does not fully represent live conversation.
What is the quality bar?
Set acceptance rules before fieldwork. Examples include minimum audio quality, allowable clipping, transcript conventions, required metadata, and when a clip must be rejected or escalated. A clear protocol prevents different teams from applying different standards later.
What will the data be used for, and excluded from?
State whether recordings support model training, testing, voice cloning, evaluation, or research. Dataset documentation should also record limitations and prohibited uses. The widely used Datasheets for Datasets framework recommends documenting a dataset's creation, composition, intended uses, maintenance, and legal or ethical considerations (Gebru et al.).
How do you recruit speakers for coverage rather than volume?
Recruit against a coverage matrix, not a headcount. The goal is not demographic collection for its own sake but a dataset that represents the people and conditions the system is intended to serve.
A coverage matrix is a table of the speaker, language, device, and environment variables that affect the use case, used to check that each combination is actually represented.
Include the variables that meaningfully affect the use case:
- Language, locale, dialect, and accent
- First language and multilingual or code-switching patterns
- Age bands and gender representation, where lawful and relevant
- Device type and recording environment
- Speaking style: careful, spontaneous, fast, soft, emotional, or interrupted
- Domain familiarity, especially for specialist vocabulary
A large speaker count does not automatically solve representation, because a dataset can be large while concentrating on one city, device type, or speech style. Research on multilingual acoustic modelling notes that useful diversity includes unique speakers, recording hardware, and recording conditions, not only hours of audio (CommonPhone). For accent coverage specifically, see collecting accented and non-native speech.
For lower-resource languages, work with local communities, language organisations, and native-language reviewers early. They can help create culturally appropriate prompts, identify borrowed words and regional terms, and prevent a collection process from mislabelling a language variety. Lifewood's low-resource speech data service is built around this kind of local sourcing.
How should you record speech that matches real use?
Use a mixture of collection modes, because each answers a different modelling need. Scripted prompts cover key phrases and rare words, scenario-based speaking covers task language, natural conversation covers dialogue behaviour, and in-product opt-in samples capture real devices and acoustics.
| Collection mode | Best for | Main limitation |
|---|---|---|
| Scripted prompts | Coverage of key phrases, rare words, numbers, and pronunciation targets | Less natural rhythm and dialogue behaviour |
| Scenario-based speaking | Customer-service requests, task instructions, and domain language | Requires careful prompt design and consent handling |
| Natural conversations | Turn-taking, interruptions, repairs, and code-switching | More difficult to de-identify and annotate |
| In-product opt-in samples | Real device and acoustic conditions | Must not become a shortcut around clear consent and governance |
Build prompts locally rather than translating an English script word for word. Native reviewers should check whether the phrase is natural, culturally appropriate, and representative of how people would actually request the task. They should also verify pronunciation-sensitive material such as names, addresses, prices, dates, and alphanumeric strings.
Record technical metadata alongside each clip: sampling rate, microphone or device class, recording environment, locale, prompt or scenario identifier, and collection version. This makes later analysis possible without exposing more personal information than necessary. A step-by-step operational view is in running a speech data collection programme.
How does human-in-the-loop transcription and annotation work?
Human-in-the-loop transcription combines automated checks with native-language people who transcribe, review, and adjudicate speech. Automation can accelerate collection, but people are still essential for judgement calls such as dialect questions, overlapping speech, and sensitive content.
Human-in-the-loop (HITL) is a workflow in which trained people review, correct, or decide on machine-processed data at defined points, so that quality does not depend on automation alone.
A practical HITL workflow
- Automated intake checks flag silence, clipping, duplicate files, wrong duration, and corrupted audio.
- Native-language transcription produces a transcript using documented spelling and code-switching rules.
- Independent review checks a sample or all high-risk clips, depending on the quality target.
- Adjudication resolves disagreements, unclear speech, dialect questions, and sensitive content.
- Feedback loops update prompts and guidance when the same errors appear repeatedly.
Annotation guidelines should define what to do with fillers, false starts, laughter, overlaps, background speech, profanity, personally identifying information, and words that cannot be reliably heard. Without these rules, annotators may be individually accurate but inconsistent as a group.
Useful quality measures include transcription agreement, word or character error rates on a held-out audited sample, rejection rate by collection channel, and disagreement rate by language. Track them by locale, not only as a global average. Lifewood applies this pattern across 100+ languages with two independent review passes through its multilingual data collection service, and the wider discussion is in human-in-the-loop multilingual data quality.
How do you build privacy, consent, and governance into speech collection?
Treat privacy as part of project design, not a legal review at the end. Voice recordings can contain personal information and, in some contexts, biometric or otherwise sensitive signals, so consent, minimisation, access control, and deletion rules must be set before recording starts.
At a minimum:
- Explain clearly what is recorded, why, who can access it, how long it will be retained, and whether it may be used for training or synthetic-voice development.
- Obtain consent appropriate to the purpose and jurisdiction, and make withdrawal and deletion procedures workable.
- Collect the minimum metadata needed for quality and fairness analysis.
- Separate identity and contact information from audio and operational labels where possible.
- Apply access controls, encryption, audit trails, and supplier obligations.
- Define what happens when recordings include accidental personal data or sensitive disclosures.
The NIST Privacy Framework is a voluntary enterprise-risk-management resource for identifying and managing privacy risk. It is a useful starting point, but teams should also obtain jurisdiction-specific legal advice for their collection locations and intended uses. Contributor terms and payment are covered in consent and pay for data contributors.
How do you validate the dataset and the deployed experience?
Validate with evaluation sets that are separate from training data and that deliberately cover the most important combinations of language, accent, task, device, and noise condition. Do not wait until model launch to discover gaps.
Report results in slices, such as:
- Speech-recognition error by locale and acoustic environment
- Intent accuracy by language and code-switching condition
- Failure and fallback rates for names, numbers, dates, and domain terms
- Human-review overrides and unresolved annotation cases
- Complaint, correction, or abandonment rates after launch
This approach distinguishes a truly broad model improvement from a result driven by the most represented language. It also gives product, data, and localisation teams a shared way to prioritise the next collection round. If you are comparing outside providers, the ranked list of top multilingual data collection companies is a useful starting shortlist.
What should a speech collection checklist include?
A speech collection checklist should confirm that scope, consent, prompts, quality standards, reviewers, and evaluation data are all settled before the first recording. Use it as a gate before fieldwork begins.
Before collection begins, confirm that you have:
- Defined the product tasks and supported locales
- Built a coverage matrix for speakers, conditions, and language behaviours
- Written consent, retention, withdrawal, and data-access procedures
- Localised and reviewed prompts with native speakers
- Set audio, metadata, transcription, and rejection standards
- Trained annotators and reviewers using the same guideline version
- Created a disagreement and escalation process
- Reserved a representative, consent-cleared evaluation set
- Prepared a dataset datasheet and version-control process
- Chosen subgroup metrics to monitor after deployment
Strong multilingual conversational AI starts with a collection plan that respects real users. Scale is valuable, but representative, well-documented data is what makes scale dependable.