Short answer. Session one and session forty have to be acoustically and performatively indistinguishable, because the model learns recording variance alongside the voice. That is why crowdsourced and home-studio capture fails for TTS: room acoustics, microphone position, vocal warmup and performance register all become part of what the model reproduces. The old standard was roughly 24 hours of single-speaker read audio, producing a voice usually described as acceptable but flat. Modern systems train one model across many speakers, each reduced to a speaker embedding — which separates language from identity, so a voice can speak languages its actor never learned.
Key takeaways
- Session one and session forty need to be acoustically and performatively indistinguishable, because the model learns recording variance alongside the intended voice.
- The old standard was roughly 24 hours of single-speaker read audio, producing a voice described as acceptable but flat; modern systems train one model on many speakers using speaker embeddings instead.
- Hours needed depend on method: zero-shot cloning from seconds of audio, fine-tuning on a multi-speaker base model at one to five hours, flagship preset voices at twenty to forty studio hours, and foundation models at 100 to 300 hours or more.
- ASR data is not TTS data: background noise is training signal for recognition and contamination for synthesis, and TTS requires precisely aligned text-speech pairs.
- WAV lossless at a target of 48 kHz is the format standard for new collection, though widely used corpora such as LibriTTS-R (roughly 585 hours) run at 24 kHz.
- Talent should be contracted for the full recording programme and consent should explicitly cover AI training use, separate from any voice-over engagement.
What casting question does TTS voice data actually answer?
The discipline exists to keep a speaker's recordings acoustically and performatively identical across a multi-week programme, because any inconsistency the performer or the room introduces becomes part of the trained voice.
The single best sentence found in researching this topic is a production requirement, not a technical spec: "Session one and session forty need to be acoustically and performatively indistinguishable." That is the whole discipline. Text-to-speech development requires hours of speech recorded under identical conditions, and the reason is stated equally plainly by the same source: crowdsourced recordings and self-directed home studio sessions introduce variance in room acoustics, microphone position, vocal warm-up and performance register, and the model learns that variance alongside the intended voice characteristics. Your inconsistency becomes part of the voice — that is what separates TTS collection from almost every other kind of speech data work.
How has the "how many hours" question changed?
It has moved from a single number to a question about architecture, because the amount of audio needed now depends on which training method is used, not on a fixed studio target.
The old question was simple: how many hours of clean single-speaker audio do you have? The standard answer was around 24 hours of LJSpeech-style read audio, which produced a model that sounded, in one practitioner's assessment, acceptable but flat. That era has ended, and the reason is architectural. Modern systems train one model on many speakers at once, with each speaker summarised as a speaker embedding — a vector a few hundred numbers long that captures what makes a voice itself, separated from the language content it is speaking. The framing that makes this click: the model is the instrument, and the embedding tells it which person to be, or, as one source puts it, the difference between a tape archive and a gifted impressionist. The archive can only replay what was recorded; the impressionist has one vocal tract and an unlimited number of identities, because identity turned out to be the small part of the problem.
Two consequences follow that matter commercially. A voice can speak languages its actor never learned, because language and identity live in different parts of the model. And voices can be blended, since somewhere between any two embeddings sits a third voice that has never existed.
So how much audio do you actually need? It depends entirely on the method, and the published ranges are wide. Zero-shot cloning — generating a new voice from a very short reference clip with no dedicated fine-tuning — works from seconds of audio. Fine-tuning on a strong multi-speaker base model typically takes one to five hours. Flagship preset voices still use twenty to forty studio hours, with a reason worth quoting: every gap in the data becomes a gap in the voice. Broad-coverage foundation models are quoted by data vendors at 100 to 300 hours or more. The practical read is that the hours question is now downstream of the architecture question — ask what is being trained before asking how much to record. Teams scoping a programme for a low-resource language should expect that question to matter even more, since a smaller pool of speakers narrows which methods are realistic.
Why can't an existing ASR dataset be reused for TTS?
Because the features that make a corpus useful for speech recognition are frequently the features that make it unusable for synthesis, most existing speech archives cannot be repurposed as TTS training data.
Background noise is the clearest example: for ASR it is training signal, teaching the model to cope with real conditions. For TTS it is contamination that the model will reproduce. TTS also requires pairs of text and corresponding speech recordings, aligned, with diverse and representative samples across speakers and speaking styles. An ASR corpus with approximate transcripts is fine for recognition and unusable for synthesis. The practical implication for anyone with an existing speech archive is that it is probably not a TTS dataset, and treating it as one produces a voice that has learned the recording's problems rather than the speaker's.
What technical specification should a TTS dataset meet?
The published standards are reasonably consistent: lossless WAV audio, a 48 kHz sample rate for new recordings, and a full metadata schema delivered alongside the audio.
Format is WAV, lossless. The reasoning is that compressed formats such as MP3 strip subtleties of intonation, pause and emotional tone, and the reported result is robotic or less expressive output. Sample rate is 48 kHz, described as the accepted standard for TTS datasets on the argument that it captures the full frequency spectrum of human speech and leaves post-processing headroom. Worth a caveat here, because published corpora vary: LibriTTS-R, a widely used multi-speaker English corpus of roughly 585 hours, is at 24 kHz. So 48 kHz is a target for new commissioned collection rather than a universal property of TTS data, and combining commissioned audio with existing corpora usually means resampling something.
Metadata is delivered alongside the audio, typically as a JSON sidecar. The published schema is worth copying in full: speaker profile (age, gender, accent, vocal range), session metadata (recording session dates), signal chain (microphone and preamp chain used), room acoustics profile, time-aligned transcriptions at both utterance and word level, and phonetic annotations. The signal chain and room profile are the entries most often omitted, and the ones that make a dataset reproducible — if session forty needs to match session one, there has to be a record of what session one actually was.
Why does phonetic script coverage matter more than script length?
Because a phoneme or phonetic context that never appears in the recording script will be synthesised by extrapolation later, and extrapolation is where audible artefacts live, not because a script needs to be longer.
Where phonetic coverage — a script's inclusion of every sound and sound-sequence a language uses, in enough contexts to train each reliably — requirements are defined, scripts should be reviewed against coverage targets before recording begins. This is also where language-specific expertise becomes non-negotiable: designing a phonetically balanced script for Bengali, Yoruba or Vietnamese requires knowing that language's phoneme inventory and its permissible sequences. A script translated from an English phonetically balanced set is balanced for English, not for the target language. One practical note worth adopting: pilot batches of five to ten finished hours delivered within five to seven working days of brief sign-off is a reasonable expectation, and running that pilot before committing to a full programme is the cheapest quality control available.
How is casting for TTS different from hiring an annotator?
Casting for TTS is closer to casting for a long-running production than to hiring an annotator, because the programme runs across weeks and the same voice has to be sustained for all of it.
Published practice is to provide audition samples from a roster so the client selects the voice, with casting available against specific demographics, accents or vocal qualities against the brand brief. Talent is then contracted for the full programme duration — a voice actor who becomes unavailable at hour twenty-five leaves a project with an unfinished voice, not a partial dataset, because a different speaker cannot be substituted into the same identity. Casting also has to account for endurance, not just tone: forty studio hours is a substantial vocal load, and a voice that is beautiful for two hours and tires by hour six produces a dataset with an audible arc in it. Emotional range, if the target is expressive synthesis, has to be decided at casting rather than inferred afterward — one commercial dataset is described as 1,000 hours across 8 languages and 43 emotional states, sourced from more than 150 professional voice artists, delivered with emotion and tone tags. Whatever the right number of states for a given application, it is a casting and scripting decision made before session one, alongside the conversational and consent design that governs how contributors are treated more broadly.
How is drift managed across a multi-week recording programme?
Drift is managed through a fixed setup and active monitoring rather than through the performer's memory of earlier sessions, since small, cumulative changes in delivery are otherwise invisible in the room and obvious to the model.
Drift — the gradual, often unnoticed change in a performer's energy, pacing or articulation across recording days — is managed with reference playback at the start of every session, so the performer re-enters the same register rather than approximating it from memory. Sessions are directed by audio producers who monitor for drift in energy, pacing and articulation across recording days, described as the subtle inconsistencies that cause artefacts in synthesised output. Physical setup — microphone position, distance, room treatment and signal chain — is fixed and documented rather than reassembled each session, and structured QA before delivery covers audio quality checks, transcript and labelling review, and validation against the technical specification. Programmes that also need diarization or timestamped transcripts alongside the voice data can draw on the same speech and audio annotation discipline used for QA on the labelling side.
Why does consent for voice data differ from consent for other data?
Consent for voice has to cover identity, not just content, because a recording that trains a model to reproduce someone's voice is a different act than recording them for a single use such as an advert.
Scraped audio is described in commercial literature as risky, inconsistent and legally uncertain, and the alternative is explicit: sourcing real voice actors with explicit consent for AI training use. One useful distinction is separating voice-over work from voice data — recordings made for a voice-over job are used only for that purpose, with voice data sourced separately through dedicated consented collection. That separation matters because a performer who recorded an advert did not consent to having their voice modelled, and conflating the two is the fastest route to a dispute. Consent should explicitly cover recording, use and licensing for AI training purposes, as distinct from consent to be recorded at all.
What does this mean for scoping a multilingual TTS programme?
Studio consistency and phonetic script design are usually the two most under-scoped parts of a TTS programme, and both need to be resolved before a studio is booked, not after.
Lifewood has run voice AI data collection since establishing dedicated voice hubs, and works across 50+ languages. Two observations apply regardless of supplier. First, studio consistency is a logistics problem before it is an audio problem: booking the same room, chain and performer across six weeks is straightforward where professional studio capacity exists, but for a Sylheti or Wolof voice the room and the performer may need to be assembled rather than hired, and that assembly is the part of a project that runs late — the same constraint that shapes broader low-resource language speech collection. Second, phonetic script design is where multilingual TTS projects most often under-specify: a client brief typically states hours, voice character and delivery format, and rarely states phonetic coverage. For a language with an existing balanced corpus that omission is survivable; for a language where no such corpus exists, someone has to build the coverage target from the phoneme inventory up, or the gaps only become visible when the synthesised voice mispronounces a common construction. Both point toward scoping the linguistics before booking the studio — the reverse of how these programmes are usually planned, and a pattern that shows up across managed speech data programmes more broadly, including work spanning accented and non-native speech and coordinated multilingual speech collection across markets. Buyers weighing TTS against a broader annotation or AI data services scope should treat voice as its own specification, not a line item inside a general speech brief.
What should a TTS data specification checklist include?
A complete specification fixes the architecture, the format, the phonetic coverage target, the metadata schema and the consent terms before recording starts, rather than leaving any of them to be decided mid-programme.
- Architecture first: zero-shot, fine-tune or foundation model determines the hours requirement before anything else.
- WAV format with a stated sample rate — 48 kHz as the target for new collection, with a resampling plan if combining with existing corpora.
- Phonetic coverage targets built from the target language's own inventory, with scripts reviewed against them before recording.
- A full metadata schema including signal chain and room profile, not just speaker demographics.
- Utterance and word-level time alignment plus phonetic annotation.
- Talent contracted for the full programme, with reference playback and directed sessions monitoring drift.
- Emotional range specified upfront if expressive synthesis is the target.
- Consent covering AI training use explicitly, separate from any voice-over engagement, plus a five-to-ten-hour pilot before the full programme.