Skip to main content
AI Data

Studio Standards and Speaker Casting for TTS Voice Data

Short answer. Session one and session forty have to be acoustically and performatively indistinguishable, because the model learns recording variance alongside the voice. That is why…

Mumu D. · September 2026 · 12 min read

Download PDF

Short answer. Session one and session forty have to be acoustically and performatively indistinguishable, because the model learns recording variance alongside the voice. That is why crowdsourced and home-studio capture fails for TTS: room acoustics, microphone position, vocal warmup and performance register all become part of what the model reproduces. The old standard was roughly 24 hours of single-speaker read audio, producing a voice usually described as acceptable but flat. Modern systems train one model across many speakers, each reduced to a speaker embedding — which separates language from identity, so a voice can speak languages its actor never learned.


Casting Does TTS Voice Data Need?

The single best sentence I found in researching this topic is a production requirement, not a technical spec:

"Session one and session forty need to be acoustically and performatively indistinguishable."

That is the whole discipline. Text-to-speech development requires hours of speech recorded under identical conditions, and the reason is stated equally plainly by the same source: crowdsourced recordings and self-directed home studio sessions introduce variance in room acoustics, microphone position, vocal warm-up and performance register, and the model learns that variance alongside the intended voice characteristics.

Your inconsistency becomes part of the voice. That is what separates TTS collection from almost every other kind of speech data work.


What changed, and what it means for how much you need

The old question was simple. How many hours of clean single-speaker audio do you have? The standard answer was around 24 hours of LJSpeech-style read audio, which produced a model that sounded, in one practitioner's assessment, acceptable but flat.

That era has ended, and the reason is architectural. Modern systems train one model on many speakers at once, with each speaker summarised as a speaker embedding: a vector a few hundred numbers long that captures what makes that voice itself.

The framing that makes this click: the model is the instrument, and the embedding tells it which person to be. Or, as the same source puts it, the difference between a tape archive and a gifted impressionist. The archive can only replay what was recorded. The impressionist has one vocal tract and an unlimited number of identities, because identity turned out to be the small part of the problem.

Two consequences follow that matter commercially. A voice can speak languages its actor never learned, because language and identity live in different parts of the model. And voices can be blended, since somewhere between any two embeddings sits a third voice that has never existed.

So how much audio do you actually need? It depends entirely on the method, and the published ranges are wide:

Zero-shot cloning works from seconds of audio. Fine-tuning on a strong multi-speaker base model typically takes one to five hours. Flagship preset voices still use twenty to forty studio hours, with a reason worth quoting: every gap in the data becomes a gap in the voice. Broad-coverage foundation models are quoted by data vendors at 100 to 300 hours or more.

The practical read is that the hours question is now downstream of the architecture question. Ask what is being trained before asking how much to record.


Why ASR data is not TTS data

This gets assumed away in scoping conversations and it should not be.

The features that make a corpus valuable for speech recognition are frequently undesirable in TTS. Background noise is the clearest example: for ASR it is training signal, teaching the model to cope with real conditions. For TTS it is contamination that the model will reproduce.

TTS also requires pairs of text and corresponding speech recordings, aligned, with diverse and representative samples across speakers and speaking styles. An ASR corpus with approximate transcripts is fine for recognition and unusable for synthesis.

The practical implication for anyone with an existing speech archive: it is probably not a TTS dataset, and treating it as one produces a voice that has learned your recording problems.


The technical specification

The published standards are reasonably consistent.

Format: WAV, lossless. The reasoning is that compressed formats such as MP3 strip subtleties of intonation, pause and emotional tone, and the reported result is robotic or less expressive output.

Sample rate: 48 kHz is described as the accepted standard for TTS datasets, on the argument that it captures the full frequency spectrum of human speech and leaves post-processing headroom.

Worth a caveat here, because published corpora vary. LibriTTS-R, a widely used multi-speaker English corpus of roughly 585 hours, is at 24 kHz. So 48 kHz is a target for new commissioned collection rather than a universal property of TTS data, and if you are combining commissioned audio with existing corpora you will be resampling something.

Metadata delivered alongside the audio, typically as a JSON sidecar. The published schema is worth copying in full:

Speaker profile: age, gender, accent, vocal range. Session metadata: recording session dates. Signal chain: microphone and preamp chain used. Room acoustics profile. Time-aligned transcriptions at both utterance and word level. Phonetic annotations.

The signal chain and room profile are the entries most often omitted and the ones that make a dataset reproducible. If session forty needs to match session one, you need a record of what session one actually was.


Phonetic coverage, and why scripts are not just scripts

A detail that separates professional TTS collection from recording someone reading for a few days.

Where phonetic coverage requirements are defined, scripts should be reviewed against coverage targets before recording begins. The reason is the same as the flagship voice argument: every gap in the data becomes a gap in the voice. A phoneme or phonetic context that never appears in the script will be synthesised by extrapolation, and extrapolation is where artefacts live.

This is also where language-specific expertise becomes non-negotiable. Designing a phonetically balanced script for Bengali, Yoruba or Vietnamese requires knowing that language's phoneme inventory and its permissible sequences. A script translated from an English phonetically balanced set is balanced for English.

One practical note from the same source: pilot batches of five to ten finished hours delivered within five to seven working days of brief sign-off is a reasonable expectation, and running that pilot before committing to a full programme is the cheapest quality control available.


Casting, which is a real discipline

Speaker casting for TTS is closer to casting for a long-running production than to hiring an annotator, and the reason is duration. A twenty to forty hour programme runs across weeks.

Audition against the brand brief. Published practice is to provide audition samples from a roster so the client selects the voice, with casting available against specific demographics, accents or vocal qualities.

Contract for the whole programme. One vendor states it plainly: talent is contracted for the full programme duration. A voice actor who becomes unavailable at hour twenty-five leaves you with an unfinished voice, not a partial dataset, because you cannot substitute a different speaker into the same identity.

Cast for endurance as well as tone. Forty studio hours is a substantial vocal load. A voice that is beautiful for two hours and tires by hour six will produce a dataset with an audible arc in it.

Consider the emotional range required upfront. Expressive TTS needs it recorded, not inferred. One commercial dataset is described as 1,000 hours across 8 languages and 43 emotional states, sourced from more than 150 professional voice artists, delivered with emotion and tone tags. Whatever the right number of states for your application, it is a casting and scripting decision made before session one.


Consistency management, which is the actual craft

Given that the model learns your variance, here is what published practice does about it.

Reference playback at the start of every session. One provider describes starting each session with reference playback from the previous session, so the performer re-enters the same register rather than approximating it from memory.

Directed sessions with drift monitoring. Sessions directed by audio producers who monitor for drift in energy, pacing and articulation across recording days, described as the subtle inconsistencies that cause artefacts in synthesised output.

That word, drift, is the right one. It is not that a performer sounds different on day nine. It is that they sound slightly different, cumulatively, in ways nobody notices in the room and the model picks up immediately.

Fixed physical setup. Microphone position, distance, room treatment and signal chain locked and documented, not reassembled each session.

Structured QA before delivery. Published practice covers audio quality checks, transcript and labelling review, and validation against the technical specification.


Consent, which is different for voice than for other data

Voice identity is personal in a way transcription is not, and the practice reflects it.

Scraped audio is described in the commercial literature as risky, inconsistent and legally uncertain, and the alternative is explicit: sourcing real voice actors with explicit consent for AI training use.

The distinction worth adopting is one provider's separation of voice-over work from voice data. Recordings made for a voice-over job are used only for that purpose, with voice data sourced separately through dedicated consented collection.

That separation matters because a performer who recorded an advert did not consent to having their voice modelled, and conflating the two is the fastest route to a dispute.

Consent should explicitly cover recording, use and licensing for AI training purposes, as distinct from consent to be recorded.


Where our own work fits

Declaring the interest: Lifewood has run voice AI data collection since establishing dedicated voice hubs, and works across 50-plus languages and dialects.

Two observations that I think are useful regardless of supplier.

The first is that studio consistency is a logistics problem before it is an audio problem. Booking the same room, the same chain and the same performer across six weeks in a market where professional studio capacity is limited is the actual constraint. In high-resource languages there is a booking market. For a Sylheti or Wolof voice, the room and the performer may need to be assembled rather than hired, and that assembly is the part of the project that runs late.

The second is that phonetic script design is where multilingual TTS projects most often under-specify. A client brief typically specifies hours, voice character and delivery format, and rarely specifies phonetic coverage. For a language with an existing balanced corpus that omission is survivable. For a language where no such corpus exists, someone has to build the coverage target from the phoneme inventory up, and if nobody does, the gaps only become visible when the synthesised voice mispronounces a common construction.

Both of these are arguments for scoping the linguistics before booking the studio, which is the reverse of how these projects are usually planned.


A specification checklist

Architecture first. Zero-shot, fine-tune or foundation model determines the hours requirement before anything else.

WAV, and state the sample rate, with 48 kHz as the target for new collection and a plan for resampling if combining with existing corpora.

Phonetic coverage targets, built from the target language's inventory, with scripts reviewed against them before recording.

Full metadata schema including signal chain and room profile, not just speaker demographics.

Utterance and word-level time alignment plus phonetic annotation.

Talent contracted for the full programme.

Reference playback and directed sessions with named responsibility for drift monitoring.

Emotional range specified upfront if expressive synthesis is the target.

Consent covering AI training use explicitly, separate from any voice-over engagement.

A five to ten hour pilot before committing the full programme.


Key takeaways

  • Session one and session forty need to be acoustically and performatively indistinguishable, because the model learns recording variance alongside the intended voice.
  • Crowdsourced and home studio recordings introduce variance in room acoustics, microphone position, vocal warmup and performance register that becomes part of the trained voice.
  • The old standard was roughly 24 hours of single-speaker read audio, producing a voice described as acceptable but flat.
  • Modern systems train one model on many speakers, with each speaker summarised as a speaker embedding, a vector a few hundred numbers long. The model is the instrument; the embedding says which person to be.
  • Because language and identity are separated in the model, a voice can speak languages its actor never learned, and voices can be blended into ones that never existed.
  • Hours depend on method: zero-shot cloning from seconds, fine-tuning on a multi-speaker base at one to five hours, flagship preset voices at twenty to forty studio hours, foundation models at 100 to 300 hours or more.
  • Every gap in the data becomes a gap in the voice, which is why flagship voices still require large studio programmes.
  • ASR data is not TTS data. Background noise is training signal for recognition and contamination for synthesis, and TTS requires aligned text-speech pairs.
  • WAV lossless is the format standard; compressed formats strip intonation and emotional subtlety and produce more robotic output.
  • 48 kHz is described as the accepted standard for new TTS datasets, though widely used corpora such as LibriTTS-R at roughly 585 hours are at 24 kHz.
  • Metadata should include speaker profile, session dates, microphone and preamp chain, room acoustics profile, utterance and word-level time alignment and phonetic annotation.
  • Scripts should be reviewed against phonetic coverage targets before recording, built from the target language's own inventory rather than translated from an English balanced set.
  • Talent should be contracted for the full programme duration, since a voice cannot be substituted mid-dataset.
  • Consistency practice includes reference playback from the previous session and directed sessions monitoring drift in energy, pacing and articulation.
  • Voice-over work and voice data should be separated, with consent explicitly covering recording, use and licensing for AI training.

Sources and further reading

Frequently asked questions

It depends on the method. Zero-shot cloning works from seconds, fine-tuning on a strong multi-speaker base model typically takes one to five hours, and flagship preset voices still use twenty to forty studio hours because every gap in the data becomes a gap in the voice.

Usually not. Background noise that helps a recognition model is contamination for synthesis, and TTS requires precisely aligned text and speech pairs rather than approximate transcripts.

WAV lossless, since compressed formats strip intonation and emotional subtlety. 48 kHz is described as the accepted standard for new collection, though established corpora such as LibriTTS-R are at 24 kHz.

Because the model cannot distinguish intended voice characteristics from recording variance. Differences in room acoustics, microphone position, vocal warm-up and register across sessions are learned as part of the voice.

Published practice starts each session with reference playback from the previous one and has audio producers direct sessions while monitoring for drift in energy, pacing and articulation.

Not without separate consent. Good practice separates voice-over work from voice data entirely, with consent explicitly covering recording, use and licensing for AI training purposes.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team