Short answer. Treat it as an audio production job, not a text-to-speech call. Model quality is rarely the limiting factor now; what makes synthetic voice sound synthetic is a short list of production defects — mispronounced names and terminology, emphasis landing on the wrong clause in long sentences, uniform pacing with no breath, flat affect across a long read, and audio never normalised to a delivery target. Four of those five are script and process problems rather than model problems, which is why changing vendor rarely fixes a synthetic-sounding read and rewriting the script often does. The fixes are a pronunciation lexicon, copy written for the ear, per-segment prosody direction, and loudness normalisation measured with ITU-R BS.1770 as EBU R 128 specifies.
Key takeaways
- Most synthetic-voice problems are script and process defects, not model limitations: mispronunciation, wrong emphasis, flat pacing, and unnormalised loudness.
- A per-language pronunciation lexicon, built once for every product name and technical term, is the single highest-return fix.
- ITU-R BS.1770 defines how programme loudness and true-peak level are measured; EBU R 128 defines the delivery practice built on that measurement.
- Consent is required whenever a synthetic voice resembles an identifiable real person, regardless of whether their recordings were used to train it.
- A hybrid approach — human voice for hero markets and flagship assets, synthetic across the long tail — usually costs less than committing entirely to either option.
What actually gives synthetic voice away?
Not timbre. On short, neutral copy, contemporary synthesis is frequently difficult to distinguish from a human read; what gives it away in production is an accumulation of small defects across a long piece.
| Tell | Cause | Fix |
|---|---|---|
| Mispronounced names and terms | No lexicon; the model guesses from spelling | Phonetic lexicon per language, applied at synthesis, brand terms locked |
| Wrong emphasis in long sentences | The model cannot infer which clause carries the point | Shorter sentences; explicit emphasis markup; restructure so the stressed word falls naturally |
| Metronomic pacing, no breath | One rate applied across the whole read | Vary rate per segment; insert pauses at meaning boundaries, not only at punctuation |
| Flat affect over a long piece | One setting applied to the entire script | Direct prosody per segment, as you would direct a voice actor |
| Wrong loudness or over-compression | No normalisation to a delivery target | Normalise the finished mix to a published target, per ITU-R BS.1770 measurement |
Only the second row is partly a model capability question. The rest are decisions somebody did not make.
How do you write a script that sounds natural when synthesized?
Synthetic prosody is driven more directly by sentence construction than by any voice setting. Three habits carry most of the improvement.
- One idea per sentence. Subordinate clauses are where emphasis goes wrong, because the model has to guess which half of the sentence is the point.
- Put the stressed word where stress naturally falls — usually near the end of the clause. Copy written to be scanned puts it at the front, and that reads as flat.
- Read the script aloud before synthesis. Anything a human stumbles over will be worse from a model, because a human recovers mid-sentence and a model does not.
Punctuation is doing real work here. A comma is a pause instruction as much as a grammatical mark, and copy punctuated for the page produces a read punctuated for the page. Building the lexicon described in the licensed voice library approach at the same time keeps pronunciation and prosody decisions in one durable asset rather than re-litigated per project.
What loudness standard should synthetic voiceover meet?
Perceived loudness is not peak level, and mixing to peak is why one asset sounds quiet and another gets turned down automatically at the destination.
ITU-R BS.1770 defines the measurement algorithm for perceived programme loudness and true-peak level. EBU R 128 defines the practice built on that measurement: normalise to a target programme loudness, with defined tolerances and a limit on true peak, so material from different sources sits at a consistent level without riding the fader per item.
Five operational rules belong in the delivery specification rather than in a mix engineer's judgement:
- Normalise to the destination's published target, not to whatever sounded right in the room. Broadcast, streaming platforms and podcast distributors publish different targets.
- Measure integrated loudness across the whole programme. Short-term measurement leads to over-compressing passages that are supposed to be quiet.
- Respect the true-peak ceiling. Inter-sample peaks that pass a sample-peak meter can still clip after lossy encoding — the most common cause of distortion that appears only after upload.
- Normalise after the full mix, once voice, music and effects are balanced. Normalising the voice track alone and then adding music invalidates the measurement.
- Re-measure per language variant. Different scripts have different durations and dynamic profiles, so a localised variant does not inherit the master's compliance.
What does a professional synthetic voiceover production pipeline look like?
Steps two and three below account for most of the quality difference between two teams using the same model.
- Write the script for the ear, per the rules above, and read it aloud.
- Build the pronunciation lexicon: every product name, person, place and technical term, with explicit phonetics per language. This is a durable brand asset, built once and reused across every asset and every language.
- Segment and direct: break the script into segments and give each a direction — pace, emphasis, energy, pause before and after. Treating a five-minute read as one synthesis call produces a five-minute read that sounds like one setting.
- Choose the voice per market, not once, because voice suitability is cultural: a read that lands as warm and authoritative in one market can read as informal or wrong in another.
- Synthesise several takes per segment and select, which is faster than parameter tuning and produces better results — as it does for image and video generation.
- Edit as audio, not as text, assembling in a DAW, adjusting pauses, fixing breaths, and sitting the voice against music and effects.
- Normalise and check true peak on the finished mix, per language variant.
- Attach consent records, marking and disclosure: which voice model, which consent applies and until when, plus machine-readable marking at export.
What consent and disclosure rules apply to synthetic voice?
Two distinct duties apply, and satisfying one does not satisfy the other.
Consent arises whenever the voice resembles an identifiable real person — whether it was cloned from their recordings or merely sounds like them. Tennessee's ELVIS Act, effective 1 July 2024, made voice a protected property right of every individual rather than only of performers, and reached the tools used to replicate it. In union production, SAG-AFTRA agreements require separate written consent for synthetic voice and digital replicas rather than treating the original engagement as covering them; the same question of who owns AI-generated video applies to the voice track inside it.
Disclosure is a labelling duty that attaches to the output. Under Article 50 of the EU AI Act, applying from 2 August 2026, synthetic audio must be marked in a machine-readable format, with disclosure of deepfake content; China's labelling measures have required explicit and implicit labelling of synthesised audio since 1 September 2025. The marking obligation sits alongside the broader question of content provenance and watermarking, and general AI content labelling law across these jurisdictions is worth reading before a launch date is set.
The scoping question to settle first is whether the voice resembles an identifiable real person. If yes, the work needs documented consent covering both the creation of the voice model and its use, with defined term and territory and an agreed disposition of the model at expiry. If no, the consent question falls away and only marking and disclosure remain.
When should you use a human voice instead of synthetic?
Synthetic voice is what makes wide language coverage economically possible; being explicit about where it is the wrong choice is more useful than enthusiasm in either direction.
- Synthetic for high-volume, factual, narration-led content with a short shelf life: product walkthroughs, training modules, catalogue video, localised variants of an approved master, internal communication.
- Human where the read carries the brand: flagship films, campaign work with emotional register, anything where a specific performance is the point, and long-form content with a multi-year shelf life.
- Human where the speaker is a real, named person. A synthetic version of a real voice raises consent and disclosure duties a real recording simply does not.
- Hybrid in the common middle: human in the top revenue markets and for hero assets, synthetic across the long tail, with the same script, terminology and loudness specification applied to both so the set stays coherent. This is the same pattern used in broader multilingual AI voice production programmes.
How does Lifewood approach synthetic voiceover production?
Voice is one layer of the master-and-layers architecture used across Lifewood's localisation work: the picture is locked, the stems are separated, and voice is swapped per language against fixed timing.
Everything above is that swap done properly: lexicon per language, direction per segment, in-market voice selection, normalisation per variant, and consent and disclosure recorded per asset. Lifewood produces both synthetic and human voice across 100+ languages with region-native reviewers assessing each variant, on a dual-layer human-in-the-loop process held to a 95%+ accuracy SLA, delivered from 40+ delivery centres across 30+ countries. The recommendation given most often is the hybrid pattern above, because it is usually a better film for less money than committing entirely to either. See AIGC video production and the wider AIGC services this voice work sits inside.