Short answer. Treat it as an audio production job, not a text-to-speech call. Model quality is rarely the limiting factor now; what makes synthetic voice sound synthetic is a short list of production defects — mispronounced names and terminology, emphasis landing on the wrong clause in long sentences, uniform pacing with no breath, flat affect across a long read, and audio never normalised to a delivery target. Four of those five are script and process problems rather than model problems, which is why changing vendor rarely fixes a synthetic-sounding read and rewriting the script often does. The fixes are a pronunciation lexicon, copy written for the ear, per-segment prosody direction, and loudness normalisation measured with ITU-R BS.1770 as EBU R 128 specifies.
This is the craft layer beneath a multilingual voice programme: what to fix when the read is technically correct and still unusable, and what the delivery specification should say so that nobody argues about it at the mix.
What actually gives synthetic voice away
Not timbre. On short, neutral copy, contemporary synthesis is frequently difficult to distinguish from a human read. What gives it away in production is an accumulation of small defects across a long piece.
| Tell | Cause | Fix |
|---|---|---|
| Mispronounced names and terms | No lexicon; the model guesses from spelling | Phonetic lexicon per language, applied at synthesis, brand terms locked |
| Wrong emphasis in long sentences | The model cannot infer which clause carries the point | Shorter sentences; explicit emphasis markup; restructure so the stressed word falls naturally |
| Metronomic pacing, no breath | One rate applied across the whole read | Vary rate per segment; insert pauses at meaning boundaries, not only at punctuation |
| Flat affect over a long piece | One setting applied to the entire script | Direct prosody per segment, as you would direct a voice actor |
| Wrong loudness or over-compression | No normalisation to a delivery target | Normalise the finished mix to a published target, per ITU-R BS.1770 measurement |
Only the second is partly a model capability question. The rest are decisions somebody did not make.
Write for the ear, not the page
Synthetic prosody is driven more directly by sentence construction than by any voice setting. Three habits carry most of the improvement:
- One idea per sentence. Subordinate clauses are where emphasis goes wrong, because the model has to guess which half of the sentence is the point.
- Put the stressed word where stress naturally falls — usually near the end of the clause. Copy written to be scanned puts it at the front, and that reads as flat.
- Read the script aloud before synthesis. Anything a human stumbles over will be worse from a model, because a human recovers mid-sentence and a model does not.
Punctuation is doing real work here. A comma is a pause instruction as much as a grammatical mark, and copy punctuated for the page produces a read punctuated for the page.
Loudness: the standard that stops the argument
Perceived loudness is not peak level, and mixing to peak is why one asset sounds quiet and another gets turned down automatically at the destination. Two documents settle this and they work together.
ITU-R BS.1770 defines the measurement algorithm for perceived programme loudness and true-peak level. EBU R 128 defines the practice built on that measurement: normalise to a target programme loudness, with defined tolerances and a limit on true peak, so material from different sources sits at a consistent level without riding the fader per item.
Five operational rules belong in the delivery specification rather than in a mix engineer's judgement:
- Normalise to the destination's published target, not to whatever sounded right in the room. Broadcast, streaming platforms and podcast distributors publish different targets.
- Measure integrated loudness across the whole programme. Short-term measurement leads to over-compressing passages that are supposed to be quiet.
- Respect the true-peak ceiling. Inter-sample peaks that pass a sample-peak meter can still clip after lossy encoding — the most common cause of distortion that appears only after upload.
- Normalise after the full mix, once voice, music and effects are balanced. Normalising the voice track alone and then adding music invalidates the measurement.
- Re-measure per language variant. Different scripts have different durations and dynamic profiles, so a localised variant does not inherit the master's compliance.
The production pipeline
Steps two and three account for most of the quality difference between two teams using the same model.
- Write the script for the ear, per the rules above, and read it aloud.
- Build the pronunciation lexicon. Every product name, person, place and technical term, with explicit phonetics per language. This is a durable brand asset: built once, reused across every asset and every language, and the highest-return item in this list.
- Segment and direct. Break the script into segments and give each a direction — pace, emphasis, energy, pause before and after. Treating a five-minute read as one synthesis call produces a five-minute read that sounds like one setting.
- Choose the voice per market, not once. Voice suitability is cultural: a read that lands as warm and authoritative in one market can read as informal or wrong in another. This is an in-market judgement, not a central one.
- Synthesise several takes per segment and select. Selection is faster than parameter tuning and produces better results, as it does for image and video generation.
- Edit as audio, not as text. Assemble in a DAW; adjust pauses, fix breaths, sit the voice against music and effects. This is ordinary audio post, and it is where synthetic reads become natural.
- Normalise and check true peak on the finished mix, per variant.
- Attach consent records, marking and disclosure. Which voice model, which consent applies and until when, plus machine-readable marking at export.
Consent and disclosure, briefly
Two distinct duties, and satisfying one does not satisfy the other.
Consent arises whenever the voice resembles an identifiable real person — whether it was cloned from their recordings or merely sounds like them. Tennessee's ELVIS Act, effective 1 July 2024, made voice a protected property right of every individual rather than only of performers, and reached the tools used to replicate it. In union production, SAG-AFTRA agreements require separate written consent for synthetic voice and digital replicas rather than treating the original engagement as covering them.
Disclosure is a labelling duty that attaches to the output. Under Article 50 of the EU AI Act, applying from 2 August 2026, synthetic audio must be marked in a machine-readable format, with disclosure of deepfake content; China's labelling Measures have required explicit and implicit labelling of synthesised audio since 1 September 2025.
The scoping question to settle first: does this voice resemble an identifiable real person? If yes, the work needs documented consent covering both the creation of the voice model and its use, with defined term and territory and an agreed disposition of the model at expiry. If no, the consent question falls away and only marking and disclosure remain.
When to use a human voice anyway
Synthetic voice is what makes wide language coverage economically possible. Being explicit about where it is the wrong choice is more useful than enthusiasm in either direction.
- Synthetic for high-volume, factual, narration-led content with a short shelf life: product walkthroughs, training modules, catalogue video, localised variants of an approved master, internal communication.
- Human where the read carries the brand: flagship films, campaign work with emotional register, anything where a specific performance is the point, and long-form content with a multi-year shelf life.
- Human where the speaker is a real, named person. A synthetic version of a real voice raises consent and disclosure duties a real recording simply does not.
- Hybrid in the common middle: human in the top revenue markets and for hero assets, synthetic across the long tail, with the same script, terminology and loudness specification applied to both so the set stays coherent.
How Lifewood approaches this
Voice is one layer of the master-and-layers architecture used across Lifewood's localisation work: the picture is locked, the stems are separated, and voice is swapped per language against fixed timing. Everything above is that swap done properly — lexicon per language, direction per segment, in-market voice selection, normalisation per variant, and consent and disclosure recorded per asset.
Lifewood produces both synthetic and human voice across 50+ languages with region-native reviewers assessing each variant, on a dual-layer human-in-the-loop process held to a 95%+ accuracy threshold, delivered from 40+ delivery centres across 30+ countries. The recommendation given most often is the hybrid pattern above, because it is usually a better film for less money than committing entirely to either.
Sources and further reading
- EBU R 128, Loudness normalisation and permitted maximum level of audio signals — European Broadcasting Union.
- Recommendation ITU-R BS.1770, Algorithms to measure audio programme loudness and true-peak audio level.
- Tennessee's ELVIS Act (HB 2091 / SB 2096), effective 1 July 2024 — voice as a protected property right.
- SAG-AFTRA, artificial intelligence provisions and member resources — consent requirements for synthetic voice and digital replicas.
- EU Artificial Intelligence Act, Article 50; China's Measures for Labeling of AI-Generated Synthetic Content, in force 1 September 2025.
- Companion guide: Multilingual AI Voice Production — the four service levels, duration drift and native-speaker review.

