LIFEWOOD
Ready100
AIGC

Synthetic Voiceover: Quality and Loudness Standards

Short answer. Treat it as an audio production job, not a text-to-speech call. Model quality is rarely the limiting factor now; what makes synthetic voice sound synthetic is a short list…

Lifewood Data Technology · August 2026 · 7 min read

Download PDF

Short answer. Treat it as an audio production job, not a text-to-speech call. Model quality is rarely the limiting factor now; what makes synthetic voice sound synthetic is a short list of production defects — mispronounced names and terminology, emphasis landing on the wrong clause in long sentences, uniform pacing with no breath, flat affect across a long read, and audio never normalised to a delivery target. Four of those five are script and process problems rather than model problems, which is why changing vendor rarely fixes a synthetic-sounding read and rewriting the script often does. The fixes are a pronunciation lexicon, copy written for the ear, per-segment prosody direction, and loudness normalisation measured with ITU-R BS.1770 as EBU R 128 specifies.

This is the craft layer beneath a multilingual voice programme: what to fix when the read is technically correct and still unusable, and what the delivery specification should say so that nobody argues about it at the mix.


What actually gives synthetic voice away

Not timbre. On short, neutral copy, contemporary synthesis is frequently difficult to distinguish from a human read. What gives it away in production is an accumulation of small defects across a long piece.

Tell Cause Fix
Mispronounced names and terms No lexicon; the model guesses from spelling Phonetic lexicon per language, applied at synthesis, brand terms locked
Wrong emphasis in long sentences The model cannot infer which clause carries the point Shorter sentences; explicit emphasis markup; restructure so the stressed word falls naturally
Metronomic pacing, no breath One rate applied across the whole read Vary rate per segment; insert pauses at meaning boundaries, not only at punctuation
Flat affect over a long piece One setting applied to the entire script Direct prosody per segment, as you would direct a voice actor
Wrong loudness or over-compression No normalisation to a delivery target Normalise the finished mix to a published target, per ITU-R BS.1770 measurement

Only the second is partly a model capability question. The rest are decisions somebody did not make.


Write for the ear, not the page

Synthetic prosody is driven more directly by sentence construction than by any voice setting. Three habits carry most of the improvement:

  • One idea per sentence. Subordinate clauses are where emphasis goes wrong, because the model has to guess which half of the sentence is the point.
  • Put the stressed word where stress naturally falls — usually near the end of the clause. Copy written to be scanned puts it at the front, and that reads as flat.
  • Read the script aloud before synthesis. Anything a human stumbles over will be worse from a model, because a human recovers mid-sentence and a model does not.

Punctuation is doing real work here. A comma is a pause instruction as much as a grammatical mark, and copy punctuated for the page produces a read punctuated for the page.


Loudness: the standard that stops the argument

Perceived loudness is not peak level, and mixing to peak is why one asset sounds quiet and another gets turned down automatically at the destination. Two documents settle this and they work together.

ITU-R BS.1770 defines the measurement algorithm for perceived programme loudness and true-peak level. EBU R 128 defines the practice built on that measurement: normalise to a target programme loudness, with defined tolerances and a limit on true peak, so material from different sources sits at a consistent level without riding the fader per item.

Five operational rules belong in the delivery specification rather than in a mix engineer's judgement:

  • Normalise to the destination's published target, not to whatever sounded right in the room. Broadcast, streaming platforms and podcast distributors publish different targets.
  • Measure integrated loudness across the whole programme. Short-term measurement leads to over-compressing passages that are supposed to be quiet.
  • Respect the true-peak ceiling. Inter-sample peaks that pass a sample-peak meter can still clip after lossy encoding — the most common cause of distortion that appears only after upload.
  • Normalise after the full mix, once voice, music and effects are balanced. Normalising the voice track alone and then adding music invalidates the measurement.
  • Re-measure per language variant. Different scripts have different durations and dynamic profiles, so a localised variant does not inherit the master's compliance.

The production pipeline

Steps two and three account for most of the quality difference between two teams using the same model.

  1. Write the script for the ear, per the rules above, and read it aloud.
  2. Build the pronunciation lexicon. Every product name, person, place and technical term, with explicit phonetics per language. This is a durable brand asset: built once, reused across every asset and every language, and the highest-return item in this list.
  3. Segment and direct. Break the script into segments and give each a direction — pace, emphasis, energy, pause before and after. Treating a five-minute read as one synthesis call produces a five-minute read that sounds like one setting.
  4. Choose the voice per market, not once. Voice suitability is cultural: a read that lands as warm and authoritative in one market can read as informal or wrong in another. This is an in-market judgement, not a central one.
  5. Synthesise several takes per segment and select. Selection is faster than parameter tuning and produces better results, as it does for image and video generation.
  6. Edit as audio, not as text. Assemble in a DAW; adjust pauses, fix breaths, sit the voice against music and effects. This is ordinary audio post, and it is where synthetic reads become natural.
  7. Normalise and check true peak on the finished mix, per variant.
  8. Attach consent records, marking and disclosure. Which voice model, which consent applies and until when, plus machine-readable marking at export.

Consent and disclosure, briefly

Two distinct duties, and satisfying one does not satisfy the other.

Consent arises whenever the voice resembles an identifiable real person — whether it was cloned from their recordings or merely sounds like them. Tennessee's ELVIS Act, effective 1 July 2024, made voice a protected property right of every individual rather than only of performers, and reached the tools used to replicate it. In union production, SAG-AFTRA agreements require separate written consent for synthetic voice and digital replicas rather than treating the original engagement as covering them.

Disclosure is a labelling duty that attaches to the output. Under Article 50 of the EU AI Act, applying from 2 August 2026, synthetic audio must be marked in a machine-readable format, with disclosure of deepfake content; China's labelling Measures have required explicit and implicit labelling of synthesised audio since 1 September 2025.

The scoping question to settle first: does this voice resemble an identifiable real person? If yes, the work needs documented consent covering both the creation of the voice model and its use, with defined term and territory and an agreed disposition of the model at expiry. If no, the consent question falls away and only marking and disclosure remain.


When to use a human voice anyway

Synthetic voice is what makes wide language coverage economically possible. Being explicit about where it is the wrong choice is more useful than enthusiasm in either direction.

  • Synthetic for high-volume, factual, narration-led content with a short shelf life: product walkthroughs, training modules, catalogue video, localised variants of an approved master, internal communication.
  • Human where the read carries the brand: flagship films, campaign work with emotional register, anything where a specific performance is the point, and long-form content with a multi-year shelf life.
  • Human where the speaker is a real, named person. A synthetic version of a real voice raises consent and disclosure duties a real recording simply does not.
  • Hybrid in the common middle: human in the top revenue markets and for hero assets, synthetic across the long tail, with the same script, terminology and loudness specification applied to both so the set stays coherent.

How Lifewood approaches this

Voice is one layer of the master-and-layers architecture used across Lifewood's localisation work: the picture is locked, the stems are separated, and voice is swapped per language against fixed timing. Everything above is that swap done properly — lexicon per language, direction per segment, in-market voice selection, normalisation per variant, and consent and disclosure recorded per asset.

Lifewood produces both synthetic and human voice across 50+ languages with region-native reviewers assessing each variant, on a dual-layer human-in-the-loop process held to a 95%+ accuracy threshold, delivered from 40+ delivery centres across 30+ countries. The recommendation given most often is the hybrid pattern above, because it is usually a better film for less money than committing entirely to either.

See AIGC video production.


Sources and further reading

  • EBU R 128, Loudness normalisation and permitted maximum level of audio signals — European Broadcasting Union.
  • Recommendation ITU-R BS.1770, Algorithms to measure audio programme loudness and true-peak audio level.
  • Tennessee's ELVIS Act (HB 2091 / SB 2096), effective 1 July 2024 — voice as a protected property right.
  • SAG-AFTRA, artificial intelligence provisions and member resources — consent requirements for synthetic voice and digital replicas.
  • EU Artificial Intelligence Act, Article 50; China's Measures for Labeling of AI-Generated Synthetic Content, in force 1 September 2025.
  • Companion guide: Multilingual AI Voice Production — the four service levels, duration drift and native-speaker review.

Frequently asked questions

Usually for reasons unrelated to the model: names pronounced from spelling rather than from a lexicon, emphasis landing on the wrong clause in long sentences, uniform pacing with no breath, and audio never normalised to a delivery target. Rewriting the script for the ear and directing prosody segment by segment typically produces a larger improvement than changing vendor.

At the target published by the destination, measured with the algorithm standardised in ITU-R BS.1770 and applied per EBU R 128 practice — integrated programme loudness on the finished mix, with the true-peak ceiling respected. Targets differ between broadcast, streaming and podcast distribution, so the delivery specification should name the destination rather than a single house number.

If the voice resembles an identifiable real person, yes — and resemblance is the trigger, not whether their recordings were used. The ELVIS Act protects the voice of every individual, and union agreements require separate written consent for synthetic voice with defined scope and compensation. A wholly synthetic voice resembling nobody does not raise the consent question, though disclosure duties may still apply.

Only with their specific, documented, revocable consent covering both the creation of the voice model and each category of use, with defined term and territory and an agreed disposition of the model at expiry. Employment does not imply consent to voice replication, and a general media release signed before generative tools existed almost certainly does not cover it.

With a per-language lexicon maintained as a brand asset. Product and brand names are the hardest cases, because the correct pronunciation is a brand decision rather than a linguistic one — decide it explicitly per market, record it phonetically, and apply it at synthesis. Built once, it is reused indefinitely.

Technically yes, legally with conditions. Disclosure obligations apply where the output constitutes a deepfake or where a market requires labelling, and consumer-protection rules on deceptive practice reach undisclosed synthetic endorsement. Separately, whether it is the right creative choice depends on whether the read is carrying the brand — for hero campaign work, human voice is usually still the better product.

Match the character rather than the timbre. A single voice identity cloned across markets often reads as wrong locally, because voice suitability is cultural. Define the character — age band, energy, formality, pace — as part of the brand specification, and select per market against that definition with an in-market reviewer.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team