Short answer. Multilingual AI voice is four distinct services sold under one name: subtitling (no voice at all), voice replacement (new narration, same visuals), synthetic dubbing (voice matched to on-screen timing and often lip movement), and voice cloning (a specific person's voice reproduced in languages they do not speak). Cost, consent requirements and failure modes differ sharply between them. The three things that decide whether the output is usable are the same across all four: a native-speaker review pass that can reject and rewrite, a pronunciation lexicon for names and product terms, and documented consent for any real person's voice. Duration drift — target languages running longer or shorter than the source — is the technical problem that catches most teams.
Voice is where a competently localised video most often gives itself away. The words are right, the delivery is not: an accent that belongs to the wrong region, a brand name pronounced as if read from a spreadsheet, a sentence squeezed into an English-shaped gap it does not fit.
This guide covers the four service levels, what each requires, and how to specify a multilingual voice programme that holds up across dozens of markets.
The four levels, and what each is for
| Level | What you get | Consent burden | Typical use |
|---|---|---|---|
| Subtitling | Timed text, original audio | None | Low-priority markets; B2B where the source language is understood |
| Voice replacement | New narration in the target language, visuals unchanged | Voice talent agreement only | The default for most markets and most content |
| Synthetic dubbing | Voice matched to on-screen timing, sometimes with lip adjustment | Talent plus, if visuals are altered, likeness consent | Presenter-led and dialogue content |
| Voice cloning | A specific person's voice speaking a language they do not | Explicit, documented, scope-limited consent from that person | Founder or spokesperson content across markets |
The most common mistake is buying level four when level two would do. Cloning a named executive's voice into eleven languages is impressive, consent-heavy, and rarely what the content needed — a well-cast local voice frequently outperforms it on trust, because listeners in that market hear someone who sounds like them rather than someone who sounds subtly wrong.
Duration drift, and why it is a production problem
The same content takes different amounts of time to say in different languages. Several European languages commonly run longer than English; several Asian languages run shorter. Over a two-minute video, that difference compounds into seconds.
Three ways it gets handled, in descending order of quality:
- Design the master with elastic sections. Segments that can stretch or compress without breaking the edit. Costs nothing at production and solves the problem everywhere downstream.
- Re-time the locale version. Adjust the edit per language. Works, and multiplies post-production effort by the number of markets.
- Compress or stretch the audio. Fast to do and audible — rushed delivery in the long languages, unnatural pauses in the short ones. This is what a cheap dubbing quote is usually buying.
Specify which one applies before production, not after the first locale comes back wrong.
The pronunciation problem
Synthetic voice mispronounces exactly the words that matter most: brand names, product names, technical terms, place names, and acronyms that are spoken as letters in one language and as a word in another.
The fix is a pronunciation lexicon — a maintained list of every term that must be said a specific way, with its phonetic representation per language. Build it once, version it, and apply it across every asset. Without one, each new video re-learns the same mistakes, and inconsistency across a campaign is more damaging than a single error, because it reads as carelessness rather than accident.
Two practical rules: include the deliberate exceptions — terms that keep their source-language pronunciation in every market — and have a native speaker confirm each entry, because a phonetic spelling that looks right to a non-speaker often is not.
Consent, likeness and voice rights
For any real person's voice, this is the gate rather than a formality.
- Explicit and scope-limited consent. What the voice will be used for, in which languages and markets, for how long, and whether it may be used for content the person has not personally reviewed. That last clause is the one people care about most once it is explained.
- Derivative and onward use. Whether the cloned voice may be reused in future campaigns, and what happens at the end of the relationship — an employee's cloned voice outliving their employment is a foreseeable dispute worth settling in advance.
- Withdrawal path that can actually be executed against specific assets, not a right in principle.
- Voice talent agreements for conventional voice work should now address synthetic use explicitly. A recording made for one campaign is not automatically training material for a voice model.
- Disclosure. Whether the synthetic nature is disclosed, decided centrally, consistent with each market's expectations — which are tightening at different speeds, so confirm the current position per market with counsel.
- Provenance per asset. Which voice, which consent record, which model and version, which reviewer.
What native-speaker review actually catches
Machine translation quality has improved enormously. It still cannot judge these, and none of them is visible in a transcript:
- Register. Formality that is correct and wrong for the audience — over-formal in a consumer ad, over-casual in a regulated context.
- Accent and regional fit. A voice that is technically the right language and audibly from somewhere else. This matters commercially in markets with strong regional identity.
- Prosody on meaning. Emphasis landing on the wrong word, turning a claim into an odd aside.
- Idiom that translated cleanly and means nothing. Common in taglines, which are the most-heard line in the asset.
- Claim legality. Whether the sentence is sayable in that market at all — a translation pass will not flag it.
- Numbers, dates and currency spoken correctly in local convention.
Require the reviewer to listen, not read. Several of these are inaudible in text and obvious in audio, which is why transcript-only QC passes assets that native speakers reject immediately.
Specifying a multilingual voice programme
| Decision | Get it in the brief |
|---|---|
| Level per market | Subtitling, voice replacement, synthetic dubbing or cloning — decided per market, not applied uniformly |
| Voice casting | Gender, age range, accent and register per market, approved before volume |
| Duration handling | Elastic master, per-locale re-timing, or audio compression — stated up front |
| Pronunciation lexicon | Who builds it, who approves each entry, how it is versioned |
| Review standard | Native-speaker listening review, with authority to reject and rewrite |
| Consent pack | Scope, duration, markets, onward use, withdrawal path |
| Deliverables | Mixed audio, stems, transcripts, captions, per-platform loudness specs |
| Provenance | Voice, consent record, model version, reviewer and date, per asset |
Red flags: a per-minute price with no review layer named; language coverage quoted as supported languages rather than reviewer headcount; no pronunciation lexicon in the process; voice cloning offered without asking who the voice belongs to; "unlimited languages" with a single QC reviewer behind it.
How Lifewood approaches this
Lifewood delivers AI-assisted voice synthesis as one stage inside a full AIGC pipeline — script and concept development, voice, visual and motion generation, brand-style transfer, assembly and final QA — rather than as a standalone dubbing service, which is what allows duration handling to be solved in the master rather than patched per locale.
The review layer is the part that decides output quality at this scale: a first-pass editor checks factual accuracy and brand voice, a second-pass reviewer validates language, cultural fit and polish, under a 95%+ accuracy SLA with timestamped approval records. Coverage runs to 50+ languages from 40+ delivery centres in 30+ countries with 56,788 contributors, so listening review is done by in-market native speakers rather than machine-translation spot checks — and the same operation runs multilingual transcription and phonetic labelling, which is exactly the capability a pronunciation lexicon depends on. The workforce behind it received 414,120 training hours during 2025.
See AIGC video production, AIGC services, multilingual data collection and low-resource speech data.
Sources and further reading
- Synthetic-media disclosure and personality-rights requirements differ by market and are changing; confirm current obligations per market with counsel.
- Companion guides: How to Scale AI Marketing Video Production in 2026 and 10 Things to Know About AI Video Production in APAC.
- Lifewood AIGC pipeline scope is published at lifewood.com/aigc-video-production.

