Short answer. You localise the source once and the variants many times, by separating what is language-independent (picture edit, music, sound design, motion graphics rig) from what is language-dependent (script, voice, on-screen text, reading speed, cultural references, legal disclosures). Cost then scales with the number of language-dependent layers, not with the number of languages: a master with burned-in text is fifty re-edits; a master with a live text layer is one edit and fifty swaps. Decide per market whether the deliverable is subtitles, voiceover, dub or a re-shot variant — four budgets, not four settings — and put an in-market reviewer on every language, because the failure mode at scale is not a wrong word but a fluent sentence saying something the brand never authorised.
Localising into three languages is a project. Localising the same asset into fifty is a system, and most teams discover the difference somewhere between the tenth and the fifteenth language. This guide is the architecture, the pipeline order, the mode-selection decision, and the review standard that makes the word "reviewed" mean something.
Why localisation breaks between the tenth and the fifteenth language
The symptoms are consistent: version drift, where one language's cut is two frames longer than the master and nobody can say why; approval deadlock, where regional offices each hold a veto on files they received at different times; and the expensive one, re-rendering, where a late change to one line of copy becomes fifty exports instead of fifty text swaps.
None of these are translation problems. They are architecture problems that surface as translation problems. The cause is almost always that the master was built as a finished film rather than as a template — text baked into the picture, voiceover glued to the timeline, a music bed ducked against an English narration track that no longer exists once the narration is Japanese.
The commercial stakes matter, because localisation budgets are argued as cost. CSA Research's third global "Can't Read, Won't Buy" survey — 8,709 consumers across 29 countries, each surveyed in their market's official language, reported via press release rather than as a published paper — found 76% of online shoppers prefer to buy with information in their own language and 40% will not buy from a site in another language at all. In an unlocalised market, a company is not reaching a smaller share of buyers; it is reaching close to none of the ones who insist on their own language.
Separate the master from the layers
The method reduces to one discipline: at build time, decide for every element whether it changes with language. What does not change is the master. What does becomes a layer with a defined swap procedure.
| Element | Language-dependent? | How to build it |
|---|---|---|
| Picture edit, cutaways, pacing | No — unless a shot is culturally unusable | Lock once; budget any market-specific replacement separately |
| Music bed and sound design | No | Deliver as stems; a mixed track cannot be re-balanced against longer narration |
| Narration and voiceover | Yes | Record or synthesise against master timing, not the source waveform |
| On-screen titles and lower thirds | Yes | Live text in a motion template with expansion headroom — never burned in |
| UI or product screens on camera | Yes, if the product is localised | Composite over a tracked placeholder so each locale swaps cleanly |
| Subtitles and captions | Yes | Sidecar files (SRT/TTML) unless the platform forces a burn-in |
| Legal disclosures, pricing, claims | Yes, and jurisdiction-dependent | A per-market claims matrix — the layer that creates real liability |
| Currency, dates, units, formats | Yes | Data-driven fields, not typed strings |
The highest-leverage row is on-screen text. Burned-in titles convert every copy change into a re-render across every language; live text converts the same change into one edit and a batch.
Build rule: give every text layer at least 30% horizontal headroom. Several languages routinely run longer than English for the same sentence, and a template that only fits the source language gets redesigned mid-project.
The eight-stage pipeline
The order is load-bearing. The two stages teams skip — terminology lock and in-context review — are the two that produce the expensive failures, because both catch errors that are invisible in a spreadsheet of strings and obvious the moment someone watches the cut.
Lock the source and freeze the picture. Nothing starts until the master edit is approved. Every change after this point multiplies by the number of target languages.
Extract a structured script with timing. Not a transcript — a segmented script with in and out timecodes, speaker attribution, on-screen text captured separately from spoken lines, and a note on every segment marking whether timing is rigid or elastic. Translators cannot respect constraints they were never told about.
Lock terminology and the claims matrix. A glossary of product names, feature names and legally controlled phrases, marking what must never be translated; alongside it, which claims are permitted in which market. A translator is not the right person to be discovering a regulatory limit.
Translate and adapt — transcreate where the line is doing work. Straight translation is correct for instructional and factual copy; hooks, humour, wordplay and taglines need transcreation from intent. Deciding per segment which applies is a five-minute job that prevents a class of failure no later QA catches.
Fit the script to time before recording anything. Adapted copy is checked against the stage-two constraints — subtitle reading speed for text, breath-and-pace length for voice. Fitting after recording means re-recording; fitting after mixing means re-mixing.
Produce voice, human, synthetic or mixed. Choose per market and per asset, and document the rights position at this stage rather than later.
Assemble, then review in context, in every language. Compose the variant and have a native speaker of that market watch the finished cut. Reviewing strings in a spreadsheet does not catch a subtitle covering a logo, a line landing after the cut, or a phrase that is correct and tonally wrong.
Deliver per platform, with provenance and labels attached. Each destination has its own aspect ratio, caption format, loudness target and metadata. Any variant carrying a synthetic voice also carries a marking obligation — attach that metadata here rather than retrofitting it per market, because delivery is the last point where one process touches every language.
Subtitle, voiceover, dub or re-shoot — pick per market
These options differ by roughly an order of magnitude in cost, and the default of dubbing everything spends the budget where it buys the least.
| Mode | Relative cost | Best fit |
|---|---|---|
| Subtitles only | Lowest | Subtitle-tolerant markets; short-shelf-life social; assets where the visual carries the message |
| Voiceover, source audible underneath | Low | Documentary, testimonial and interview content where the speaker's authenticity matters |
| Full voice replacement, not lip-synced | Medium | Narration-led explainers, training, walkthroughs — most enterprise video |
| Lip-synced dub | High | On-camera presenters in dubbing-preferring markets; long-shelf-life brand films |
| Market-specific re-shoot | Highest | Where casting, setting or a regulated claim makes the source unusable |
Subtitle timing is where intentions meet arithmetic. Netflix publishes its English timed-text specification openly, and it is a reasonable reference even for teams delivering elsewhere: a maximum of 42 characters per line, no more than two lines on screen, a minimum event duration of five-sixths of a second, a maximum of seven seconds, and an adult reading speed of 20 characters per second. Languages expand against that ceiling at different rates, so a sentence that sits comfortably in one language is unreadable in another at identical timing. That is a script-fitting problem to solve before recording, not a subtitling problem to solve at the end.
What "reviewed" has to mean
Two references make the quality conversation concrete rather than adjectival.
ISO 17100, the international standard for translation services, has revision as its central process requirement: after translation, a second competent person who is not the translator compares the target against the source. A workflow where a model translates and the same model or the same person checks its own output does not meet that bar, whatever the deliverable is called.
MQM — Multidimensional Quality Metrics — supplies a hierarchical error typology rather than a single score. Errors are classified by dimension (accuracy, fluency, terminology, style, locale conventions) and by severity, which turns "the German is bad" into a count of specific, arguable defects.
Four operational rules follow:
- Define the pass threshold before work starts, in errors per thousand words at each severity, and make it contractual.
- Sample honestly — randomised across the whole delivery, not the first ten minutes of each file.
- Escalate to full review on failure, re-reviewing the failed batch rather than accepting a corrected sample.
- Keep the reviewer in-market. A fluent speaker abroad catches grammar and misses register, and register is what a brand is buying.
How Lifewood approaches this
At three to five languages with stable messaging, this pipeline runs in-house: the tooling is commodity, the reviewer network is manageable directly, and a vendor adds coordination cost against a problem that does not need it. The honest recommendation there is to fix the master-and-layers architecture and keep the work.
The case for a partner is language count plus recurrence. At fifty languages the binding constraint is not translation — it is having a qualified, in-market reviewer available for every language on every release, indefinitely. That is why localisation programmes quietly shrink to the eight languages a team can actually review.
That is the side Lifewood operates: 50+ languages with region-native reviewers, 40+ delivery centres across 30+ countries, 56,788 registered contributors, and a dual-layer human-in-the-loop review process held to a 95%+ accuracy threshold. See AIGC video production and AIGC services. If the constraint is tooling rather than reviewer coverage, buy tooling — that is genuinely the cheaper fix.
Sources and further reading
- CSA Research, "Can't Read, Won't Buy" third global survey (8,709 consumers, 29 countries) — reported via press release and secondary coverage rather than as a published paper.
- Netflix Partner Help Center, English Timed Text Style Guide — line length, reading speed and duration limits.
- ISO 17100:2015, Translation services requirements; MQM Council, the MQM error typology.
- EU Artificial Intelligence Act, Article 50, with the European Commission's transparency FAQ; China's Measures for Labeling of AI-Generated Synthetic Content.

