Short answer. You localise the source once and the variants many times, by separating what is language-independent (picture edit, music, sound design, motion graphics rig) from what is language-dependent (script, voice, on-screen text, reading speed, cultural references, legal disclosures). Cost then scales with the number of language-dependent layers, not with the number of languages. Decide per market whether the deliverable is subtitles, voiceover, dub or a re-shot variant, and put an in-market reviewer on every language, because the failure mode at scale is a fluent sentence saying something the brand never authorised.
Key takeaways
- A master built with live text layers and separated audio stems turns a copy change into one edit and fifty swaps, instead of fifty re-renders.
- CSA Research's third global survey of 8,709 consumers in 29 countries found 76% prefer buying with information in their own language and 40% will not buy from a site in another language at all.
- Netflix's public subtitle specification caps lines at 42 characters, two lines on screen, a 20 characters-per-second reading speed, and a seven-second maximum duration.
- ISO 17100 requires a second, independent person to review every translation against the source; a model checking its own output does not meet that bar.
- The EU's AI Act Article 50 requires machine-readable labelling of synthetic audio, image, video and text from 2 August 2026; China has had comparable rules since 1 September 2025.
Why does localisation break between the tenth and the fifteenth language?
Localisation programmes break when the master file was built as a finished film rather than as a template, so every later change multiplies across every language instead of updating once.
The symptoms are consistent: version drift, where one language's cut is two frames longer than the master and nobody can say why; approval deadlock, where regional offices each hold a veto on files they received at different times; and re-rendering, where a late change to one line of copy becomes fifty exports instead of fifty text swaps. None of these are translation problems — they are architecture problems that surface as translation problems, usually because text was baked into the picture, voiceover was glued to the timeline, or a music bed was ducked against an English narration track that no longer exists once the narration is Japanese.
The commercial stakes matter, because localisation budgets are argued as cost. CSA Research's third global "Can't Read, Won't Buy" survey — 8,709 consumers across 29 countries, each surveyed in their market's official language — found 76% of online shoppers prefer to buy with information in their own language and 40% will not buy from a site in another language at all. In an unlocalised market, a company is reaching close to none of the buyers who insist on their own language.
How do you separate the master from the language layers?
The master is every element of a video that does not change with language — picture edit, music bed, sound design, and the motion graphics rig; a layer is any element with a defined swap procedure per market, such as voiceover, on-screen text or subtitles. At build time, every element gets sorted into one or the other.
| Element | Language-dependent? | How to build it |
|---|---|---|
| Picture edit, cutaways, pacing | No — unless a shot is culturally unusable | Lock once; budget any market-specific replacement separately |
| Music bed and sound design | No | Deliver as stems; a mixed track cannot be re-balanced against longer narration |
| Narration and voiceover | Yes | Record or synthesise against master timing, not the source waveform |
| On-screen titles and lower thirds | Yes | Live text in a motion template with expansion headroom — never burned in |
| UI or product screens on camera | Yes, if the product is localised | Composite over a tracked placeholder so each locale swaps cleanly |
| Subtitles and captions | Yes | Sidecar files (SRT/TTML) unless the platform forces a burn-in |
| Legal disclosures, pricing, claims | Yes, and jurisdiction-dependent | A per-market claims matrix — the layer that creates real liability |
| Currency, dates, units, formats | Yes | Data-driven fields, not typed strings |
The highest-leverage row is on-screen text. Burned-in titles convert every copy change into a re-render across every language; live text converts the same change into one edit and a batch. As a build rule, give every text layer at least 30% horizontal headroom, because several languages routinely run longer than English for the same sentence, and a template that only fits the source language gets redesigned mid-project.
What is the eight-stage localisation pipeline?
The pipeline runs from locking the source picture through translation, timing, voice production and in-context review to platform delivery, and the two stages teams skip — terminology lock and in-context review — are the two that produce the most expensive failures.
- Lock the source and freeze the picture. Nothing starts until the master edit is approved, because every change after this point multiplies by the number of target languages.
- Extract a structured script with timing. Not a transcript — a segmented script with in and out timecodes, speaker attribution, on-screen text captured separately from spoken lines, and a note on every segment marking whether timing is rigid or elastic.
- Lock terminology and the claims matrix. A glossary of product and feature names and legally controlled phrases, alongside which claims are permitted in which market — a translator is not the right person to be discovering a regulatory limit mid-project.
- Translate and adapt, transcreating where the line is doing work. Straight translation suits instructional and factual copy; hooks, humour, wordplay and taglines need transcreation from intent, decided per segment.
- Fit the script to time before recording anything. Adapted copy is checked against subtitle reading speed and breath-and-pace length before recording, because fitting after recording means re-recording, and fitting after mixing means re-mixing.
- Produce voice — human, synthetic or mixed — chosen per market and per asset, with the rights position documented at this stage.
- Assemble, then review in context, in every language, with a native speaker of that market watching the finished cut, because reviewing strings in a spreadsheet does not catch a subtitle covering a logo or a phrase that is correct and tonally wrong.
- Deliver per platform, with provenance and labels attached. Each destination has its own aspect ratio, caption format, loudness target and metadata, and any variant carrying a synthetic voice needs its marking obligation attached at this stage rather than retrofitted per market.
Which localisation mode should you use for each market?
The choice is subtitles, voiceover, full voice replacement, lip-synced dub, or a market-specific re-shoot, and these differ by roughly an order of magnitude in cost, so defaulting to dubbing everything spends the budget where it buys the least.
| Mode | Relative cost | Best fit |
|---|---|---|
| Subtitles only | Lowest | Subtitle-tolerant markets; short-shelf-life social; assets where the visual carries the message |
| Voiceover, source audible underneath | Low | Documentary, testimonial and interview content where the speaker's authenticity matters |
| Full voice replacement, not lip-synced | Medium | Narration-led explainers, training, walkthroughs — most enterprise video |
| Lip-synced dub | High | On-camera presenters in dubbing-preferring markets; long-shelf-life brand films |
| Market-specific re-shoot | Highest | Where casting, setting or a regulated claim makes the source unusable |
Subtitle timing is where intentions meet arithmetic. Netflix publishes its English timed-text specification openly, and it is a reasonable reference even for teams delivering elsewhere: a maximum of 42 characters per line, no more than two lines on screen, a minimum event duration of five-sixths of a second, a maximum of seven seconds, and an adult reading speed of 20 characters per second. Languages expand against that ceiling at different rates, so a sentence that sits comfortably in one language is unreadable in another at identical timing — a script-fitting problem to solve before recording, not a subtitling problem to solve at the end.
What does a properly reviewed localisation actually look like?
A properly reviewed localisation has a second, independent person check the target against the source before delivery, following a defined error typology and a pass threshold agreed before work starts.
ISO 17100, the international standard for translation services, has revision as its central process requirement: after translation, a second competent person who is not the translator compares the target against the source. A workflow where a model translates and the same model or the same person checks its own output does not meet that bar, whatever the deliverable is called. MQM — Multidimensional Quality Metrics — supplies a hierarchical error typology rather than a single score: errors are classified by dimension (accuracy, fluency, terminology, style, locale conventions) and by severity, which turns "the German is bad" into a count of specific, arguable defects.
Four operational rules follow: define the pass threshold before work starts, in errors per thousand words at each severity, and make it contractual; sample honestly, randomised across the whole delivery rather than the first ten minutes of each file; escalate to full review on failure, re-reviewing the failed batch rather than accepting a corrected sample; and keep the reviewer in-market, because a fluent speaker abroad catches grammar and misses register, and register is what a brand is buying. Programmes weighing AI dubbing quality against human voice actors run into this same review question from the audio side.
Should you build this in-house or use a partner?
Build it in-house at three to five languages with stable messaging, where the tooling is commodity and the reviewer network is manageable directly; bring in a partner once language count and recurrence make reviewer coverage the binding constraint rather than translation itself.
At fifty languages the binding constraint is not translation — it is having a qualified, in-market reviewer available for every language on every release, indefinitely. That is why localisation programmes quietly shrink to the eight languages a team can actually review. Lifewood's AIGC video production work runs on the partner side of that line: 100+ languages with region-native reviewers, 40+ delivery centres across 30+ countries, 56,000+ registered contributors, and a dual-layer human-in-the-loop review process held to a 95%+ accuracy threshold, part of the wider AIGC services portfolio. Bangladesh alone accounts for 414,120 of Lifewood's training hours in 2025, reflecting the scale of a single in-country reviewer pool. Teams evaluating that trade-off in more detail can compare providers in the buyer's guide to AIGC video production companies or review the broader key factors in AI video localization for 2026. If the constraint is tooling rather than reviewer coverage, buying tooling is genuinely the cheaper fix.