LIFEWOOD
Ready100
AIGC

Localizing One Video Into 50 Languages

Short answer. You localise the source once and the variants many times, by separating what is language-independent (picture edit, music, sound design, motion graphics rig) from what is…

Lifewood Data Technology · August 2026 · 8 min read

Download PDF

Short answer. You localise the source once and the variants many times, by separating what is language-independent (picture edit, music, sound design, motion graphics rig) from what is language-dependent (script, voice, on-screen text, reading speed, cultural references, legal disclosures). Cost then scales with the number of language-dependent layers, not with the number of languages: a master with burned-in text is fifty re-edits; a master with a live text layer is one edit and fifty swaps. Decide per market whether the deliverable is subtitles, voiceover, dub or a re-shot variant — four budgets, not four settings — and put an in-market reviewer on every language, because the failure mode at scale is not a wrong word but a fluent sentence saying something the brand never authorised.

Localising into three languages is a project. Localising the same asset into fifty is a system, and most teams discover the difference somewhere between the tenth and the fifteenth language. This guide is the architecture, the pipeline order, the mode-selection decision, and the review standard that makes the word "reviewed" mean something.


Why localisation breaks between the tenth and the fifteenth language

The symptoms are consistent: version drift, where one language's cut is two frames longer than the master and nobody can say why; approval deadlock, where regional offices each hold a veto on files they received at different times; and the expensive one, re-rendering, where a late change to one line of copy becomes fifty exports instead of fifty text swaps.

None of these are translation problems. They are architecture problems that surface as translation problems. The cause is almost always that the master was built as a finished film rather than as a template — text baked into the picture, voiceover glued to the timeline, a music bed ducked against an English narration track that no longer exists once the narration is Japanese.

The commercial stakes matter, because localisation budgets are argued as cost. CSA Research's third global "Can't Read, Won't Buy" survey — 8,709 consumers across 29 countries, each surveyed in their market's official language, reported via press release rather than as a published paper — found 76% of online shoppers prefer to buy with information in their own language and 40% will not buy from a site in another language at all. In an unlocalised market, a company is not reaching a smaller share of buyers; it is reaching close to none of the ones who insist on their own language.


Separate the master from the layers

The method reduces to one discipline: at build time, decide for every element whether it changes with language. What does not change is the master. What does becomes a layer with a defined swap procedure.

Element Language-dependent? How to build it
Picture edit, cutaways, pacing No — unless a shot is culturally unusable Lock once; budget any market-specific replacement separately
Music bed and sound design No Deliver as stems; a mixed track cannot be re-balanced against longer narration
Narration and voiceover Yes Record or synthesise against master timing, not the source waveform
On-screen titles and lower thirds Yes Live text in a motion template with expansion headroom — never burned in
UI or product screens on camera Yes, if the product is localised Composite over a tracked placeholder so each locale swaps cleanly
Subtitles and captions Yes Sidecar files (SRT/TTML) unless the platform forces a burn-in
Legal disclosures, pricing, claims Yes, and jurisdiction-dependent A per-market claims matrix — the layer that creates real liability
Currency, dates, units, formats Yes Data-driven fields, not typed strings

The highest-leverage row is on-screen text. Burned-in titles convert every copy change into a re-render across every language; live text converts the same change into one edit and a batch.

Build rule: give every text layer at least 30% horizontal headroom. Several languages routinely run longer than English for the same sentence, and a template that only fits the source language gets redesigned mid-project.


The eight-stage pipeline

The order is load-bearing. The two stages teams skip — terminology lock and in-context review — are the two that produce the expensive failures, because both catch errors that are invisible in a spreadsheet of strings and obvious the moment someone watches the cut.

  1. Lock the source and freeze the picture. Nothing starts until the master edit is approved. Every change after this point multiplies by the number of target languages.

  2. Extract a structured script with timing. Not a transcript — a segmented script with in and out timecodes, speaker attribution, on-screen text captured separately from spoken lines, and a note on every segment marking whether timing is rigid or elastic. Translators cannot respect constraints they were never told about.

  3. Lock terminology and the claims matrix. A glossary of product names, feature names and legally controlled phrases, marking what must never be translated; alongside it, which claims are permitted in which market. A translator is not the right person to be discovering a regulatory limit.

  4. Translate and adapt — transcreate where the line is doing work. Straight translation is correct for instructional and factual copy; hooks, humour, wordplay and taglines need transcreation from intent. Deciding per segment which applies is a five-minute job that prevents a class of failure no later QA catches.

  5. Fit the script to time before recording anything. Adapted copy is checked against the stage-two constraints — subtitle reading speed for text, breath-and-pace length for voice. Fitting after recording means re-recording; fitting after mixing means re-mixing.

  6. Produce voice, human, synthetic or mixed. Choose per market and per asset, and document the rights position at this stage rather than later.

  7. Assemble, then review in context, in every language. Compose the variant and have a native speaker of that market watch the finished cut. Reviewing strings in a spreadsheet does not catch a subtitle covering a logo, a line landing after the cut, or a phrase that is correct and tonally wrong.

  8. Deliver per platform, with provenance and labels attached. Each destination has its own aspect ratio, caption format, loudness target and metadata. Any variant carrying a synthetic voice also carries a marking obligation — attach that metadata here rather than retrofitting it per market, because delivery is the last point where one process touches every language.


Subtitle, voiceover, dub or re-shoot — pick per market

These options differ by roughly an order of magnitude in cost, and the default of dubbing everything spends the budget where it buys the least.

Mode Relative cost Best fit
Subtitles only Lowest Subtitle-tolerant markets; short-shelf-life social; assets where the visual carries the message
Voiceover, source audible underneath Low Documentary, testimonial and interview content where the speaker's authenticity matters
Full voice replacement, not lip-synced Medium Narration-led explainers, training, walkthroughs — most enterprise video
Lip-synced dub High On-camera presenters in dubbing-preferring markets; long-shelf-life brand films
Market-specific re-shoot Highest Where casting, setting or a regulated claim makes the source unusable

Subtitle timing is where intentions meet arithmetic. Netflix publishes its English timed-text specification openly, and it is a reasonable reference even for teams delivering elsewhere: a maximum of 42 characters per line, no more than two lines on screen, a minimum event duration of five-sixths of a second, a maximum of seven seconds, and an adult reading speed of 20 characters per second. Languages expand against that ceiling at different rates, so a sentence that sits comfortably in one language is unreadable in another at identical timing. That is a script-fitting problem to solve before recording, not a subtitling problem to solve at the end.


What "reviewed" has to mean

Two references make the quality conversation concrete rather than adjectival.

ISO 17100, the international standard for translation services, has revision as its central process requirement: after translation, a second competent person who is not the translator compares the target against the source. A workflow where a model translates and the same model or the same person checks its own output does not meet that bar, whatever the deliverable is called.

MQM — Multidimensional Quality Metrics — supplies a hierarchical error typology rather than a single score. Errors are classified by dimension (accuracy, fluency, terminology, style, locale conventions) and by severity, which turns "the German is bad" into a count of specific, arguable defects.

Four operational rules follow:

  • Define the pass threshold before work starts, in errors per thousand words at each severity, and make it contractual.
  • Sample honestly — randomised across the whole delivery, not the first ten minutes of each file.
  • Escalate to full review on failure, re-reviewing the failed batch rather than accepting a corrected sample.
  • Keep the reviewer in-market. A fluent speaker abroad catches grammar and misses register, and register is what a brand is buying.

How Lifewood approaches this

At three to five languages with stable messaging, this pipeline runs in-house: the tooling is commodity, the reviewer network is manageable directly, and a vendor adds coordination cost against a problem that does not need it. The honest recommendation there is to fix the master-and-layers architecture and keep the work.

The case for a partner is language count plus recurrence. At fifty languages the binding constraint is not translation — it is having a qualified, in-market reviewer available for every language on every release, indefinitely. That is why localisation programmes quietly shrink to the eight languages a team can actually review.

That is the side Lifewood operates: 50+ languages with region-native reviewers, 40+ delivery centres across 30+ countries, 56,788 registered contributors, and a dual-layer human-in-the-loop review process held to a 95%+ accuracy threshold. See AIGC video production and AIGC services. If the constraint is tooling rather than reviewer coverage, buy tooling — that is genuinely the cheaper fix.


Sources and further reading

  • CSA Research, "Can't Read, Won't Buy" third global survey (8,709 consumers, 29 countries) — reported via press release and secondary coverage rather than as a published paper.
  • Netflix Partner Help Center, English Timed Text Style Guide — line length, reading speed and duration limits.
  • ISO 17100:2015, Translation services requirements; MQM Council, the MQM error typology.
  • EU Artificial Intelligence Act, Article 50, with the European Commission's transparency FAQ; China's Measures for Labeling of AI-Generated Synthetic Content.

Frequently asked questions

The gating factor is review capacity, not translation or synthesis. Extraction, terminology lock and adaptation for a short corporate video typically run one to two weeks, and voice production and assembly run in parallel batches. In-context review is the stage that does not compress, because it needs one qualified reviewer per language watching a finished cut. Very short quoted timelines at high language counts have usually removed that stage.

For narration-led, factual, high-volume content with a short shelf life, synthetic voice is defensible and is what makes fifty-language coverage affordable. For on-camera performance, emotional register and brand films with multi-year shelf life, human voice remains the better product. The useful question is which assets carry brand risk if the read is merely competent.

Sidecar files — SRT or TTML — wherever the platform supports them: editable without re-rendering, indexable, and switchable by the viewer. Burn in only where the destination requires it, and treat those as an extra render pass rather than the default.

It is a process standard for translation services. Its defining requirement is revision: after translation, a second competent person who is not the translator compares target against source. It also sets competence requirements for translators, revisers and reviewers. It certifies that a process was followed, not the quality of any individual translation.

In the EU, from 2 August 2026, Article 50 requires synthetic audio, image, video and text outputs to be marked in a machine-readable way, with disclosure where content reproduces a real person. China has had comparable explicit and implicit labelling obligations since 1 September 2025. Because variants are produced once and shipped everywhere, the practical answer is to mark all of them rather than maintain per-market exceptions.

Check whether the master has live text layers and separated audio stems. If it does, ten more languages is adaptation, voice and review. If text is burned in and the mix is a single stereo file, the cheapest path is usually to rebuild the master once as a template — the rebuild pays for itself around the fourth or fifth new language and keeps paying on every subsequent copy change.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team