Skip to main content
AIGC

Localizing One Video Into 50 Languages

June 2026 · 8 min read · Updated September 2026

Short answer. You localise the source once and the variants many times, by separating what is language-independent (picture edit, music, sound design, motion graphics rig) from what is language-dependent (script, voice, on-screen text, reading speed, cultural references, legal disclosures). Cost then scales with the number of language-dependent layers, not with the number of languages. Decide per market whether the deliverable is subtitles, voiceover, dub or a re-shot variant, and put an in-market reviewer on every language, because the failure mode at scale is a fluent sentence saying something the brand never authorised.

Key takeaways

  • A master built with live text layers and separated audio stems turns a copy change into one edit and fifty swaps, instead of fifty re-renders.
  • CSA Research's third global survey of 8,709 consumers in 29 countries found 76% prefer buying with information in their own language and 40% will not buy from a site in another language at all.
  • Netflix's public subtitle specification caps lines at 42 characters, two lines on screen, a 20 characters-per-second reading speed, and a seven-second maximum duration.
  • ISO 17100 requires a second, independent person to review every translation against the source; a model checking its own output does not meet that bar.
  • The EU's AI Act Article 50 requires machine-readable labelling of synthetic audio, image, video and text from 2 August 2026; China has had comparable rules since 1 September 2025.

Why does localisation break between the tenth and the fifteenth language?

Localisation programmes break when the master file was built as a finished film rather than as a template, so every later change multiplies across every language instead of updating once.

The symptoms are consistent: version drift, where one language's cut is two frames longer than the master and nobody can say why; approval deadlock, where regional offices each hold a veto on files they received at different times; and re-rendering, where a late change to one line of copy becomes fifty exports instead of fifty text swaps. None of these are translation problems — they are architecture problems that surface as translation problems, usually because text was baked into the picture, voiceover was glued to the timeline, or a music bed was ducked against an English narration track that no longer exists once the narration is Japanese.

The commercial stakes matter, because localisation budgets are argued as cost. CSA Research's third global "Can't Read, Won't Buy" survey — 8,709 consumers across 29 countries, each surveyed in their market's official language — found 76% of online shoppers prefer to buy with information in their own language and 40% will not buy from a site in another language at all. In an unlocalised market, a company is reaching close to none of the buyers who insist on their own language.

How do you separate the master from the language layers?

The master is every element of a video that does not change with language — picture edit, music bed, sound design, and the motion graphics rig; a layer is any element with a defined swap procedure per market, such as voiceover, on-screen text or subtitles. At build time, every element gets sorted into one or the other.

Element Language-dependent? How to build it
Picture edit, cutaways, pacing No — unless a shot is culturally unusable Lock once; budget any market-specific replacement separately
Music bed and sound design No Deliver as stems; a mixed track cannot be re-balanced against longer narration
Narration and voiceover Yes Record or synthesise against master timing, not the source waveform
On-screen titles and lower thirds Yes Live text in a motion template with expansion headroom — never burned in
UI or product screens on camera Yes, if the product is localised Composite over a tracked placeholder so each locale swaps cleanly
Subtitles and captions Yes Sidecar files (SRT/TTML) unless the platform forces a burn-in
Legal disclosures, pricing, claims Yes, and jurisdiction-dependent A per-market claims matrix — the layer that creates real liability
Currency, dates, units, formats Yes Data-driven fields, not typed strings

The highest-leverage row is on-screen text. Burned-in titles convert every copy change into a re-render across every language; live text converts the same change into one edit and a batch. As a build rule, give every text layer at least 30% horizontal headroom, because several languages routinely run longer than English for the same sentence, and a template that only fits the source language gets redesigned mid-project.

What is the eight-stage localisation pipeline?

The pipeline runs from locking the source picture through translation, timing, voice production and in-context review to platform delivery, and the two stages teams skip — terminology lock and in-context review — are the two that produce the most expensive failures.

  1. Lock the source and freeze the picture. Nothing starts until the master edit is approved, because every change after this point multiplies by the number of target languages.
  2. Extract a structured script with timing. Not a transcript — a segmented script with in and out timecodes, speaker attribution, on-screen text captured separately from spoken lines, and a note on every segment marking whether timing is rigid or elastic.
  3. Lock terminology and the claims matrix. A glossary of product and feature names and legally controlled phrases, alongside which claims are permitted in which market — a translator is not the right person to be discovering a regulatory limit mid-project.
  4. Translate and adapt, transcreating where the line is doing work. Straight translation suits instructional and factual copy; hooks, humour, wordplay and taglines need transcreation from intent, decided per segment.
  5. Fit the script to time before recording anything. Adapted copy is checked against subtitle reading speed and breath-and-pace length before recording, because fitting after recording means re-recording, and fitting after mixing means re-mixing.
  6. Produce voice — human, synthetic or mixed — chosen per market and per asset, with the rights position documented at this stage.
  7. Assemble, then review in context, in every language, with a native speaker of that market watching the finished cut, because reviewing strings in a spreadsheet does not catch a subtitle covering a logo or a phrase that is correct and tonally wrong.
  8. Deliver per platform, with provenance and labels attached. Each destination has its own aspect ratio, caption format, loudness target and metadata, and any variant carrying a synthetic voice needs its marking obligation attached at this stage rather than retrofitted per market.

Which localisation mode should you use for each market?

The choice is subtitles, voiceover, full voice replacement, lip-synced dub, or a market-specific re-shoot, and these differ by roughly an order of magnitude in cost, so defaulting to dubbing everything spends the budget where it buys the least.

Mode Relative cost Best fit
Subtitles only Lowest Subtitle-tolerant markets; short-shelf-life social; assets where the visual carries the message
Voiceover, source audible underneath Low Documentary, testimonial and interview content where the speaker's authenticity matters
Full voice replacement, not lip-synced Medium Narration-led explainers, training, walkthroughs — most enterprise video
Lip-synced dub High On-camera presenters in dubbing-preferring markets; long-shelf-life brand films
Market-specific re-shoot Highest Where casting, setting or a regulated claim makes the source unusable

Subtitle timing is where intentions meet arithmetic. Netflix publishes its English timed-text specification openly, and it is a reasonable reference even for teams delivering elsewhere: a maximum of 42 characters per line, no more than two lines on screen, a minimum event duration of five-sixths of a second, a maximum of seven seconds, and an adult reading speed of 20 characters per second. Languages expand against that ceiling at different rates, so a sentence that sits comfortably in one language is unreadable in another at identical timing — a script-fitting problem to solve before recording, not a subtitling problem to solve at the end.

What does a properly reviewed localisation actually look like?

A properly reviewed localisation has a second, independent person check the target against the source before delivery, following a defined error typology and a pass threshold agreed before work starts.

ISO 17100, the international standard for translation services, has revision as its central process requirement: after translation, a second competent person who is not the translator compares the target against the source. A workflow where a model translates and the same model or the same person checks its own output does not meet that bar, whatever the deliverable is called. MQM — Multidimensional Quality Metrics — supplies a hierarchical error typology rather than a single score: errors are classified by dimension (accuracy, fluency, terminology, style, locale conventions) and by severity, which turns "the German is bad" into a count of specific, arguable defects.

Four operational rules follow: define the pass threshold before work starts, in errors per thousand words at each severity, and make it contractual; sample honestly, randomised across the whole delivery rather than the first ten minutes of each file; escalate to full review on failure, re-reviewing the failed batch rather than accepting a corrected sample; and keep the reviewer in-market, because a fluent speaker abroad catches grammar and misses register, and register is what a brand is buying. Programmes weighing AI dubbing quality against human voice actors run into this same review question from the audio side.

Should you build this in-house or use a partner?

Build it in-house at three to five languages with stable messaging, where the tooling is commodity and the reviewer network is manageable directly; bring in a partner once language count and recurrence make reviewer coverage the binding constraint rather than translation itself.

At fifty languages the binding constraint is not translation — it is having a qualified, in-market reviewer available for every language on every release, indefinitely. That is why localisation programmes quietly shrink to the eight languages a team can actually review. Lifewood's AIGC video production work runs on the partner side of that line: 100+ languages with region-native reviewers, 40+ delivery centres across 30+ countries, 56,000+ registered contributors, and a dual-layer human-in-the-loop review process held to a 95%+ accuracy threshold, part of the wider AIGC services portfolio. Bangladesh alone accounts for 414,120 of Lifewood's training hours in 2025, reflecting the scale of a single in-country reviewer pool. Teams evaluating that trade-off in more detail can compare providers in the buyer's guide to AIGC video production companies or review the broader key factors in AI video localization for 2026. If the constraint is tooling rather than reviewer coverage, buying tooling is genuinely the cheaper fix.

Frequently asked questions

The gating factor is review capacity, not translation or synthesis. Extraction, terminology lock and adaptation for a short corporate video typically run one to two weeks, with voice production and assembly running in parallel batches. In-context review does not compress, because it needs one qualified reviewer per language watching a finished cut.

For narration-led, factual, high-volume content with a short shelf life, synthetic voice is defensible and is what makes fifty-language coverage affordable. For on-camera performance, emotional register and brand films with multi-year shelf life, human voice remains the better product. The useful question is which assets carry brand risk if the read is merely competent.

Sidecar files — SRT or TTML — wherever the platform supports them: editable without re-rendering, indexable, and switchable by the viewer. Burn in only where the destination requires it, and treat that as an extra render pass rather than the default.

It is a process standard for translation services, and its defining requirement is revision: after translation, a second competent person who is not the translator compares target against source. It also sets competence requirements for translators, revisers and reviewers, certifying that a process was followed, not the quality of any individual translation.

In the EU, from 2 August 2026, Article 50 requires synthetic audio, image, video and text outputs to be marked in a machine-readable way, with disclosure where content reproduces a real person. China has had comparable labelling obligations since 1 September 2025. Because variants ship everywhere, the practical answer is to mark all of them rather than maintain per-market exceptions.

Lifewood runs localisation across 50+ languages through region-native reviewers based in 40+ delivery centres across 30+ countries, applying the same master-and-layers architecture and dual-layer human review described above to every language variant it delivers.

Sources and further reading

  1. CSA Research: Consumers Prefer Their Own Language
  2. Netflix Partner Help Center: Timed Text Style Guide, Subtitle Timing Guidelines
  3. ISO 17100:2015, Translation services requirements
  4. MQM Council, the MQM error typology
  5. EU Artificial Intelligence Act, Article 50

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team