Short answer. Fix in post anything that must be exact or replaceable per market — logos, typography, legal copy, colour matching, timing, stabilisation, captions and audio — and regenerate only when the defect lives in the content of the frame: anatomy, object geometry, motion, lighting direction or temporal drift. Regeneration is stochastic and changes the whole shot; a post fix is deterministic and changes only what you point it at. Trim and reframe before you regenerate.
Key takeaways
- Regeneration is stochastic: re-prompting a shot to correct one defect produces a new shot in which everything else has also changed slightly, so the whole shot has to be re-reviewed and may no longer cut with its neighbours.
- A post fix is deterministic: it changes exactly what it is pointed at and nothing else, which makes post the correct home for logos, typography, legal copy, colour, timing, captions and audio.
- Only defects inside the content of the frame — anatomy, object geometry, implausible motion, lighting that contradicts the adjacent shot, or objects that change shape mid-shot — justify regenerating.
- The order of repair is trim, then reframe, then regenerate, because each step is roughly an order of magnitude cheaper than the next.
- Burned-in text is the single most common reason a localised version of a video has to be rebuilt from scratch, so language elements belong in replaceable layers.
What is the difference between fixing a defect in the prompt and fixing it in post?
Regeneration replaces the whole shot with a new, slightly different one, while a post fix alters only the element it targets. That difference in predictability decides which defects go where.
In generative video production the single most consequential craft decision is which defects you fix by regenerating and which you fix in post. Re-prompting a shot to correct one problem produces a new shot in which everything else has also changed slightly, so the fix has to be re-reviewed in full and may break continuity with the shots either side. A post fix changes exactly what you point it at and nothing else. Teams that try to prompt their way to a finished asset spend most of their budget re-reviewing shots they had already approved.
| Criterion | Fix by regeneration (the prompt) | Fix in post (the edit) |
|---|---|---|
| Nature of the change | Stochastic: a new shot in which everything moves a little | Deterministic: changes only the targeted element |
| What it can correct | Content of the frame: anatomy, geometry, motion, lighting, temporal drift | Exact and layered elements: logos, type, legal copy, grade, timing, captions, audio |
| Review burden after the fix | Full re-review of the shot against the continuity contract | Check the changed element only |
| Continuity risk to adjacent shots | High: identity, lighting and camera behaviour can drift at both seams | None while the fix sits in a layer above the pixels |
| Replaceable per market | No: the result is baked into the pixels | Yes: swap the layer without touching the master |
| Relative cost | Highest: generation time plus a full re-review | Lowest: standard tools with predictable turnaround |
| When to use it | Only after trimming and reframing have been ruled out | Always, for anything that must be exact |
The published guidance on scaled AI video production is largely about the supply chain: gates, throughput, brand locking, localisation economics. This piece is about the layer underneath it — how an individual sequence is actually produced, reviewed and repaired — which is where the per-asset cost is decided. For the vendor-level view of who runs that layer well, see the comparison of AIGC video production providers.
What has to be locked before any generation?
Everything that can be decided in advance should be decided before the first shot is generated, because changing it later means regenerating shots rather than editing them. Generation is the expensive, irreversible step.
- Communication objective, audience, runtime, aspect ratio and channel. These constrain shot count and pacing, and reworking them after generation invalidates the whole sequence.
- Visual style and mandatory brand elements, as reference frames rather than adjectives. A style guide written in prose produces a different interpretation on every generation.
- The approved claim set. What the video is permitted to assert, agreed before anyone writes narration.
- A storyboard or shot list. Not optional. Generating before there is a shot list is how a production discovers in the edit that it has forty beautiful clips and no sequence; AI storyboarding and previsualisation is the cheapest place to find that out.
- The continuity contract — the explicit list of what must remain identical across shots: character identity, product appearance, wardrobe, palette, environment, camera language, lighting direction, and any on-screen text.
The continuity contract is the artefact most productions skip and the one that determines the review burden. Anything on it becomes a checkable property at review; anything not on it becomes a matter of opinion at review, which is slower and less consistent.
Note also the structural constraint that makes shot-based production necessary in the first place: generative video is produced as a coherent block, and coherence degrades as the block gets longer, so long runtimes are reached by chaining rather than by one generation. OpenAI's Sora API documentation, for example, describes single generations of 16 or 20 seconds, with longer videos built by extending a source clip up to six times to a maximum of 120 seconds. Chaining is a splice rather than a memory, and identity, lighting and camera behaviour drift at every seam. The consequence — write in shots, generate each shot separately against locked references, assemble in an editor — is covered in why AI video clips have length limits.
What belongs in post, always?
Anything that must be exact, and anything that must be replaceable per market, belongs in a layer rather than in a pixel.
| Element | Why it does not belong in the generation | Post treatment |
|---|---|---|
| Logos and brand marks | Generators approximate; approximation of a trademark is unusable | Composited from the asset file |
| Exact typography and titles | Character-level accuracy is not reliable, and legibility is a design decision | Titled in the edit, in a text layer |
| Legal copy and disclaimers | Must be exact and often market-specific | Composited, per market |
| Colour matching across shots | Each generation lands in its own grade | Graded as a sequence |
| Timing and pacing | Only judgeable against the cut, not the clip | Trimmed and re-ordered in the edit |
| Stabilisation and minor cleanup | Deterministic repairs | Standard post tools |
| Captions and subtitles | Must be timed, accurate and replaceable per market | Separate layer |
| Audio, voice and mix | Pronunciation, emotion, rights and brand fit are their own workflow | Sound design, separately |
The rule underneath the table — exact or replaceable means a layer, not a pixel — applies to anything not listed in it as well. Burned-in text is the single most common reason a locale version has to be rebuilt from scratch. This is also the mechanism behind consistent output at volume, which is why professional studios keep AI video consistent by compositing exact elements rather than prompting for them.
What genuinely requires regeneration?
Only defects that live in the content of the frame and cannot be composited away justify regenerating a shot. Even then, trimming or reframing should be tried first.
Defects that cannot be fixed in post:
- Subject anatomy — hands, faces, limb counts, eye lines
- Object geometry and physical plausibility
- Motion that reads as wrong: sliding feet, impossible articulation, unnatural weight
- Lighting direction inconsistent with the scene or with the adjacent shot
- Scene composition that does not cut with what precedes it
- Temporal inconsistency within the shot: an object that changes shape, a garment that shifts
Before regenerating, ask the cheaper question first: can the shot be shortened or reframed instead? A defect that occupies the last eight frames is removed by a trim. A defect in the corner is removed by a punch-in. Both are deterministic and neither risks the rest of the shot. Trim, reframe, then regenerate — in that order, because each step is an order of magnitude cheaper than the next.
When regeneration is unavoidable, hold the references and the seed constant where the model supports it, change one variable, and re-review the shot against the continuity contract in full rather than checking only the thing you fixed. The changed variable is not the only thing that moved.
How should shots be selected?
Shots should be selected on a timeline with the neighbouring shots in place, against the continuity contract rather than against taste. Generating several candidates per shot is correct; the mistake is usually in how the picking is done.
Do not judge shots in isolation. A clip that is beautiful alone can be unusable in sequence — the light comes from the wrong side, the subject faces the wrong way after a cut, the energy is wrong for the beat it lands on, the product looks a different colour than in the shot before. Selection should happen on a timeline with the neighbouring shots in place, even as rough placeholders.
Select against the contract, not against taste. The continuity contract makes selection a checkable process that two different people perform the same way. Without it, selection is a series of individual preferences and the sequence drifts.
Keep the runners-up. The second-choice take is what you reach for when a later shot changes and the first choice no longer cuts. Discarding candidates to save storage is a false economy given how much a regeneration costs.
What does a shot-level review check?
A shot-level review checks each shot by defect class rather than as a general impression, because a general impression finds the striking problems and misses the systematic ones.
| Class | Check |
|---|---|
| Subject integrity | Hands, faces, anatomy, eye line, identity consistent with the reference |
| Object and scene | Product accuracy, geometry, reflections, physical plausibility |
| Temporal | Does anything change shape, colour or position without cause across the shot |
| Text | Any text in frame is either correct or absent — never approximate |
| Continuity | Lighting direction, palette, wardrobe, environment against the adjacent shots |
| Motion | Weight, contact, articulation, camera behaviour |
| Audio sync | Lip sync where applicable; effects landing on the right frames |
| Brand | Palette, clear space, product representation, prohibited elements |
| Claims and culture | Anything asserted, and anything a specific market would read differently |
Then a separate delivery pass on the finished asset: resolution, frame rate, loudness, colour space, safe areas, caption timing, file format and platform specification. These are deterministic and should be automated; the classes above are not and should not be. The delivery pass is also where provenance metadata is attached: NIST AI 100-4, the NIST report on reducing risks from synthetic content, treats recording a piece of content's source and its history of changes as the primary technical approach to transparency for generated media.
The claims check carries weight beyond legal exposure: Google's published guidance on AI-generated content judges material on quality and helpfulness rather than on how it was produced, so a generated video is held to the same standard as a filmed one.
High-risk content — regulated claims, likeness, anything with legal exposure — needs a named approver from legal, compliance or a subject-matter role on top of both passes.
How do you keep language layers modular?
Keep the visual master fixed and put every language-dependent element in a replaceable layer, so a market can be added without regenerating anything. One video usually becomes many.
- Subtitles, captions, voiceover, on-screen text and end cards live in replaceable layers. If any of them is baked into the render, every market pays for a re-render.
- Design the master with elastic sections. Narration duration varies by language, and a master timed frame-for-frame to one language forces a re-time for every other.
- Native review at every level. Pronunciation, timing, terminology, register, cultural cues — and whether the visual itself is appropriate for that market, which is a question no translation step asks.
- Version naming that survives the campaign. One concept across several aspect ratios, durations, languages and offers produces a matrix, and asset management is what stops the matrix becoming a folder nobody can navigate.
The economics of this layer are covered in producing AI marketing video at scale, and the market-by-market mechanics in AI video localisation for global markets.
How does Lifewood approach prompt-versus-post decisions?
Lifewood produces AIGC video as a managed service on a shot-based architecture, with exact elements composited in post rather than prompted and every shot reviewed by defect class before it is assembled. Localisation is handled as adaptation from a signed-off master, with language kept in replaceable layers.
In practice the continuity contract and reference frames are locked before generation, candidates are selected on a timeline rather than in isolation, and shot-level review is run by defect class with a named reviewer recorded against the asset. This is the workflow behind Lifewood's AIGC video production service and the wider AIGC services line, where human review sits inside production rather than at the end of it.
Localisation is where the delivery footprint decides what is practical: 100+ languages, 40+ delivery centres across 30+ countries and 56,000+ registered contributors give in-market native review in markets where the alternative is machine translation with a spot check. Engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement. Lifewood was founded in 2004 and brings over two decades of AI-data delivery to the review discipline described here.