LIFEWOOD
Ready100
AIGC

AIGC Video: What to Fix in the Prompt and What to Fix in Post

Short answer. In generative video production the single most consequential craft decision is which defects you fix by regenerating and which you fix in post — and the rule is that…

Lifewood Data Technology · August 2026 · 8 min read

Download PDF

Short answer. In generative video production the single most consequential craft decision is which defects you fix by regenerating and which you fix in post — and the rule is that regeneration is stochastic while post is deterministic. Re-prompting a shot to correct one problem produces a new shot in which everything else has also changed slightly, so the fix has to be re-reviewed in full and may break continuity with the shots either side. A post fix changes exactly what you point it at and nothing else. That makes post the correct home for anything localised and exact — logos, typography, legal copy, colour matching, timing, stabilisation, captions, audio — and regeneration the correct response only to defects that live in the content of the frame itself. Teams that try to prompt their way to a finished asset spend most of their budget re-reviewing shots they had already approved.

The published guidance on scaled AI video production is largely about the supply chain: gates, throughput, brand locking, localisation economics. This piece is about the layer underneath it — how an individual sequence is actually produced, reviewed and repaired — which is where the per-asset cost is decided.


What has to be locked before any generation

Generation is the expensive, irreversible step. Everything that can be decided in advance should be, because changing it later means regenerating shots rather than editing them.

  • Communication objective, audience, runtime, aspect ratio and channel. These constrain shot count and pacing, and reworking them after generation invalidates the whole sequence.
  • Visual style and mandatory brand elements, as reference frames rather than adjectives. A style guide written in prose produces a different interpretation on every generation.
  • The approved claim set. What the video is permitted to assert, agreed before anyone writes narration.
  • A storyboard or shot list. Not optional. Generating before there is a shot list is how a production discovers in the edit that it has forty beautiful clips and no sequence.
  • The continuity contract — the explicit list of what must remain identical across shots: character identity, product appearance, wardrobe, palette, environment, camera language, lighting direction, and any on-screen text.

The continuity contract is the artefact most productions skip and the one that determines the review burden. Anything on it becomes a checkable property at review; anything not on it becomes a matter of opinion at review, which is slower and less consistent.

Note also the structural constraint that makes shot-based production necessary in the first place: generative video is produced as a coherent block, and coherence degrades as the block gets longer, so long runtimes are reached by chaining rather than by one generation. Chaining is a splice rather than a memory, and identity, lighting and camera behaviour drift at every seam. The consequence — write in shots, generate each shot separately against locked references, assemble in an editor — is covered in why AI video clips have length limits.


What belongs in post, always

Element Why it does not belong in the generation Post treatment
Logos and brand marks Generators approximate; approximation of a trademark is unusable Composited from the asset file
Exact typography and titles Character-level accuracy is not reliable, and legibility is a design decision Titled in the edit, in a text layer
Legal copy and disclaimers Must be exact and often market-specific Composited, per market
Colour matching across shots Each generation lands in its own grade Graded as a sequence
Timing and pacing Only judgeable against the cut, not the clip Trimmed and re-ordered in the edit
Stabilisation and minor cleanup Deterministic repairs Standard post tools
Captions and subtitles Must be timed, accurate and replaceable per market Separate layer
Audio, voice and mix Pronunciation, emotion, rights and brand fit are their own workflow Sound design, separately

The rule underneath the table: anything that must be exact, and anything that must be replaceable per market, belongs in a layer rather than in a pixel. Burned-in text is the single most common reason a locale version has to be rebuilt from scratch.


What genuinely requires regeneration

Some defects live in the content of the frame and cannot be composited away:

  • Subject anatomy — hands, faces, limb counts, eye lines
  • Object geometry and physical plausibility
  • Motion that reads as wrong: sliding feet, impossible articulation, unnatural weight
  • Lighting direction inconsistent with the scene or with the adjacent shot
  • Scene composition that does not cut with what precedes it
  • Temporal inconsistency within the shot: an object that changes shape, a garment that shifts

Before regenerating, ask the cheaper question first: can the shot be shortened or reframed instead? A defect that occupies the last eight frames is removed by a trim. A defect in the corner is removed by a punch-in. Both are deterministic and neither risks the rest of the shot. Trim, reframe, then regenerate — in that order, because each step is an order of magnitude cheaper than the next.

When regeneration is unavoidable, hold the references and the seed constant where the model supports it, change one variable, and re-review the shot against the continuity contract in full rather than checking only the thing you fixed. The changed variable is not the only thing that moved.


How to select shots

Teams generate several candidates per shot and pick the strongest, which is correct. The mistake is in how the picking is done.

Do not judge shots in isolation. A clip that is beautiful alone can be unusable in sequence — the light comes from the wrong side, the subject faces the wrong way after a cut, the energy is wrong for the beat it lands on, the product looks a different colour than in the shot before. Selection should happen on a timeline with the neighbouring shots in place, even as rough placeholders.

Select against the contract, not against taste. The continuity contract makes selection a checkable process that two different people perform the same way. Without it, selection is a series of individual preferences and the sequence drifts.

Keep the runners-up. The second-choice take is what you reach for when a later shot changes and the first choice no longer cuts. Discarding candidates to save storage is a false economy given how much a regeneration costs.


The shot-level review checklist

Review by defect class rather than as a general impression, because a general impression finds the striking problems and misses the systematic ones.

Class Check
Subject integrity Hands, faces, anatomy, eye line, identity consistent with the reference
Object and scene Product accuracy, geometry, reflections, physical plausibility
Temporal Does anything change shape, colour or position without cause across the shot
Text Any text in frame is either correct or absent — never approximate
Continuity Lighting direction, palette, wardrobe, environment against the adjacent shots
Motion Weight, contact, articulation, camera behaviour
Audio sync Lip sync where applicable; effects landing on the right frames
Brand Palette, clear space, product representation, prohibited elements
Claims and culture Anything asserted, and anything a specific market would read differently

Then a separate delivery pass on the finished asset: resolution, frame rate, loudness, colour space, safe areas, caption timing, file format and platform specification. These are deterministic and should be automated; the classes above are not and should not be.

High-risk content — regulated claims, likeness, anything with legal exposure — needs a named approver from legal, compliance or a subject-matter role on top of both.


Keeping language layers modular

One video usually becomes many. Separate the visual master from the language layers so a market can be added without regenerating anything:

  • Subtitles, captions, voiceover, on-screen text and end cards live in replaceable layers. If any of them is baked into the render, every market pays for a re-render.
  • Design the master with elastic sections. Narration duration varies by language, and a master timed frame-for-frame to one language forces a re-time for every other.
  • Native review at every level. Pronunciation, timing, terminology, register, cultural cues — and whether the visual itself is appropriate for that market, which is a question no translation step asks.
  • Version naming that survives the campaign. One concept across several aspect ratios, durations, languages and offers produces a matrix, and asset management is what stops the matrix becoming a folder nobody can navigate.

The economics of this layer are covered in producing AI marketing video at scale.


How Lifewood approaches this

Lifewood produces AIGC video as a managed service on a shot-based architecture: continuity contract and reference frames locked before generation, candidates selected on a timeline rather than in isolation, exact elements composited in post rather than prompted, and shot-level review run by defect class with a named reviewer recorded against the asset.

Localisation is handled as adaptation from a signed-off master with language kept in replaceable layers, which is where the delivery footprint decides what is practical: 50+ languages, 40+ delivery centres across 30+ countries and 56,788 registered contributors give in-market native review in markets where the alternative is machine translation with a spot check. Engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement. The AI-data heritage runs to 2004, with the current company established in 2018.

See AIGC video production, AIGC services and multilingual data collection.


Sources and further reading

  • Google Search Central, Guidance on AI-generated content.
  • NIST, Reducing Risks Posed by Synthetic Content — on provenance and transparency for generated media.
  • Companion guides: Why AI Video Clips Have Length Limits and Producing AI Marketing Video at Scale.

Frequently asked questions

No. Generative models approximate glyphs and marks, and an approximated trademark or an approximate disclaimer is unusable regardless of how good it looks. Composite them in post from the actual asset files, in a layer that can be replaced per market — which also removes the most common reason a locale version has to be rebuilt.

When the defect is in the content of the frame — anatomy, object geometry, implausible motion, lighting that contradicts the adjacent shot, temporal inconsistency within the shot. Try trimming and reframing first, because both are deterministic and cheap. Regeneration is stochastic: it changes everything in the shot, not only the problem, so it requires a full re-review and can break continuity either side.

Some tools will produce a long sequence from one instruction, and the result is generally unusable for enterprise delivery. Brand consistency, exact text, claim accuracy, continuity across cuts and platform delivery specifications all require staged production, and none of them are properties a single generation can be asked to guarantee.

Inconsistency across time and across cuts — identity, objects, lighting, palette and text that hold within a shot and drift between shots. It is the reason a clip can look impressive alone and fail in a finished sequence, and it is why selection should happen on a timeline with the neighbouring shots in place rather than clip by clip.

By defect class rather than by general impression: subject integrity, object and scene accuracy, temporal consistency, text, continuity against adjacent shots, motion, audio sync, brand conformance, and claims or cultural reading. Then a separate automated delivery pass for resolution, frame rate, loudness, safe areas, caption timing and platform specification.

Keep the visual master fixed and vary only the language layers — subtitles, voiceover, on-screen text and end cards — none of which should be baked into the render. Design the master with elastic sections, because narration length differs by language, and have a native speaker review each version for pronunciation, register, terminology and whether the visual itself reads correctly in that market.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team