Skip to main content
AIGC

AIGC Video: What to Fix in the Prompt and What to Fix in Post

June 2026 · 11 min read · Updated September 2026

Short answer. Fix in post anything that must be exact or replaceable per market — logos, typography, legal copy, colour matching, timing, stabilisation, captions and audio — and regenerate only when the defect lives in the content of the frame: anatomy, object geometry, motion, lighting direction or temporal drift. Regeneration is stochastic and changes the whole shot; a post fix is deterministic and changes only what you point it at. Trim and reframe before you regenerate.

Key takeaways

  • Regeneration is stochastic: re-prompting a shot to correct one defect produces a new shot in which everything else has also changed slightly, so the whole shot has to be re-reviewed and may no longer cut with its neighbours.
  • A post fix is deterministic: it changes exactly what it is pointed at and nothing else, which makes post the correct home for logos, typography, legal copy, colour, timing, captions and audio.
  • Only defects inside the content of the frame — anatomy, object geometry, implausible motion, lighting that contradicts the adjacent shot, or objects that change shape mid-shot — justify regenerating.
  • The order of repair is trim, then reframe, then regenerate, because each step is roughly an order of magnitude cheaper than the next.
  • Burned-in text is the single most common reason a localised version of a video has to be rebuilt from scratch, so language elements belong in replaceable layers.

What is the difference between fixing a defect in the prompt and fixing it in post?

Regeneration replaces the whole shot with a new, slightly different one, while a post fix alters only the element it targets. That difference in predictability decides which defects go where.

In generative video production the single most consequential craft decision is which defects you fix by regenerating and which you fix in post. Re-prompting a shot to correct one problem produces a new shot in which everything else has also changed slightly, so the fix has to be re-reviewed in full and may break continuity with the shots either side. A post fix changes exactly what you point it at and nothing else. Teams that try to prompt their way to a finished asset spend most of their budget re-reviewing shots they had already approved.

Criterion Fix by regeneration (the prompt) Fix in post (the edit)
Nature of the change Stochastic: a new shot in which everything moves a little Deterministic: changes only the targeted element
What it can correct Content of the frame: anatomy, geometry, motion, lighting, temporal drift Exact and layered elements: logos, type, legal copy, grade, timing, captions, audio
Review burden after the fix Full re-review of the shot against the continuity contract Check the changed element only
Continuity risk to adjacent shots High: identity, lighting and camera behaviour can drift at both seams None while the fix sits in a layer above the pixels
Replaceable per market No: the result is baked into the pixels Yes: swap the layer without touching the master
Relative cost Highest: generation time plus a full re-review Lowest: standard tools with predictable turnaround
When to use it Only after trimming and reframing have been ruled out Always, for anything that must be exact

The published guidance on scaled AI video production is largely about the supply chain: gates, throughput, brand locking, localisation economics. This piece is about the layer underneath it — how an individual sequence is actually produced, reviewed and repaired — which is where the per-asset cost is decided. For the vendor-level view of who runs that layer well, see the comparison of AIGC video production providers.

What has to be locked before any generation?

Everything that can be decided in advance should be decided before the first shot is generated, because changing it later means regenerating shots rather than editing them. Generation is the expensive, irreversible step.

  • Communication objective, audience, runtime, aspect ratio and channel. These constrain shot count and pacing, and reworking them after generation invalidates the whole sequence.
  • Visual style and mandatory brand elements, as reference frames rather than adjectives. A style guide written in prose produces a different interpretation on every generation.
  • The approved claim set. What the video is permitted to assert, agreed before anyone writes narration.
  • A storyboard or shot list. Not optional. Generating before there is a shot list is how a production discovers in the edit that it has forty beautiful clips and no sequence; AI storyboarding and previsualisation is the cheapest place to find that out.
  • The continuity contract — the explicit list of what must remain identical across shots: character identity, product appearance, wardrobe, palette, environment, camera language, lighting direction, and any on-screen text.

The continuity contract is the artefact most productions skip and the one that determines the review burden. Anything on it becomes a checkable property at review; anything not on it becomes a matter of opinion at review, which is slower and less consistent.

Note also the structural constraint that makes shot-based production necessary in the first place: generative video is produced as a coherent block, and coherence degrades as the block gets longer, so long runtimes are reached by chaining rather than by one generation. OpenAI's Sora API documentation, for example, describes single generations of 16 or 20 seconds, with longer videos built by extending a source clip up to six times to a maximum of 120 seconds. Chaining is a splice rather than a memory, and identity, lighting and camera behaviour drift at every seam. The consequence — write in shots, generate each shot separately against locked references, assemble in an editor — is covered in why AI video clips have length limits.

What belongs in post, always?

Anything that must be exact, and anything that must be replaceable per market, belongs in a layer rather than in a pixel.

Element Why it does not belong in the generation Post treatment
Logos and brand marks Generators approximate; approximation of a trademark is unusable Composited from the asset file
Exact typography and titles Character-level accuracy is not reliable, and legibility is a design decision Titled in the edit, in a text layer
Legal copy and disclaimers Must be exact and often market-specific Composited, per market
Colour matching across shots Each generation lands in its own grade Graded as a sequence
Timing and pacing Only judgeable against the cut, not the clip Trimmed and re-ordered in the edit
Stabilisation and minor cleanup Deterministic repairs Standard post tools
Captions and subtitles Must be timed, accurate and replaceable per market Separate layer
Audio, voice and mix Pronunciation, emotion, rights and brand fit are their own workflow Sound design, separately

The rule underneath the table — exact or replaceable means a layer, not a pixel — applies to anything not listed in it as well. Burned-in text is the single most common reason a locale version has to be rebuilt from scratch. This is also the mechanism behind consistent output at volume, which is why professional studios keep AI video consistent by compositing exact elements rather than prompting for them.

What genuinely requires regeneration?

Only defects that live in the content of the frame and cannot be composited away justify regenerating a shot. Even then, trimming or reframing should be tried first.

Defects that cannot be fixed in post:

  • Subject anatomy — hands, faces, limb counts, eye lines
  • Object geometry and physical plausibility
  • Motion that reads as wrong: sliding feet, impossible articulation, unnatural weight
  • Lighting direction inconsistent with the scene or with the adjacent shot
  • Scene composition that does not cut with what precedes it
  • Temporal inconsistency within the shot: an object that changes shape, a garment that shifts

Before regenerating, ask the cheaper question first: can the shot be shortened or reframed instead? A defect that occupies the last eight frames is removed by a trim. A defect in the corner is removed by a punch-in. Both are deterministic and neither risks the rest of the shot. Trim, reframe, then regenerate — in that order, because each step is an order of magnitude cheaper than the next.

When regeneration is unavoidable, hold the references and the seed constant where the model supports it, change one variable, and re-review the shot against the continuity contract in full rather than checking only the thing you fixed. The changed variable is not the only thing that moved.

How should shots be selected?

Shots should be selected on a timeline with the neighbouring shots in place, against the continuity contract rather than against taste. Generating several candidates per shot is correct; the mistake is usually in how the picking is done.

Do not judge shots in isolation. A clip that is beautiful alone can be unusable in sequence — the light comes from the wrong side, the subject faces the wrong way after a cut, the energy is wrong for the beat it lands on, the product looks a different colour than in the shot before. Selection should happen on a timeline with the neighbouring shots in place, even as rough placeholders.

Select against the contract, not against taste. The continuity contract makes selection a checkable process that two different people perform the same way. Without it, selection is a series of individual preferences and the sequence drifts.

Keep the runners-up. The second-choice take is what you reach for when a later shot changes and the first choice no longer cuts. Discarding candidates to save storage is a false economy given how much a regeneration costs.

What does a shot-level review check?

A shot-level review checks each shot by defect class rather than as a general impression, because a general impression finds the striking problems and misses the systematic ones.

Class Check
Subject integrity Hands, faces, anatomy, eye line, identity consistent with the reference
Object and scene Product accuracy, geometry, reflections, physical plausibility
Temporal Does anything change shape, colour or position without cause across the shot
Text Any text in frame is either correct or absent — never approximate
Continuity Lighting direction, palette, wardrobe, environment against the adjacent shots
Motion Weight, contact, articulation, camera behaviour
Audio sync Lip sync where applicable; effects landing on the right frames
Brand Palette, clear space, product representation, prohibited elements
Claims and culture Anything asserted, and anything a specific market would read differently

Then a separate delivery pass on the finished asset: resolution, frame rate, loudness, colour space, safe areas, caption timing, file format and platform specification. These are deterministic and should be automated; the classes above are not and should not be. The delivery pass is also where provenance metadata is attached: NIST AI 100-4, the NIST report on reducing risks from synthetic content, treats recording a piece of content's source and its history of changes as the primary technical approach to transparency for generated media.

The claims check carries weight beyond legal exposure: Google's published guidance on AI-generated content judges material on quality and helpfulness rather than on how it was produced, so a generated video is held to the same standard as a filmed one.

High-risk content — regulated claims, likeness, anything with legal exposure — needs a named approver from legal, compliance or a subject-matter role on top of both passes.

How do you keep language layers modular?

Keep the visual master fixed and put every language-dependent element in a replaceable layer, so a market can be added without regenerating anything. One video usually becomes many.

  • Subtitles, captions, voiceover, on-screen text and end cards live in replaceable layers. If any of them is baked into the render, every market pays for a re-render.
  • Design the master with elastic sections. Narration duration varies by language, and a master timed frame-for-frame to one language forces a re-time for every other.
  • Native review at every level. Pronunciation, timing, terminology, register, cultural cues — and whether the visual itself is appropriate for that market, which is a question no translation step asks.
  • Version naming that survives the campaign. One concept across several aspect ratios, durations, languages and offers produces a matrix, and asset management is what stops the matrix becoming a folder nobody can navigate.

The economics of this layer are covered in producing AI marketing video at scale, and the market-by-market mechanics in AI video localisation for global markets.

How does Lifewood approach prompt-versus-post decisions?

Lifewood produces AIGC video as a managed service on a shot-based architecture, with exact elements composited in post rather than prompted and every shot reviewed by defect class before it is assembled. Localisation is handled as adaptation from a signed-off master, with language kept in replaceable layers.

In practice the continuity contract and reference frames are locked before generation, candidates are selected on a timeline rather than in isolation, and shot-level review is run by defect class with a named reviewer recorded against the asset. This is the workflow behind Lifewood's AIGC video production service and the wider AIGC services line, where human review sits inside production rather than at the end of it.

Localisation is where the delivery footprint decides what is practical: 100+ languages, 40+ delivery centres across 30+ countries and 56,000+ registered contributors give in-market native review in markets where the alternative is machine translation with a spot check. Engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement. Lifewood was founded in 2004 and brings over two decades of AI-data delivery to the review discipline described here.

Frequently asked questions

No. Generative models approximate glyphs and marks, and an approximated trademark or an approximate disclaimer is unusable regardless of how good it looks. Composite them in post from the actual asset files, in a layer that can be replaced per market, which also removes the most common reason a locale version has to be rebuilt.

When the defect is in the content of the frame: anatomy, object geometry, implausible motion, lighting that contradicts the adjacent shot, or temporal inconsistency within the shot. Try trimming and reframing first, because both are deterministic and cheap. Regeneration changes everything in the shot, so it requires a full re-review.

Some tools will produce a long sequence from one instruction, and the result is generally unusable for enterprise delivery. Brand consistency, exact text, claim accuracy, continuity across cuts and platform delivery specifications all require staged production, and none of them are properties a single generation can be asked to guarantee.

Inconsistency across time and across cuts: identity, objects, lighting, palette and text that hold within a shot and drift between shots. It is the reason a clip can look impressive alone and fail in a finished sequence, and it is why selection should happen on a timeline with neighbouring shots in place.

By defect class rather than by general impression: subject integrity, object and scene accuracy, temporal consistency, text, continuity against adjacent shots, motion, audio sync, brand conformance, and claims or cultural reading. Then a separate automated delivery pass covers resolution, frame rate, loudness, safe areas, caption timing and platform specification.

Lifewood Data Technology produces AI-generated video as a managed service with human review built into the process: shot-level review by defect class, a named reviewer recorded against each asset, exact elements composited in post, and native-speaker review of each language version across 50+ languages from 40+ delivery centres in 30+ countries.

Sources and further reading

  1. Google Search's guidance about AI-generated content — quality and helpfulness judged regardless of how content is produced.
  2. NIST AI 100-4, Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency — provenance tracking and transparency for generated media.
  3. OpenAI, Video generation with Sora (API guide) — 16- and 20-second generations, extended up to six times to a 120-second maximum.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team