Skip to main content
AIGC

How to Run an AIGC Pilot That Actually Predicts Something

July 2026 · 8 min read · Updated September 2026

Short answer. Design the pilot so it can fail. Most AI content pilots succeed and predict nothing because they use flagship content that gets senior attention no production run will sustain, and because they are judged on whether the output looks good rather than against a written threshold agreed beforehand. A pilot that predicts includes your hardest source material and measures review effort as a primary output, not overhead.

Key takeaways

  • A pilot that only uses clean, flagship content and senior-team attention will pass regardless of whether the production system works.
  • Pass criteria — an error threshold by severity, a maximum review-hours budget, a maximum rejection rate, and delivery compliance — must be written down before anything is produced.
  • Review hours per asset, not generation cost, is usually the number that decides whether a programme is staffable at scale.
  • Reviewer agreement between two independent reviewers determines whether every other pilot metric can be trusted or compared.
  • An ambiguous pilot result (quality passed, effort failed) should narrow the scope of the rollout, not lower the quality bar.

Why do most pilots succeed and predict nothing?

Five specific artefacts distort a pilot's result, in rough order of how much damage they do.

The attention artefact. A pilot receives senior attention per asset that a production run cannot sustain. This is the single largest source of over-prediction, and it is invisible unless you ask who did the work.

The easy-case artefact. Pilots use clean source material and flagship content. Catalogues contain the awkward third of the inventory nobody wants to brief, and that third is where quality goes.

Review effort treated as overhead. Review hours are the variable that decides whether volume is feasible. Pilots routinely absorb them into general enthusiasm rather than counting them, and vendor quotes usually omit the line entirely.

No stated threshold. Judged on impression, every pilot passes. Judged against a written error threshold at a defined sample rate, many do not — which is the information you were trying to buy.

Generation tested in isolation. Rights clearance, provenance, platform specifications and delivery are where scale actually hurts. A pilot that stops at "the video looks good" has not touched any of them. Provenance here means the record of how and when an asset was generated, reviewed and approved — the paper trail a regulator or platform can ask for later.

How do you design a pilot that can fail?

Steps one and two are non-negotiable. Without them, the rest is a demonstration.

  1. Write the pass criteria first. Error threshold by severity at a defined sample rate; maximum review hours per asset; maximum rejection rate; first-submission delivery compliance. Sign it off before briefing anyone. If you cannot state what failure looks like, you are not running a test.
  2. Select the hard cases deliberately. Your worst-formatted source material, your smallest or most linguistically distant market, your most regulated claim, your tightest brand constraint, and one asset type nobody enjoys producing. Add two easy items as a control, not as the sample.
  3. Brief the way you actually brief. Use your real template at the real level of detail, including its ambiguities. A pilot briefed better than production tests a process you will not run.
  4. Cap the attention. Agree a per-asset time budget consistent with production volume, and ask the supplier to staff the pilot with the people who will staff the volume. Ask explicitly; the answer is informative either way.
  5. Run the full chain. Brief, produce, clear rights, review, apply provenance and labels, deliver to real platform specifications. Every step you skip is a cost you discover after signing.
  6. Score blind against the rubric. Two reviewers, vendor or model identity concealed where possible, scored by error type and severity using a defined typology such as MQM. Check reviewer agreement before believing any of the scores.
  7. Measure effort, not only output. Review hours per asset, rework hours, takes per usable asset, brief-to-delivery time. Multiply by planned volume.
  8. Write the decision down against the criteria. State which were met, which were not, and what would have to change. A recommendation with no comparison to the stated thresholds has quietly reverted to impression.

What should a pilot measure?

A pilot needs to measure seven things together, and the ones that predict staffability matter more than the ones that measure polish.

Metric How to capture it What it predicts
Errors per asset, by severity Rubric-scored review of every pilot asset Whether the quality bar is reachable at all
Review hours per asset Timed honestly, including the reading you would skip under pressure Whether the volume is staffable — usually the binding constraint
Takes per usable asset Generated candidates against accepted outputs Real production cost, as opposed to quoted unit cost
Rejection and rework rate Assets returned for regeneration Brief quality as much as supplier quality
Brief-to-delivery time Wall clock, including approvals Whether the programme fits your publishing cadence
Delivery compliance Assets meeting platform specs, captions, loudness and labels on first submission How much hidden work sits after "the content is done"
Reviewer agreement Two reviewers on overlapping items Whether any of the other numbers are trustworthy

Reviewer agreement is a measure of how often two independent reviewers reach the same error judgement on the same asset; it is the load-bearing row, because if two reviewers disagree about what counts as an error, every other metric in the table is provisional and incomparable between batches or suppliers.

The extrapolation that decides most programmes takes two minutes:

Annual review load (FTE) = Review hours per asset × Planned annual assets ÷ Annual working hours per reviewer

If the answer exceeds the headcount you have, the programme does not scale at the quality bar you set, and no amount of cheaper generation changes that. This single calculation prevents the most common failure mode in AI content operations, and it is why review effort belongs in the pilot's outputs rather than its overheads. Tracking that load over time is also the core of measuring whether an AI content programme is working.

What should you ask while the pilot is running?

A pilot also tests how a supplier answers questions under scrutiny, which predicts the working relationship better than the deliverables do.

  • Who worked on this, and will they work on our volume? The most useful question available, and the one most likely to produce a revealing pause.
  • Show us an asset that failed your own QA, and why. A supplier with no failures has no QA, or is unwilling to show it.
  • How do you handle a market where the model is weak? The answer separates suppliers who have genuinely operated in difficult markets from those who have not, a distinction covered in more depth when choosing an AIGC video production provider.
  • What would you refuse to produce? A supplier with no boundaries will accept a brief that creates a rights or compliance problem you own.

Ask the review questions verbatim too — who reviews, against what rubric, at what sample rate, and what happens on failure. Those are covered in more depth in the guides on human-in-the-loop AIGC review and quality control at scale.

What do you do with an ambiguous result?

Pilots rarely produce a clean pass or fail, and that ambiguous result is still useful data — it should not be resolved by rounding up.

Result The right response
Quality passed, effort failed Narrow the scope, not the quality bar — fewer assets, fewer languages, or a tighter template that reduces review load per asset
Most markets passed, one failed Tier that market: heavier review, human authorship, or exclusion. Language-capability tiering, discovered empirically
Passed, but with the supplier's best people Re-run a smaller second round staffed as production would be. The single most predictive follow-up available
Failed on brief ambiguity, not production Fix the brief template and re-run. Rework traceable to unclear briefs is your problem, and a different supplier will not solve it

The one response that is always wrong is to approve the rollout on the basis that the numbers were "close enough" to criteria that were written down precisely so that closeness would not be a judgement call. Getting the brief itself right the first time, covered in writing a brief an AIGC team can produce from, removes one of the most common sources of an ambiguous result.

How does Lifewood approach AIGC pilots?

Lifewood runs pilots on this structure and prefers the hard-case version, because a pilot that passes on easy content produces a programme that disappoints on real content — a worse outcome for a supplier than an early no.

The brief asked for most often is the most awkward asset in the catalogue, in the market the buyer is least confident about, with the review line quoted separately so it can be argued with rather than absorbed. The reason the hard case is affordable to test is delivery footprint: 100+ languages and 40+ delivery centres across 30+ countries mean the smallest market in a pilot can be staffed in-market rather than approximated. See AIGC services and AIGC video production for how that structure carries through into a full production programme.

Frequently asked questions

Small in volume and hard in composition — commonly ten to twenty assets. Composition matters far more than size: worst-formatted source material, smallest or most linguistically distant market, most regulated claim, tightest brand constraint, plus a couple of easy items as a control. Twenty representative-average assets will pass and predict nothing.

Written before the pilot starts, covering four things: an error threshold by severity at a defined sample rate, a maximum review-hours-per-asset budget, a maximum rejection rate, and first-submission delivery compliance. Criteria decided after seeing the output are a description of the output, not a test.

Almost always the attention artefact plus the easy-case artefact. The pilot received per-asset senior attention that production cannot sustain, on content chosen because it was clean. Re-run a small second round on awkward content, staffed the way production will be staffed, and the gap usually becomes visible immediately.

Yes, on an identical brief and an identical hard-case set, scored blind against the same rubric. Sequential pilots are hard to compare because the brief and the reviewers' expectations both drift. Running them in parallel is more work in one week and saves considerably more later.

Review hours per asset at your required quality bar. Generation cost is quoted and comparable; review effort is neither, and it decides whether the volume you are planning is staffable. Multiply it by planned annual volume before making any decision.

Long enough to include the full chain — brief, production, rights clearance, review, provenance and delivery to real platform specifications — which is usually two to four weeks. Pilots compressed into a few days test generation only, and generation is not where programmes fail at scale.

Sources and further reading

  1. MQM error typology, severity-weighted error scoring — MQM Council.
  2. Landis, J.R. and Koch, G.G., "The Measurement of Observer Agreement for Categorical Data", Biometrics 33(1), 1977 — the basis for interpreting reviewer agreement.
  3. AI Risk Management Framework (AI RMF 1.0) — National Institute of Standards and Technology, January 2023, on measurement and documentation practice.
  4. ISO 17100:2015, translation services requirements — , including revision by a second person — International Organization for Standardization.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team