LIFEWOOD
Ready100
AIGC

How to Run an AIGC Pilot That Actually Predicts Something

Short answer. Design the pilot so it can fail. The standard AI content pilot succeeds and predicts nothing, for two reasons: it uses flagship content that receives senior attention no…

Lifewood Data Technology · August 2026 · 7 min read

Download PDF

Short answer. Design the pilot so it can fail. The standard AI content pilot succeeds and predicts nothing, for two reasons: it uses flagship content that receives senior attention no production run will ever get, and it is judged on whether the output looks good rather than against a written threshold agreed before anything is produced. A pilot that predicts includes your worst-formatted source material, your smallest market and your most regulated claim; it runs at production attention levels; it measures review effort as a primary output, not as overhead; and it states in advance what result would cause you to walk away.

The pattern is consistent enough to be worth naming. A pilot is scoped as a handful of representative assets, briefed carefully, produced by whoever is best on the team, reviewed by people who are genuinely interested because it is new, and judged on whether the results look good. They do. The programme is approved at fifty times the volume and, within a quarter, the review queue is a bottleneck, the small markets are producing rejects, and someone is asking why the output does not resemble the pilot.

Nothing dishonest happened. The pilot measured a different system from the one that was later built. This guide covers how to design a pilot that measures the right system, what to record while it runs, and what to do with the ambiguous result you will most likely get.


Why do most pilots succeed and predict nothing?

Five specific artefacts, in rough order of how much damage they do.

The attention artefact. A pilot receives senior attention per asset that a production run cannot sustain. This is the single largest source of over-prediction, and it is invisible unless you ask who did the work.

The easy-case artefact. Pilots use clean source material and flagship content. Catalogues contain the awkward third of the inventory nobody wants to brief, and that third is where quality goes.

Review effort treated as overhead. Review hours are the variable that decides whether volume is feasible. Pilots routinely absorb them into general enthusiasm rather than counting them, and vendor quotes usually omit the line entirely.

No stated threshold. Judged on impression, every pilot passes. Judged against a written error threshold at a defined sample rate, many do not — which is the information you were trying to buy.

Generation tested in isolation. Rights clearance, provenance, platform specifications and delivery are where scale actually hurts. A pilot that stops at "the video looks good" has not touched any of them.


How do you design a pilot that can fail?

Steps one and two are non-negotiable. Without them, the rest is a demonstration.

  1. Write the pass criteria first. Error threshold by severity at a defined sample rate; maximum review hours per asset; maximum rejection rate; first-submission delivery compliance. Sign it off before briefing anyone. If you cannot state what failure looks like, you are not running a test.
  2. Select the hard cases deliberately. Your worst-formatted source material, your smallest or most linguistically distant market, your most regulated claim, your tightest brand constraint, and one asset type nobody enjoys producing. Add two easy items as a control, not as the sample.
  3. Brief the way you actually brief. Use your real template at the real level of detail, including its ambiguities. A pilot briefed better than production tests a process you will not run.
  4. Cap the attention. Agree a per-asset time budget consistent with production volume, and ask the supplier to staff the pilot with the people who will staff the volume. Ask explicitly; the answer is informative either way.
  5. Run the full chain. Brief, produce, clear rights, review, apply provenance and labels, deliver to real platform specifications. Every step you skip is a cost you discover after signing.
  6. Score blind against the rubric. Two reviewers, vendor or model identity concealed where possible, scored by error type and severity using a defined typology such as MQM. Check reviewer agreement before believing any of the scores.
  7. Measure effort, not only output. Review hours per asset, rework hours, takes per usable asset, brief-to-delivery time. Multiply by planned volume.
  8. Write the decision down against the criteria. State which were met, which were not, and what would have to change. A recommendation with no comparison to the stated thresholds has quietly reverted to impression.

What should a pilot measure?

Metric How to capture it What it predicts
Errors per asset, by severity Rubric-scored review of every pilot asset Whether the quality bar is reachable at all
Review hours per asset Timed honestly, including the reading you would skip under pressure Whether the volume is staffable — usually the binding constraint
Takes per usable asset Generated candidates against accepted outputs Real production cost, as opposed to quoted unit cost
Rejection and rework rate Assets returned for regeneration Brief quality as much as supplier quality
Brief-to-delivery time Wall clock, including approvals Whether the programme fits your publishing cadence
Delivery compliance Assets meeting platform specs, captions, loudness and labels on first submission How much hidden work sits after "the content is done"
Reviewer agreement Two reviewers on overlapping items Whether any of the other numbers are trustworthy

Reviewer agreement is the load-bearing row. If two reviewers disagree about what counts as an error, error counts are incomparable between batches and between suppliers, and every other metric in the table is provisional.

The extrapolation that decides most programmes takes two minutes:

Annual review load (FTE) = Review hours per asset × Planned annual assets ÷ Annual working hours per reviewer

If the answer exceeds the headcount you have, the programme does not scale at the quality bar you set, and no amount of cheaper generation changes that. This single calculation prevents the most common failure mode in AI content operations, and it is why review effort belongs in the pilot's outputs rather than its overheads.


What to ask while the pilot is running

A pilot also tests how a supplier answers questions under scrutiny, which predicts the working relationship better than the deliverables do. Four probes are specific to pilots rather than to procurement generally.

  • Who worked on this, and will they work on our volume? The most useful question available, and the one most likely to produce a revealing pause.
  • Show us an asset that failed your own QA, and why. A supplier with no failures has no QA, or is unwilling to show it.
  • How do you handle a market where the model is weak? The answer separates suppliers who have operated in low-resource languages from those who have not.
  • What would you refuse to produce? A supplier with no boundaries will accept a brief that creates a rights or compliance problem you own.

Ask the review questions verbatim too — who reviews, against what rubric, at what sample rate, and what happens on failure. Those four are covered in more depth in the guides on human-in-the-loop AIGC review and annotation accuracy standards and SLAs.


What do you do with an ambiguous result?

Pilots rarely produce a clean pass or fail. The common outcome is that quality met the bar while review effort exceeded the budget, or that most markets passed and one did not. That is a useful result, and it should not be resolved by rounding up.

Result The right response
Quality passed, effort failed Narrow the scope, not the quality bar — fewer assets, fewer languages, or a tighter template that reduces review load per asset
Most markets passed, one failed Tier that market: heavier review, human authorship, or exclusion. Language-capability tiering, discovered empirically
Passed, but with the supplier's best people Re-run a smaller second round staffed as production would be. The single most predictive follow-up available
Failed on brief ambiguity, not production Fix the brief template and re-run. Rework traceable to unclear briefs is your problem, and a different supplier will not solve it

The one response that is always wrong is to approve the rollout on the basis that the numbers were "close enough" to criteria that were written down precisely so that closeness would not be a judgement call.


How Lifewood approaches this

Lifewood runs pilots on this structure and prefers the hard-case version, because a pilot that passes on easy content produces a programme that disappoints on real content — a worse outcome for a supplier than an early no. The brief we ask for most often is the most awkward asset in the catalogue, in the market the buyer is least confident about, with the review line quoted separately so it can be argued with rather than absorbed.

The reason the hard case is affordable to test is delivery footprint: 50+ languages and 40+ delivery centres across 30+ countries mean the smallest market in a pilot can be staffed in-market rather than approximated. See AIGC services, the QA process and the delivery methodology.


Sources and further reading

  • The MQM error typology, severity-weighted error scoring — MQM Council.
  • Landis, J.R. and Koch, G.G., "The Measurement of Observer Agreement for Categorical Data", Biometrics 33(1), 1977 — the basis for interpreting reviewer agreement.
  • AI Risk Management Framework (AI RMF 1.0) — National Institute of Standards and Technology, January 2023, on measurement and documentation practice.
  • ISO 17100:2015, translation services requirements, including revision by a second person — International Organization for Standardization.

Frequently asked questions

Small in volume and hard in composition — commonly ten to twenty assets. Composition matters far more than size: worst-formatted source material, smallest or most linguistically distant market, most regulated claim, tightest brand constraint, plus a couple of easy items as a control. Twenty representative-average assets will pass and predict nothing.

Written before the pilot starts, covering four things: an error threshold by severity at a defined sample rate, a maximum review-hours-per-asset budget, a maximum rejection rate, and first-submission delivery compliance. Criteria decided after seeing the output are a description of the output.

Almost always the attention artefact plus the easy-case artefact. The pilot received per-asset senior attention that production cannot sustain, on content chosen because it was clean. Re-run a small second round on awkward content, staffed the way production will be staffed, and the gap usually becomes visible immediately.

Yes, on an identical brief and an identical hard-case set, scored blind against the same rubric. Sequential pilots are hard to compare because the brief and the reviewers' expectations both drift. Running them in parallel is more work in one week and saves considerably more later.

Review hours per asset at your required quality bar. Generation cost is quoted and comparable; review effort is neither, and it decides whether the volume you are planning is staffable. Multiply it by planned annual volume before making any decision.

Long enough to include the full chain — brief, production, rights clearance, review, provenance and delivery to real platform specifications — which is usually two to four weeks. Pilots compressed into a few days test generation only, and generation is not where programmes fail at scale.

Because an overall rating cannot be argued with or compared. A severity-weighted typology such as MQM makes each judgement attributable to a specific error class, which lets two reviewers disagree productively and lets two suppliers be compared on the same axis.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team