Short answer. Design the pilot so it can fail. The standard AI content pilot succeeds and predicts nothing, for two reasons: it uses flagship content that receives senior attention no production run will ever get, and it is judged on whether the output looks good rather than against a written threshold agreed before anything is produced. A pilot that predicts includes your worst-formatted source material, your smallest market and your most regulated claim; it runs at production attention levels; it measures review effort as a primary output, not as overhead; and it states in advance what result would cause you to walk away.
The pattern is consistent enough to be worth naming. A pilot is scoped as a handful of representative assets, briefed carefully, produced by whoever is best on the team, reviewed by people who are genuinely interested because it is new, and judged on whether the results look good. They do. The programme is approved at fifty times the volume and, within a quarter, the review queue is a bottleneck, the small markets are producing rejects, and someone is asking why the output does not resemble the pilot.
Nothing dishonest happened. The pilot measured a different system from the one that was later built. This guide covers how to design a pilot that measures the right system, what to record while it runs, and what to do with the ambiguous result you will most likely get.
Why do most pilots succeed and predict nothing?
Five specific artefacts, in rough order of how much damage they do.
The attention artefact. A pilot receives senior attention per asset that a production run cannot sustain. This is the single largest source of over-prediction, and it is invisible unless you ask who did the work.
The easy-case artefact. Pilots use clean source material and flagship content. Catalogues contain the awkward third of the inventory nobody wants to brief, and that third is where quality goes.
Review effort treated as overhead. Review hours are the variable that decides whether volume is feasible. Pilots routinely absorb them into general enthusiasm rather than counting them, and vendor quotes usually omit the line entirely.
No stated threshold. Judged on impression, every pilot passes. Judged against a written error threshold at a defined sample rate, many do not — which is the information you were trying to buy.
Generation tested in isolation. Rights clearance, provenance, platform specifications and delivery are where scale actually hurts. A pilot that stops at "the video looks good" has not touched any of them.
How do you design a pilot that can fail?
Steps one and two are non-negotiable. Without them, the rest is a demonstration.
- Write the pass criteria first. Error threshold by severity at a defined sample rate; maximum review hours per asset; maximum rejection rate; first-submission delivery compliance. Sign it off before briefing anyone. If you cannot state what failure looks like, you are not running a test.
- Select the hard cases deliberately. Your worst-formatted source material, your smallest or most linguistically distant market, your most regulated claim, your tightest brand constraint, and one asset type nobody enjoys producing. Add two easy items as a control, not as the sample.
- Brief the way you actually brief. Use your real template at the real level of detail, including its ambiguities. A pilot briefed better than production tests a process you will not run.
- Cap the attention. Agree a per-asset time budget consistent with production volume, and ask the supplier to staff the pilot with the people who will staff the volume. Ask explicitly; the answer is informative either way.
- Run the full chain. Brief, produce, clear rights, review, apply provenance and labels, deliver to real platform specifications. Every step you skip is a cost you discover after signing.
- Score blind against the rubric. Two reviewers, vendor or model identity concealed where possible, scored by error type and severity using a defined typology such as MQM. Check reviewer agreement before believing any of the scores.
- Measure effort, not only output. Review hours per asset, rework hours, takes per usable asset, brief-to-delivery time. Multiply by planned volume.
- Write the decision down against the criteria. State which were met, which were not, and what would have to change. A recommendation with no comparison to the stated thresholds has quietly reverted to impression.
What should a pilot measure?
| Metric | How to capture it | What it predicts |
|---|---|---|
| Errors per asset, by severity | Rubric-scored review of every pilot asset | Whether the quality bar is reachable at all |
| Review hours per asset | Timed honestly, including the reading you would skip under pressure | Whether the volume is staffable — usually the binding constraint |
| Takes per usable asset | Generated candidates against accepted outputs | Real production cost, as opposed to quoted unit cost |
| Rejection and rework rate | Assets returned for regeneration | Brief quality as much as supplier quality |
| Brief-to-delivery time | Wall clock, including approvals | Whether the programme fits your publishing cadence |
| Delivery compliance | Assets meeting platform specs, captions, loudness and labels on first submission | How much hidden work sits after "the content is done" |
| Reviewer agreement | Two reviewers on overlapping items | Whether any of the other numbers are trustworthy |
Reviewer agreement is the load-bearing row. If two reviewers disagree about what counts as an error, error counts are incomparable between batches and between suppliers, and every other metric in the table is provisional.
The extrapolation that decides most programmes takes two minutes:
Annual review load (FTE) = Review hours per asset × Planned annual assets ÷ Annual working hours per reviewer
If the answer exceeds the headcount you have, the programme does not scale at the quality bar you set, and no amount of cheaper generation changes that. This single calculation prevents the most common failure mode in AI content operations, and it is why review effort belongs in the pilot's outputs rather than its overheads.
What to ask while the pilot is running
A pilot also tests how a supplier answers questions under scrutiny, which predicts the working relationship better than the deliverables do. Four probes are specific to pilots rather than to procurement generally.
- Who worked on this, and will they work on our volume? The most useful question available, and the one most likely to produce a revealing pause.
- Show us an asset that failed your own QA, and why. A supplier with no failures has no QA, or is unwilling to show it.
- How do you handle a market where the model is weak? The answer separates suppliers who have operated in low-resource languages from those who have not.
- What would you refuse to produce? A supplier with no boundaries will accept a brief that creates a rights or compliance problem you own.
Ask the review questions verbatim too — who reviews, against what rubric, at what sample rate, and what happens on failure. Those four are covered in more depth in the guides on human-in-the-loop AIGC review and annotation accuracy standards and SLAs.
What do you do with an ambiguous result?
Pilots rarely produce a clean pass or fail. The common outcome is that quality met the bar while review effort exceeded the budget, or that most markets passed and one did not. That is a useful result, and it should not be resolved by rounding up.
| Result | The right response |
|---|---|
| Quality passed, effort failed | Narrow the scope, not the quality bar — fewer assets, fewer languages, or a tighter template that reduces review load per asset |
| Most markets passed, one failed | Tier that market: heavier review, human authorship, or exclusion. Language-capability tiering, discovered empirically |
| Passed, but with the supplier's best people | Re-run a smaller second round staffed as production would be. The single most predictive follow-up available |
| Failed on brief ambiguity, not production | Fix the brief template and re-run. Rework traceable to unclear briefs is your problem, and a different supplier will not solve it |
The one response that is always wrong is to approve the rollout on the basis that the numbers were "close enough" to criteria that were written down precisely so that closeness would not be a judgement call.
How Lifewood approaches this
Lifewood runs pilots on this structure and prefers the hard-case version, because a pilot that passes on easy content produces a programme that disappoints on real content — a worse outcome for a supplier than an early no. The brief we ask for most often is the most awkward asset in the catalogue, in the market the buyer is least confident about, with the review line quoted separately so it can be argued with rather than absorbed.
The reason the hard case is affordable to test is delivery footprint: 50+ languages and 40+ delivery centres across 30+ countries mean the smallest market in a pilot can be staffed in-market rather than approximated. See AIGC services, the QA process and the delivery methodology.
Sources and further reading
- The MQM error typology, severity-weighted error scoring — MQM Council.
- Landis, J.R. and Koch, G.G., "The Measurement of Observer Agreement for Categorical Data", Biometrics 33(1), 1977 — the basis for interpreting reviewer agreement.
- AI Risk Management Framework (AI RMF 1.0) — National Institute of Standards and Technology, January 2023, on measurement and documentation practice.
- ISO 17100:2015, translation services requirements, including revision by a second person — International Organization for Standardization.

