Short answer. Three different jobs wear the name: verification (is this true, and against what source), compliance (are we allowed to say this, here), and editorial judgement (is this worth publishing at all). They need different people, and programmes that collapse them into one role keep verification and quietly lose the other two. What makes the layer defensible rather than decorative is a written rubric, separation of producer from reviewer, a sample rate set by risk, reviewer calibration measured with a chance-corrected agreement statistic, a failure rule decided in advance, and a retained record. Without those, "human-in-the-loop" describes someone clicking approve.
Fluency is essentially solved, which is what makes current failures hard to catch: the defects that survive to publication are not garbled sentences but well-formed sentences that are wrong, not permitted, or pointless. A reviewer briefed to look for bad writing passes all three.
Three jobs, three different catches
| Job | The question answered | What it catches that the others miss |
|---|---|---|
| Verification | Is this true, and against what source? | Confident fabrications — invented statistics, misattributed quotes, plausible citations that resolve to nothing |
| Compliance | Are we allowed to say this, here? | Regulated claims, unlicensed assets, missing disclosures, market-specific prohibitions, legal red lines |
| Editorial judgement | Is this worth publishing? | Fluent, accurate, permitted content that is generic, off-strategy, or adds nothing a reader could not get elsewhere |
A subject-matter verifier is not a compliance reviewer, and neither is an editor. The three roles can be held by the same person on low-risk content, but the checks have to be separately specified, because each has a different pass condition and a different qualification behind it.
What a defensible process specifies
Each step exists because its absence produces a specific, observed failure.
- Write the rubric before the first batch. List the error dimensions that matter for this content type and define severity levels with examples. A rubric written after the first delivery is a description of what the supplier already did, and it will never fail them.
- Separate producer from reviewer. The reviewer must not be the person — or the model — that generated the draft. This is the core process requirement of ISO 17100 for translation, and it generalises: self-review reliably misses the errors that come from the producer's own assumptions.
- Set the sample rate by risk, not by convenience. Regulated claims, material factual substance and anything naming a real person get full review. High-volume low-risk content gets randomised sampling across the whole batch. Sampling the first items of each file measures the beginnings of files.
- Calibrate reviewers against each other. Have two or more reviewers score the same sample independently and measure agreement. Low agreement means the rubric is ambiguous, not that one reviewer is wrong — and an ambiguous rubric produces quality numbers that cannot be compared between batches.
- Define the failure rule in advance. State the error threshold at which a batch is rejected, and state that rejection returns the whole batch for rework rather than the sampled items. Without this, sampling becomes a way to find and fix exactly the defects that were sampled.
- Keep the record. Who reviewed what, when, against which rubric version, and what changed. This is the artefact that answers a client audit or a rights question, and it is the one part of the process that is nearly impossible to reconstruct afterwards.
Metrics that survive contact with a supplier
Two frameworks turn a subjective conversation into an arguable one. Both predate generative AI, which is their advantage — they were built for holding an outsourced language process to a standard.
MQM — score the errors, not the vibe. Multidimensional Quality Metrics originated in the EU-funded QTLaunchPad project and provides a hierarchical catalogue of error types from which an implementer selects a subset appropriate to the content. Errors are classified by dimension — accuracy, fluency, terminology, style, locale conventions — and by severity, and the score is computed from the counts. "The quality is poor" becomes "eleven major accuracy errors and four critical terminology errors per thousand words", which a supplier can dispute item by item and which can therefore be settled.
Inter-annotator agreement — measure whether the rubric is real. A quality score means something only if two competent reviewers applying the same rubric to the same content produce similar results. Cohen's kappa measures that while correcting for agreement expected by chance. Landis and Koch's 1977 paper proposed the interpretation scale still in common use — 0.61 to 0.80 described as substantial agreement, above 0.80 as almost perfect. The authors presented these as arbitrary benchmarks rather than statistical thresholds, and they remain the standard reference point. If reviewers cannot reach substantial agreement, fix the rubric before questioning the reviewers.
The trap: an accuracy percentage with no stated rubric, sample rate or severity model is not a metric. "99% accurate" can describe a character-level match, a spot check of ten items, or a full MQM pass — three things differing by orders of magnitude in cost and meaning. Ask what the denominator is.
The human contribution has a legal function, not only a quality one
In January 2025 the US Copyright Office published Part 2 of its report on copyright and artificial intelligence, addressing copyrightability. Its conclusions are narrow and consequential: human authorship remains a requirement; works generated entirely by AI are not copyrightable; and, on the basis of currently available technology, prompts alone do not give a user sufficient control over the expressive elements of the output to make them the author. What is protectable is the human contribution perceptible in the result — creative selection, coordination and arrangement of AI-generated material, and creative modification of outputs.
That reframes the review layer for a content operation. Editing is not only how a draft becomes publishable; it is a substantial part of what makes the published work ownable. A pipeline that generates and publishes behind a compliance checkbox produces assets with a weak copyright position. A pipeline where a human meaningfully selects, arranges and revises produces assets with a documented human contribution — provided the contribution was recorded rather than merely performed, which is the part teams skip.
The Office declined to create a separate registration regime for AI-assisted works, so the ordinary rules and the ordinary evidentiary burden apply. Keeping the edit history is the cheapest insurance available against that burden landing later.
What it costs, and where the cost goes
Review is usually the largest line in an AIGC programme after production, and the instinct to compress it is strong precisely because generation got cheap. The useful framing: generation cost fell substantially while review cost did not fall at all — a human still reads at human speed — so review's share of the total rises even as the total falls. A programme budgeting review as a fixed percentage of production will systematically underfund it.
- Volume drives review cost linearly; risk drives it steeply. Doubling output doubles reading time. Adding a regulated market can multiply the per-item cost, because the reviewer pool shrinks to people qualified in that jurisdiction.
- Language count multiplies the reviewer problem, not the reviewing problem. Finding one qualified reviewer per language and keeping them available is the constraint that caps most multilingual programmes.
- Rework is the hidden line. A failed batch has to be regenerated and re-reviewed. Programmes measuring only first-pass cost misprice the ones with weak briefs.
- Calibration is a real cost and a cheap one. Half a day of reviewer calibration per quarter is materially cheaper than a quarter of incomparable quality numbers.
The demand side is not speculative: Grand View Research estimates the global data collection and labelling market at USD 6.3 billion for 2026, growing at a 28.4% compound annual rate through 2030 — a market that exists because model builders concluded human-verified data is worth paying for. Market-sizing figures are research-firm estimates rather than audited totals and should be read as directional.
How Lifewood approaches this
Lifewood applies a dual-layer process — production review followed by an independent QA pass — against a 95%+ accuracy threshold, across 50+ languages with region-native reviewers in 40+ delivery centres across 30+ countries. The structure is inherited from its AI training-data work, where the deliverable is the annotation itself and there is nowhere for an unreviewed error to hide; the same two-layer separation is applied to generated content.
That is a claim, and the right response to any such claim — this one included — is to ask who reviews (named role, qualification, market), against what written rubric, at what sample rate chosen how, and what happens on failure. Vendors who answer crisply are usually doing the work; vendors who answer with adjectives usually are not, and the distinction is available in a single meeting.
See AIGC services.
Sources and further reading
- MQM Council, the MQM error typology — originating in the EU-funded QTLaunchPad project.
- Landis and Koch (1977), "The Measurement of Observer Agreement for Categorical Data" — the kappa interpretation scale.
- ISO 17100:2015, Translation services — the producer-reviewer separation requirement.
- US Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability, January 2025.
- Grand View Research, data collection and labelling market sizing — a research-firm estimate, directional rather than audited.

