LIFEWOOD
Ready100
AIGC

What Human-in-the-Loop Review Actually Does

Short answer. Three different jobs wear the name: verification (is this true, and against what source), compliance (are we allowed to say this, here), and editorial judgement (is this…

Lifewood Data Technology · August 2026 · 7 min read

Download PDF

Short answer. Three different jobs wear the name: verification (is this true, and against what source), compliance (are we allowed to say this, here), and editorial judgement (is this worth publishing at all). They need different people, and programmes that collapse them into one role keep verification and quietly lose the other two. What makes the layer defensible rather than decorative is a written rubric, separation of producer from reviewer, a sample rate set by risk, reviewer calibration measured with a chance-corrected agreement statistic, a failure rule decided in advance, and a retained record. Without those, "human-in-the-loop" describes someone clicking approve.

Fluency is essentially solved, which is what makes current failures hard to catch: the defects that survive to publication are not garbled sentences but well-formed sentences that are wrong, not permitted, or pointless. A reviewer briefed to look for bad writing passes all three.


Three jobs, three different catches

Job The question answered What it catches that the others miss
Verification Is this true, and against what source? Confident fabrications — invented statistics, misattributed quotes, plausible citations that resolve to nothing
Compliance Are we allowed to say this, here? Regulated claims, unlicensed assets, missing disclosures, market-specific prohibitions, legal red lines
Editorial judgement Is this worth publishing? Fluent, accurate, permitted content that is generic, off-strategy, or adds nothing a reader could not get elsewhere

A subject-matter verifier is not a compliance reviewer, and neither is an editor. The three roles can be held by the same person on low-risk content, but the checks have to be separately specified, because each has a different pass condition and a different qualification behind it.


What a defensible process specifies

Each step exists because its absence produces a specific, observed failure.

  1. Write the rubric before the first batch. List the error dimensions that matter for this content type and define severity levels with examples. A rubric written after the first delivery is a description of what the supplier already did, and it will never fail them.
  2. Separate producer from reviewer. The reviewer must not be the person — or the model — that generated the draft. This is the core process requirement of ISO 17100 for translation, and it generalises: self-review reliably misses the errors that come from the producer's own assumptions.
  3. Set the sample rate by risk, not by convenience. Regulated claims, material factual substance and anything naming a real person get full review. High-volume low-risk content gets randomised sampling across the whole batch. Sampling the first items of each file measures the beginnings of files.
  4. Calibrate reviewers against each other. Have two or more reviewers score the same sample independently and measure agreement. Low agreement means the rubric is ambiguous, not that one reviewer is wrong — and an ambiguous rubric produces quality numbers that cannot be compared between batches.
  5. Define the failure rule in advance. State the error threshold at which a batch is rejected, and state that rejection returns the whole batch for rework rather than the sampled items. Without this, sampling becomes a way to find and fix exactly the defects that were sampled.
  6. Keep the record. Who reviewed what, when, against which rubric version, and what changed. This is the artefact that answers a client audit or a rights question, and it is the one part of the process that is nearly impossible to reconstruct afterwards.

Metrics that survive contact with a supplier

Two frameworks turn a subjective conversation into an arguable one. Both predate generative AI, which is their advantage — they were built for holding an outsourced language process to a standard.

MQM — score the errors, not the vibe. Multidimensional Quality Metrics originated in the EU-funded QTLaunchPad project and provides a hierarchical catalogue of error types from which an implementer selects a subset appropriate to the content. Errors are classified by dimension — accuracy, fluency, terminology, style, locale conventions — and by severity, and the score is computed from the counts. "The quality is poor" becomes "eleven major accuracy errors and four critical terminology errors per thousand words", which a supplier can dispute item by item and which can therefore be settled.

Inter-annotator agreement — measure whether the rubric is real. A quality score means something only if two competent reviewers applying the same rubric to the same content produce similar results. Cohen's kappa measures that while correcting for agreement expected by chance. Landis and Koch's 1977 paper proposed the interpretation scale still in common use — 0.61 to 0.80 described as substantial agreement, above 0.80 as almost perfect. The authors presented these as arbitrary benchmarks rather than statistical thresholds, and they remain the standard reference point. If reviewers cannot reach substantial agreement, fix the rubric before questioning the reviewers.

The trap: an accuracy percentage with no stated rubric, sample rate or severity model is not a metric. "99% accurate" can describe a character-level match, a spot check of ten items, or a full MQM pass — three things differing by orders of magnitude in cost and meaning. Ask what the denominator is.


The human contribution has a legal function, not only a quality one

In January 2025 the US Copyright Office published Part 2 of its report on copyright and artificial intelligence, addressing copyrightability. Its conclusions are narrow and consequential: human authorship remains a requirement; works generated entirely by AI are not copyrightable; and, on the basis of currently available technology, prompts alone do not give a user sufficient control over the expressive elements of the output to make them the author. What is protectable is the human contribution perceptible in the result — creative selection, coordination and arrangement of AI-generated material, and creative modification of outputs.

That reframes the review layer for a content operation. Editing is not only how a draft becomes publishable; it is a substantial part of what makes the published work ownable. A pipeline that generates and publishes behind a compliance checkbox produces assets with a weak copyright position. A pipeline where a human meaningfully selects, arranges and revises produces assets with a documented human contribution — provided the contribution was recorded rather than merely performed, which is the part teams skip.

The Office declined to create a separate registration regime for AI-assisted works, so the ordinary rules and the ordinary evidentiary burden apply. Keeping the edit history is the cheapest insurance available against that burden landing later.


What it costs, and where the cost goes

Review is usually the largest line in an AIGC programme after production, and the instinct to compress it is strong precisely because generation got cheap. The useful framing: generation cost fell substantially while review cost did not fall at all — a human still reads at human speed — so review's share of the total rises even as the total falls. A programme budgeting review as a fixed percentage of production will systematically underfund it.

  • Volume drives review cost linearly; risk drives it steeply. Doubling output doubles reading time. Adding a regulated market can multiply the per-item cost, because the reviewer pool shrinks to people qualified in that jurisdiction.
  • Language count multiplies the reviewer problem, not the reviewing problem. Finding one qualified reviewer per language and keeping them available is the constraint that caps most multilingual programmes.
  • Rework is the hidden line. A failed batch has to be regenerated and re-reviewed. Programmes measuring only first-pass cost misprice the ones with weak briefs.
  • Calibration is a real cost and a cheap one. Half a day of reviewer calibration per quarter is materially cheaper than a quarter of incomparable quality numbers.

The demand side is not speculative: Grand View Research estimates the global data collection and labelling market at USD 6.3 billion for 2026, growing at a 28.4% compound annual rate through 2030 — a market that exists because model builders concluded human-verified data is worth paying for. Market-sizing figures are research-firm estimates rather than audited totals and should be read as directional.


How Lifewood approaches this

Lifewood applies a dual-layer process — production review followed by an independent QA pass — against a 95%+ accuracy threshold, across 50+ languages with region-native reviewers in 40+ delivery centres across 30+ countries. The structure is inherited from its AI training-data work, where the deliverable is the annotation itself and there is nowhere for an unreviewed error to hide; the same two-layer separation is applied to generated content.

That is a claim, and the right response to any such claim — this one included — is to ask who reviews (named role, qualification, market), against what written rubric, at what sample rate chosen how, and what happens on failure. Vendors who answer crisply are usually doing the work; vendors who answer with adjectives usually are not, and the distinction is available in a single meeting.

See AIGC services.


Sources and further reading

  • MQM Council, the MQM error typology — originating in the EU-funded QTLaunchPad project.
  • Landis and Koch (1977), "The Measurement of Observer Agreement for Categorical Data" — the kappa interpretation scale.
  • ISO 17100:2015, Translation services — the producer-reviewer separation requirement.
  • US Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability, January 2025.
  • Grand View Research, data collection and labelling market sizing — a research-firm estimate, directional rather than audited.

Frequently asked questions

Not necessarily. The phrase spans a range from a person clicking approve to a two-pass editorial process with a documented rubric, and both are described in marketing material with the same words. The distinguishing questions are whether a rubric exists in writing, whether the reviewer is independent of the producer, and whether failure returns the whole batch for rework.

The number matters less than the denominator. Require the rubric, the severity model, the sample rate and the sampling method alongside any percentage, and state the threshold as errors per thousand units at each severity rather than as a single accuracy figure. A threshold with no rubric behind it cannot be failed, which is why it is usually offered.

For some checks, usefully — mechanical ones such as forbidden terms, missing disclosures, figures absent from an approved source pack, and format compliance. Not for verification or editorial judgement, where the failure mode is a fluent, plausible error of exactly the kind a model is prone to producing and therefore poor at detecting. Model review is a boundary filter that saves reviewer time, not a replacement for the reviewer.

It can make the human contribution protectable. The US Copyright Office's January 2025 position is that human authorship is required, that fully AI-generated material is not copyrightable, and that prompts alone do not supply authorship — but that creative selection, arrangement and modification perceptible in the result are protectable. The practical requirement is recording the contribution, not merely making it.

It depends entirely on content type, risk tier and whether the reviewer is verifying facts against sources or only checking language. The useful planning move is to measure your own throughput per content type in the first month and plan from that, because published benchmarks assume a rubric and a risk profile that are unlikely to be yours.

Reviewer identity and qualification, rubric version, sample selection method and rate, the scored results, what was changed, the failure decision, and the date. Retained per batch and independent of the file, because file metadata does not survive an ordinary distribution pipeline.

Usually because the rubric is ambiguous rather than because quality moved. Have two reviewers score the same sample independently and compute agreement; if it falls below the substantial range, the numbers from different batches were never comparable in the first place, and the fix is a clearer rubric with worked examples.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team