Short answer. "99% accuracy" is not a standard; it is a number with no denominator, no task definition and no audit method behind it. A real annotation accuracy standard names four things per task type: the metric matched to the task (F1, IoU, word error rate or chance-corrected agreement), the threshold, the audit method (sample size, who draws it, gold-set injection rate) and the consequence of falling below it. Set the threshold from the downstream cost of error, not from what sounds impressive.
Key takeaways
- A usable annotation accuracy standard specifies, for every task type, the metric, the threshold, the audit method and the consequence of a miss.
- A single accuracy percentage has no denominator, no error-type breakdown and no chance correction, so the same delivery can be 95% at object level and 0% at frame level.
- The threshold should follow the downstream cost of error: safety-critical perception and product tagging do not deserve the same bar.
- Audit sample size grows with the square of the precision required, so halving the margin of error quadruples the audit cost, and rare error classes need stratified sampling.
- Low inter-annotator agreement usually signals ambiguous guidelines rather than poor annotators, and it is the earliest available warning that a taxonomy needs fixing.
Why does a single accuracy percentage mean nothing?
A single accuracy percentage means nothing because it never says what was measured, on what sample, judged by whom. Annotation contracts routinely specify a quality percentage and nothing else; both parties sign, and the disagreement arrives at the first disputed batch.
An annotation accuracy standard is a written specification that names, per task type, the metric, the threshold, the audit method and the consequence of falling below it. A bare percentage fails that definition in three specific ways.
No denominator. Accuracy of what: per item, per object, per attribute, per frame? A frame containing forty objects with two errors is 95% at object level and 0% at frame level. Both figures are defensible, and they support opposite conclusions.
No error-type breakdown. A dataset with 3% error concentrated entirely in one rare class is far worse for a model than 3% spread evenly, and identical on the headline figure.
No chance correction. On unbalanced tasks, raw agreement is inflated by the majority class. Two annotators labelling a 95%-negative dataset can agree 95% of the time while distinguishing nothing at all. Cohen's kappa is a chance-corrected agreement statistic that measures how much two raters agree beyond what random labelling would produce. It was introduced by Jacob Cohen in 1960 and is computed as:
Cohen's kappa = (Observed agreement − Expected agreement) ÷ (1 − Expected agreement)
A kappa near zero with 95% raw agreement is the classic signature, and it is invisible in any contract that specifies only "accuracy". The arithmetic behind kappa and its multi-rater cousin is worked through in inter-annotator agreement: Cohen's kappa and Krippendorff's alpha.
Which accuracy metric should an annotation SLA use for each task?
The metric should match the task: F1 with separate precision and recall for detection and extraction, IoU for boxes and cuboids, word error rate for transcription, and a chance-corrected agreement statistic for judgement tasks. One blended figure across a programme compares nothing to anything.
| Task type | Metric | Report as |
|---|---|---|
| Detection, extraction | F1, with precision and recall separately | Per class, plus confusion matrix |
| Bounding boxes, cuboids | IoU at a stated threshold | Separately for near and far range |
| Segmentation | Mean IoU | Per class, never overall only |
| Tracking | Identity switches per sequence | Sequence level, not frame level |
| Transcription | Word error rate, with a stated convention | Per language and per acoustic condition |
| Classification | F1 per class | With confusion matrix |
| Ranking, preference, RLHF | Chance-corrected agreement | Per task family |
| Sentiment, intent, moderation | Cohen's kappa or Krippendorff's alpha | Per policy category |
Two conventions must be written down or the metric is not comparable between vendors. For transcription: what counts as an error for disfluencies, numerals, code-switching and proper nouns. For detection: how truncated, occluded and ambiguous objects are treated. Perception-specific thresholds for these cases are set out in the guide to autonomous driving data annotation requirements.
Precision and recall should almost never be collapsed into F1 alone in the contract. They trade against each other, and which one you care about is a business decision. A moderation pipeline that must not miss violations wants recall; one that must not over-remove wants precision. State the priority, and set separate floors.
How should the accuracy threshold be set?
Work backwards from the consequence of an error, not forwards from ambition. The downstream use of the labels decides how high the bar must be, and paying for the highest bar everywhere wastes budget that the hard cases needed.
| Downstream use | Typical bar | Reasoning |
|---|---|---|
| Safety-critical perception | Highest, with per-class floors on rare classes | Systematic error becomes a safety failure |
| Model evaluation and benchmarks | Very high | Errors here corrupt every decision made from the benchmark |
| Foundation-model training corpus | High on consistency, tolerant of some noise | Volume partially averages random error; systematic error does not average out |
| Product metadata, tagging, search | Moderate | Errors are visible and cheap to correct downstream |
| Exploratory or internal analytics | Lower | Rework cost exceeds the value of precision |
The distinction that governs all of them: random error is diluted by volume; systematic error is amplified by it. A guideline ambiguity that makes every annotator label the same case the same wrong way does not average out; it teaches the model a rule. So the threshold should be paired with a requirement that error be characterised, not just counted. A vendor reporting error rate without error type is reporting half the measurement.
How should the quality audit be designed?
Four decisions belong in the contract: who draws the sample, how large it is, whether measurement is continuous or at delivery, and who adjudicates disputes. Leaving any of them to be settled later turns the first disagreement into a commercial argument.
Who draws the sample. If the vendor selects the audit sample, the audit measures the vendor's selection. Sample selection should be random, and drawn or verifiable by you.
How large the sample is. For a proportion estimate at a given confidence and margin of error, the standard formula is:
n ≈ z² × p(1 − p) ÷ e²
where z is the confidence multiplier (1.96 for 95%), p the expected error proportion, and e the margin of error you will accept. Two consequences worth knowing before negotiating: the required sample grows as the square of the precision you want, so halving the margin of error quadruples the audit cost; and estimating a rare error class precisely requires a much larger sample than estimating the overall rate. Rare classes usually need targeted stratified sampling rather than a bigger random draw.
Continuous or at delivery. Gold-set injection is the practice of seeding known-answer items into live annotation work at a defined rate so that quality is measured continuously rather than only at delivery. It catches drift within days. Delivery-gate auditing catches it at the end of a batch, when rework is most expensive. Require both: injection for control, gate audit for acceptance. The trade-offs between the three approaches are compared in gold sets, audit sampling and consensus.
Who adjudicates disputes. Name the process before you need it: re-audit by a third reviewer, an agreed arbiter, and a defined window. Without this, the first genuine disagreement becomes a commercial argument rather than a technical one.
What should the quality clause in an annotation SLA contain?
A workable quality clause has six parts: task definitions and guideline version, metric and threshold per task type, audit protocol, acceptance and rejection rules, a root-cause requirement, and change management. Vague versions of any of them are where disputes originate.
- Task definitions and guideline version. Quality is only measurable against a specific guideline version. Versioning is not administrative overhead here; it is the definition of the deliverable. How to write guidelines annotators can actually apply is covered in writing annotation guidelines.
- Metric, threshold and denominator per task type, including per-class floors where rare classes matter.
- Audit protocol: sample size and selection method, gold-set injection rate, who runs it, what evidence is produced.
- Acceptance and rejection. What happens to a batch below threshold: full rework, partial rework, or acceptance with credit. State turnaround for rework, because schedule impact is usually the larger cost.
- Root-cause requirement. Below-threshold batches trigger a written cause analysis and a guideline or training change, not a silent re-do. This is the clause that converts a supplier into a partner.
- Change management. How mid-project taxonomy changes are versioned, whether prior data is re-labelled or marked as an earlier version, who pays, and how the quality baseline is re-established afterwards.
Add two protective clauses that are cheap to agree at signature and impossible to obtain later: per-language or per-class reporting rather than aggregates, and retention of the audit evidence in an exportable format for the life of the engagement plus an agreed period. These clauses sit alongside the broader selection checklist in 9 criteria for choosing AI annotation services.
What should a good vendor push back on?
A vendor accepting every number you propose without discussion is a warning rather than a convenience. A good vendor challenges thresholds that cannot be met on specific classes, sample sizes that cannot detect rare errors, and guidelines that are ambiguous.
Reasonable pushback sounds like:
- "That threshold is achievable on this class and not on that one; here is why, and here is what we propose instead."
- "That sample size will not detect a 1% error rate on a rare class; you need stratified sampling."
- "Your guideline is ambiguous on this case, and agreement will be capped until it is resolved."
The third is the most valuable thing a vendor can tell you. Low inter-annotator agreement usually means the guidelines are ambiguous, not that the annotators are poor, and it is the earliest available signal that a taxonomy needs fixing, before a whole batch is labelled inconsistently. When shortlisting, it is worth testing each candidate on this behaviour; the vendors compared in top 10 AI data annotation companies vary considerably in how they handle it.
How does Lifewood set annotation quality standards?
Lifewood defines the quality standard per task at scoping rather than applying a single blended figure across a programme, because the metric that fits a cuboid does not fit a transcription and neither fits an RLHF preference ranking.
Two operational choices follow from that. Quality is measured continuously through gold-set injection rather than only at delivery, so drift is caught in days rather than at batch close. Lifewood's published standard is a 95%+ accuracy SLA and a 95%+ inter-annotator agreement threshold measured against a customer-approved gold set, with two independent review passes and timestamped approval records; the mechanics are described on the AI data validation page.
The workforce is a managed one in owned delivery centres rather than an open crowd, which matters for standards specifically, because a stable team is what makes a per-language, per-class agreement figure meaningful over time instead of a snapshot of whoever was available. Scope spans LLM data including RLHF and response evaluation across 100+ languages, computer vision, speech and NLP, conversational AI, content moderation and field collection, delivered from 40+ delivery centres across 30+ countries with 56,000+ registered contributors. The full range of AI data services is set out on the service page.