Skip to main content
AI Data

9 Criteria for Choosing AI Annotation Services

June 2026 · 11 min read · Updated September 2026

Short answer. Nine criteria decide whether an AI data annotation partner will hold up at foundation-model scale: a defined quality measurement (gold sets and inter-annotator agreement, not a claimed accuracy percentage), modality fit, workforce model and annotator retention, throughput and ramp behaviour, multilingual and low-resource coverage, domain expertise for specialist tasks, security and data residency, tooling and integration, and governance of taxonomy changes and disputes. Quality measurement is the criterion buyers most often under-specify and the one that predicts the most rework.

Key takeaways

  • A quality claim without a denominator, such as "99% accuracy", is not a measurement; require a gold-set protocol and a task-matched metric such as F1, IoU, word error rate or chance-corrected agreement.
  • The workforce model predicts quality stability more than the tooling does: a managed workforce in owned centres retains the learning curve that an open crowd pays for repeatedly.
  • Effective throughput is items delivered multiplied by first-pass acceptance rate; a vendor delivering 100,000 items a week at 70% acceptance is a 70,000-item vendor charging for 100,000.
  • Multilingual coverage should be measured as native-speaker annotator headcount per language, with location, not as a count of supported languages.
  • A paid pilot of several thousand items, including the hardest edge cases and at least one difficult language, reveals ramp behaviour and escalation quality that no proposal can.

Why is buying AI annotation services easy to get wrong?

Annotation buying is easy to get wrong because the deliverable looks the same whether it is good or not. A labelled dataset arrives on schedule, passes a spot-check, and the cost of the errors inside it surfaces months later as a model that underperforms in exactly the segment where the labels were weakest.

The nine criteria below are the ones that separate vendors after the demo. They are written to be used as an RFP structure, with a specific artefact to request under each, and they pair naturally with an annotation vendor consolidation and RFP guide when several suppliers are being compared at once. A shortlist of large-scale AI data annotation companies is a reasonable place to start; the criteria are how to narrow it.

1. Is the vendor's quality measurement actually defined?

"99% accuracy" is not a quality claim; it is a number with no denominator. Ask which measure it refers to, because the candidate measures test different things and differ by a lot on the same dataset.

Measure What it tells you When to require it
Gold-set accuracy Agreement with a trusted reference set Always; the baseline check
Inter-annotator agreement (IAA) Whether two competent people labelling the same item agree Any subjective or judgement-heavy task
F1 / precision / recall against gold Detection quality, split by error type Detection, segmentation, extraction
IoU thresholds Geometric tightness of boxes and masks Computer vision
Word error rate (WER) Transcription quality Speech
Kappa / Krippendorff's alpha Agreement corrected for chance Sentiment, ranking, moderation, RLHF

A gold set is a reference sample labelled by trusted experts, against which a vendor's routine output is scored. Inter-annotator agreement (IAA) is the degree to which two or more competent annotators, labelling the same item independently, produce the same label.

The correction for chance matters more than it sounds. On a binary task with an unbalanced class distribution, raw agreement can look excellent while the chance-corrected figure shows the annotators are barely distinguishing anything:

Cohen's kappa = (Observed agreement − Expected agreement) ÷ (1 − Expected agreement)
F1            = 2 × (Precision × Recall) ÷ (Precision + Recall)

Cohen's kappa was introduced by Jacob Cohen in 1960 for two raters on nominal categories; Krippendorff's alpha generalises the idea to any number of raters, missing data and several scale types. The arithmetic behind kappa and alpha, and what the numbers mean is worth understanding before a threshold goes into a contract.

What to request: the vendor's quality definition per task type, the gold-set protocol (who builds it, how often it refreshes, what share of work is gold-injected), and last quarter's figures for a comparable project. A vendor without a gold-set protocol is inspecting output rather than measuring it.

Also ask what happens below threshold. A defined rework policy, covering who pays, at what turnaround, and how the root cause is fed back into training, separates a supplier from a subcontractor. Gold sets, audit sampling and consensus are three distinct ways to QA annotated data, and a mature vendor can explain which of them it uses on which task.

2. Does the vendor fit your modality and task?

Annotation is not one skill, and a vendor strong in 2D bounding boxes may be weak in LiDAR while one strong in transcription may have no RLHF practice at all. Map your actual needs against the modalities the vendor has delivered at volume.

  • Computer vision: 2D/3D bounding boxes, semantic and instance segmentation, keypoints, tracking across frames, LiDAR point-cloud and sensor-fusion labelling.
  • Speech and audio: transcription, speaker diarisation, phonetic labelling, prosody, accent and dialect coverage.
  • Text and NLP: entity recognition, intent classification, sentiment, summarisation quality.
  • LLM and generative: RLHF preference ranking, SFT demonstration writing, red-teaming, response evaluation, data distillation.
  • Content moderation: policy application at scale, with the wellbeing considerations that come with it.

Ask for a reference project in your modality, at your volume. Adjacent experience is a much weaker signal here than in most categories.

3. What workforce model does the vendor run, and does it retain annotators?

The workforce model determines quality stability more than the tooling does. Three broad models exist, and most enterprise programmes need a managed workforce with specialist contractors layered in for review.

Model Strength Weakness
Open crowd Elastic, cheap, fast to start High churn, weak on domain tasks, variable IAA
Managed workforce in owned centres Trainable, retainable, auditable, secure Slower to scale into a brand-new skill
Specialist contractors Deep domain competence Expensive, limited throughput

The question that reveals the truth: "what is your annotator retention rate on a project of our length, and what happens to quality when a project team turns over?" Every complex taxonomy has a learning curve; a vendor with high churn pays that curve repeatedly, and you pay for it in rework.

4. How does throughput behave during ramp and at peak?

Steady-state throughput is the easy number; ramp time, peak behaviour and parallelism limits are the numbers that matter. Ask for effective throughput rather than delivered volume.

  • Ramp time to full quality at your volume, including the period where throughput exists but IAA has not stabilised.
  • Peak behaviour. What happens when you triple volume for six weeks? Quality tracks reviewer load with a lag of roughly one cycle.
  • Parallelism limits. How many distinct tasks can run at once without competing for the same trained pool?
Effective throughput = Items delivered × First-pass acceptance rate ÷ Cycle time

Effective throughput is the number of items delivered, multiplied by the first-pass acceptance rate, divided by cycle time. A vendor delivering 100,000 items a week at 70% acceptance is a 70,000-item vendor charging for 100,000, which is also why large-scale annotation cost should be compared on accepted items, not headline unit price.

5. Does the vendor cover your languages, including low-resource ones?

For foundation-model work, language coverage is frequently the binding constraint. Every vendor covers English, Mandarin, Spanish, French and German; programmes are decided in Thai, Vietnamese, Bahasa Indonesia, Tagalog, Bengali, Swahili, the Arabic dialects and the long tail beyond.

Measure coverage as native-speaker annotator headcount per language, with location, not as a supported-language count. For speech work in particular, ask about dialect and accent coverage inside a language: a "Vietnamese" capability that is entirely Hanoi-based is not general Vietnamese coverage, and the resulting model will show it. Managed low-resource speech data programmes exist precisely because this long tail cannot be served from a general-purpose roster.

Low-resource languages carry a second requirement: the vendor needs a sourcing method, not just a roster. Ask how they recruit and validate speakers in a language they do not currently cover, and how long it takes.

For foundation-model training data, coverage and consistency matter more than raw volume. A corpus that is broad across languages, dialects, demographics and edge cases, labelled consistently against versioned guidelines, outperforms a larger corpus assembled from whatever was easiest to source. Ask any prospective vendor how they measure coverage, not just how much they can deliver.

6. Does the vendor have domain expertise for specialist tasks?

Medical imaging, legal, financial, engineering and safety-critical driving scenarios need annotators who understand the content, not just the tool. Ask how domain reviewers are qualified, whether qualification is verified or self-declared, and what the escalation path is when an annotator is unsure.

The escalation path is the informative part. A programme with no defined route for "I don't know" produces confident wrong labels, which are more damaging than gaps because they are invisible in an acceptance check.

7. How does the vendor handle security, privacy and data residency?

Ask for evidence, not badges. Every security claim should be answerable with a document, a named facility or a named sub-processor.

  • Where is data stored and processed, and can work be confined to a named jurisdiction or a specific facility?
  • Which sub-processors touch it, and are they named?
  • Physical controls where the data warrants them: secure rooms, no personal devices, no removable media.
  • Access model: least privilege, revocation on rotation, audit logs.
  • Certifications: ask for the certificate and the scope statement, not the logo. An ISO/IEC 27001 certificate only covers the information security management system described in its scope, so a certification covering a corporate head office says nothing about the delivery centre doing your work.
  • PII handling and the deletion path at project end, with confirmation.

The security, privacy and compliance requirements for enterprise annotation run deeper than this checklist, but a vendor that cannot answer these six points cleanly will not answer the deeper ones either.

8. Whose tooling runs the work, and how does data move?

Two viable models exist: the vendor's platform, or your platform operated by their workforce. Both are fine; ambiguity is not.

Clarify who owns the annotation tool, whether your team gets read access to work in progress, and how the data actually moves: formats, schema versioning, API or bulk transfer, and how a mid-project taxonomy change propagates.

Ask whether the vendor can operate your tooling. A vendor who can only work inside their own platform creates a switching cost you are agreeing to at signature.

9. How does the vendor govern change, edge cases and disputes?

Governance is the criterion that separates a two-year partner from a one-project supplier. It covers how taxonomy changes are versioned, where ambiguous items go, who owns the guidelines and how disagreements about a batch are resolved.

  • Taxonomy change management. Real projects change definitions mid-flight. Ask how a change is versioned, whether previously labelled data is re-labelled or marked as a prior version, and who pays.
  • Edge-case handling. Where do ambiguous items go, who adjudicates, and how do decisions become guideline updates rather than tribal knowledge?
  • Guideline ownership. Written, versioned, and shared, or held in a project manager's head?
  • Dispute resolution. When you disagree about whether a batch meets spec, what is the process before it becomes a commercial argument?

How should the nine criteria be weighted?

Weight quality measurement highest, at 20%, then modality fit, workforce retention and multilingual coverage at 15% each, with throughput and security at 10% and the remaining three at 5%. The weights are a starting frame to adjust for your binding constraint, not a universal ranking.

Criterion Weight Artefact to request
Quality measurement 20% Quality definition per task, gold-set protocol, last-quarter figures
Modality and task fit 15% Reference project in your modality at your volume
Workforce model and retention 15% Retention rate; quality behaviour across team turnover
Throughput and ramp 10% Ramp curve; peak behaviour; effective throughput
Multilingual coverage 15% Annotator headcount per language, with location
Domain expertise 5% Qualification method; escalation path
Security and residency 10% Certificates with scope statements; residency options
Tooling and integration 5% Format and schema handling; ability to run your tooling
Governance 5% Guideline versioning; taxonomy change and dispute process

Run a paid pilot before committing volume, covering several thousand items including your hardest edge cases and at least one difficult language, and score the pilot on the same frame. A pilot reveals ramp behaviour and escalation quality, which no proposal can.

How does Lifewood approach these criteria?

Lifewood delivers annotation through a managed workforce in owned delivery centres rather than an open crowd, which is the model that suits long-running programmes with complex taxonomies because the learning curve is paid once and retained.

Scope spans the modalities above: LLM work including RLHF, SFT and response evaluation; computer vision including 2D/3D boxes, segmentation and keypoints for autonomous driving and medical imaging; speech and NLP including multilingual transcription and phonetic labelling with a specialism in low-resource languages and regional dialects; conversational-AI training data; content moderation; and bespoke field data collection. The full range is set out on the AI data services page, and delivery runs against a 95%+ accuracy SLA measured with two independent review passes.

The multilingual position is the structural one: 100+ languages, 40+ delivery centres across 30+ countries, and a global pool of 56,000+ registered contributors, with region-native annotators rather than remote approximations. The AI-data heritage runs to 2004; engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement.

Frequently asked questions

The category has three tiers. Large established providers such as Appen and Scale AI are the reference points for general-purpose volume. Specialist providers focus on one modality, such as LiDAR, medical imaging or speech, and win on depth. Managed multilingual providers such as Lifewood combine owned delivery centres with broad language coverage. Shortlist by binding constraint, then pilot.

Ask for a reference project in your exact modality, whether 2D boxes, segmentation, keypoints, video tracking or LiDAR, at your volume. Require IoU thresholds and F1 against a gold set rather than a headline accuracy figure, confirm the vendor can operate your tooling and export your schema, and pilot with your hardest edge cases before committing volume.

Against a gold set, with the measure matched to the task: F1 and IoU for detection and segmentation, word error rate for transcription, and chance-corrected agreement, meaning Cohen's kappa or Krippendorff's alpha, for anything involving judgement. A single accuracy percentage with no denominator, no gold-set protocol and no error-type breakdown is not a measurement.

Inter-annotator agreement is the degree to which two competent annotators labelling the same item produce the same answer. Low agreement usually means the guidelines are ambiguous rather than the annotators are poor, which makes it the earliest available signal that a taxonomy needs fixing, before a whole batch is labelled inconsistently.

It varies by modality, complexity, language and quality requirement. Simple image bounding boxes sit at the low end; LiDAR 3D annotation, RLHF ranking and multilingual speech transcription scale substantially higher. Compare on effective cost, meaning price divided by first-pass acceptance rate, rather than headline unit price, since rework is paid in schedule as well as money.

Crowd models are elastic and cheap and suit simple, high-volume, low-ambiguity tasks. Managed workforces in owned centres suit complex taxonomies, long programmes, sensitive data and specialist domains, because the training investment is retained. Many programmes use both, with the managed workforce handling the parts where retention and security matter most.

Sources and further reading

  1. Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement
  2. Krippendorff, K. (2011). Computing Krippendorff's Alpha-Reliability. University of Pennsylvania ScholarlyCommons
  3. ISO/IEC 27001:2022 Information security management systems
  4. Appen: expert-validated data for frontier models
  5. Scale AI: data for frontier AI
  6. Lifewood AI data services — service scope, 50+ languages, 40+ delivery centres, 95%+ accuracy SLA
  7. Lifewood Data Technology — company-reported delivery figures including 56,000+ registered contributors

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team