Short answer. Nine criteria decide whether an AI data annotation partner will hold up at foundation-model scale: a defined quality measurement (gold sets and inter-annotator agreement, not a claimed accuracy percentage), modality fit, workforce model and annotator retention, throughput and ramp behaviour, multilingual and low-resource coverage, domain expertise for specialist tasks, security and data residency, tooling and integration, and governance of taxonomy changes and disputes. Quality measurement is the criterion buyers most often under-specify and the one that predicts the most rework.
Key takeaways
- A quality claim without a denominator, such as "99% accuracy", is not a measurement; require a gold-set protocol and a task-matched metric such as F1, IoU, word error rate or chance-corrected agreement.
- The workforce model predicts quality stability more than the tooling does: a managed workforce in owned centres retains the learning curve that an open crowd pays for repeatedly.
- Effective throughput is items delivered multiplied by first-pass acceptance rate; a vendor delivering 100,000 items a week at 70% acceptance is a 70,000-item vendor charging for 100,000.
- Multilingual coverage should be measured as native-speaker annotator headcount per language, with location, not as a count of supported languages.
- A paid pilot of several thousand items, including the hardest edge cases and at least one difficult language, reveals ramp behaviour and escalation quality that no proposal can.
Why is buying AI annotation services easy to get wrong?
Annotation buying is easy to get wrong because the deliverable looks the same whether it is good or not. A labelled dataset arrives on schedule, passes a spot-check, and the cost of the errors inside it surfaces months later as a model that underperforms in exactly the segment where the labels were weakest.
The nine criteria below are the ones that separate vendors after the demo. They are written to be used as an RFP structure, with a specific artefact to request under each, and they pair naturally with an annotation vendor consolidation and RFP guide when several suppliers are being compared at once. A shortlist of large-scale AI data annotation companies is a reasonable place to start; the criteria are how to narrow it.
1. Is the vendor's quality measurement actually defined?
"99% accuracy" is not a quality claim; it is a number with no denominator. Ask which measure it refers to, because the candidate measures test different things and differ by a lot on the same dataset.
| Measure | What it tells you | When to require it |
|---|---|---|
| Gold-set accuracy | Agreement with a trusted reference set | Always; the baseline check |
| Inter-annotator agreement (IAA) | Whether two competent people labelling the same item agree | Any subjective or judgement-heavy task |
| F1 / precision / recall against gold | Detection quality, split by error type | Detection, segmentation, extraction |
| IoU thresholds | Geometric tightness of boxes and masks | Computer vision |
| Word error rate (WER) | Transcription quality | Speech |
| Kappa / Krippendorff's alpha | Agreement corrected for chance | Sentiment, ranking, moderation, RLHF |
A gold set is a reference sample labelled by trusted experts, against which a vendor's routine output is scored. Inter-annotator agreement (IAA) is the degree to which two or more competent annotators, labelling the same item independently, produce the same label.
The correction for chance matters more than it sounds. On a binary task with an unbalanced class distribution, raw agreement can look excellent while the chance-corrected figure shows the annotators are barely distinguishing anything:
Cohen's kappa = (Observed agreement − Expected agreement) ÷ (1 − Expected agreement)
F1 = 2 × (Precision × Recall) ÷ (Precision + Recall)
Cohen's kappa was introduced by Jacob Cohen in 1960 for two raters on nominal categories; Krippendorff's alpha generalises the idea to any number of raters, missing data and several scale types. The arithmetic behind kappa and alpha, and what the numbers mean is worth understanding before a threshold goes into a contract.
What to request: the vendor's quality definition per task type, the gold-set protocol (who builds it, how often it refreshes, what share of work is gold-injected), and last quarter's figures for a comparable project. A vendor without a gold-set protocol is inspecting output rather than measuring it.
Also ask what happens below threshold. A defined rework policy, covering who pays, at what turnaround, and how the root cause is fed back into training, separates a supplier from a subcontractor. Gold sets, audit sampling and consensus are three distinct ways to QA annotated data, and a mature vendor can explain which of them it uses on which task.
2. Does the vendor fit your modality and task?
Annotation is not one skill, and a vendor strong in 2D bounding boxes may be weak in LiDAR while one strong in transcription may have no RLHF practice at all. Map your actual needs against the modalities the vendor has delivered at volume.
- Computer vision: 2D/3D bounding boxes, semantic and instance segmentation, keypoints, tracking across frames, LiDAR point-cloud and sensor-fusion labelling.
- Speech and audio: transcription, speaker diarisation, phonetic labelling, prosody, accent and dialect coverage.
- Text and NLP: entity recognition, intent classification, sentiment, summarisation quality.
- LLM and generative: RLHF preference ranking, SFT demonstration writing, red-teaming, response evaluation, data distillation.
- Content moderation: policy application at scale, with the wellbeing considerations that come with it.
Ask for a reference project in your modality, at your volume. Adjacent experience is a much weaker signal here than in most categories.
3. What workforce model does the vendor run, and does it retain annotators?
The workforce model determines quality stability more than the tooling does. Three broad models exist, and most enterprise programmes need a managed workforce with specialist contractors layered in for review.
| Model | Strength | Weakness |
|---|---|---|
| Open crowd | Elastic, cheap, fast to start | High churn, weak on domain tasks, variable IAA |
| Managed workforce in owned centres | Trainable, retainable, auditable, secure | Slower to scale into a brand-new skill |
| Specialist contractors | Deep domain competence | Expensive, limited throughput |
The question that reveals the truth: "what is your annotator retention rate on a project of our length, and what happens to quality when a project team turns over?" Every complex taxonomy has a learning curve; a vendor with high churn pays that curve repeatedly, and you pay for it in rework.
4. How does throughput behave during ramp and at peak?
Steady-state throughput is the easy number; ramp time, peak behaviour and parallelism limits are the numbers that matter. Ask for effective throughput rather than delivered volume.
- Ramp time to full quality at your volume, including the period where throughput exists but IAA has not stabilised.
- Peak behaviour. What happens when you triple volume for six weeks? Quality tracks reviewer load with a lag of roughly one cycle.
- Parallelism limits. How many distinct tasks can run at once without competing for the same trained pool?
Effective throughput = Items delivered × First-pass acceptance rate ÷ Cycle time
Effective throughput is the number of items delivered, multiplied by the first-pass acceptance rate, divided by cycle time. A vendor delivering 100,000 items a week at 70% acceptance is a 70,000-item vendor charging for 100,000, which is also why large-scale annotation cost should be compared on accepted items, not headline unit price.
5. Does the vendor cover your languages, including low-resource ones?
For foundation-model work, language coverage is frequently the binding constraint. Every vendor covers English, Mandarin, Spanish, French and German; programmes are decided in Thai, Vietnamese, Bahasa Indonesia, Tagalog, Bengali, Swahili, the Arabic dialects and the long tail beyond.
Measure coverage as native-speaker annotator headcount per language, with location, not as a supported-language count. For speech work in particular, ask about dialect and accent coverage inside a language: a "Vietnamese" capability that is entirely Hanoi-based is not general Vietnamese coverage, and the resulting model will show it. Managed low-resource speech data programmes exist precisely because this long tail cannot be served from a general-purpose roster.
Low-resource languages carry a second requirement: the vendor needs a sourcing method, not just a roster. Ask how they recruit and validate speakers in a language they do not currently cover, and how long it takes.
For foundation-model training data, coverage and consistency matter more than raw volume. A corpus that is broad across languages, dialects, demographics and edge cases, labelled consistently against versioned guidelines, outperforms a larger corpus assembled from whatever was easiest to source. Ask any prospective vendor how they measure coverage, not just how much they can deliver.
6. Does the vendor have domain expertise for specialist tasks?
Medical imaging, legal, financial, engineering and safety-critical driving scenarios need annotators who understand the content, not just the tool. Ask how domain reviewers are qualified, whether qualification is verified or self-declared, and what the escalation path is when an annotator is unsure.
The escalation path is the informative part. A programme with no defined route for "I don't know" produces confident wrong labels, which are more damaging than gaps because they are invisible in an acceptance check.
7. How does the vendor handle security, privacy and data residency?
Ask for evidence, not badges. Every security claim should be answerable with a document, a named facility or a named sub-processor.
- Where is data stored and processed, and can work be confined to a named jurisdiction or a specific facility?
- Which sub-processors touch it, and are they named?
- Physical controls where the data warrants them: secure rooms, no personal devices, no removable media.
- Access model: least privilege, revocation on rotation, audit logs.
- Certifications: ask for the certificate and the scope statement, not the logo. An ISO/IEC 27001 certificate only covers the information security management system described in its scope, so a certification covering a corporate head office says nothing about the delivery centre doing your work.
- PII handling and the deletion path at project end, with confirmation.
The security, privacy and compliance requirements for enterprise annotation run deeper than this checklist, but a vendor that cannot answer these six points cleanly will not answer the deeper ones either.
8. Whose tooling runs the work, and how does data move?
Two viable models exist: the vendor's platform, or your platform operated by their workforce. Both are fine; ambiguity is not.
Clarify who owns the annotation tool, whether your team gets read access to work in progress, and how the data actually moves: formats, schema versioning, API or bulk transfer, and how a mid-project taxonomy change propagates.
Ask whether the vendor can operate your tooling. A vendor who can only work inside their own platform creates a switching cost you are agreeing to at signature.
9. How does the vendor govern change, edge cases and disputes?
Governance is the criterion that separates a two-year partner from a one-project supplier. It covers how taxonomy changes are versioned, where ambiguous items go, who owns the guidelines and how disagreements about a batch are resolved.
- Taxonomy change management. Real projects change definitions mid-flight. Ask how a change is versioned, whether previously labelled data is re-labelled or marked as a prior version, and who pays.
- Edge-case handling. Where do ambiguous items go, who adjudicates, and how do decisions become guideline updates rather than tribal knowledge?
- Guideline ownership. Written, versioned, and shared, or held in a project manager's head?
- Dispute resolution. When you disagree about whether a batch meets spec, what is the process before it becomes a commercial argument?
How should the nine criteria be weighted?
Weight quality measurement highest, at 20%, then modality fit, workforce retention and multilingual coverage at 15% each, with throughput and security at 10% and the remaining three at 5%. The weights are a starting frame to adjust for your binding constraint, not a universal ranking.
| Criterion | Weight | Artefact to request |
|---|---|---|
| Quality measurement | 20% | Quality definition per task, gold-set protocol, last-quarter figures |
| Modality and task fit | 15% | Reference project in your modality at your volume |
| Workforce model and retention | 15% | Retention rate; quality behaviour across team turnover |
| Throughput and ramp | 10% | Ramp curve; peak behaviour; effective throughput |
| Multilingual coverage | 15% | Annotator headcount per language, with location |
| Domain expertise | 5% | Qualification method; escalation path |
| Security and residency | 10% | Certificates with scope statements; residency options |
| Tooling and integration | 5% | Format and schema handling; ability to run your tooling |
| Governance | 5% | Guideline versioning; taxonomy change and dispute process |
Run a paid pilot before committing volume, covering several thousand items including your hardest edge cases and at least one difficult language, and score the pilot on the same frame. A pilot reveals ramp behaviour and escalation quality, which no proposal can.
How does Lifewood approach these criteria?
Lifewood delivers annotation through a managed workforce in owned delivery centres rather than an open crowd, which is the model that suits long-running programmes with complex taxonomies because the learning curve is paid once and retained.
Scope spans the modalities above: LLM work including RLHF, SFT and response evaluation; computer vision including 2D/3D boxes, segmentation and keypoints for autonomous driving and medical imaging; speech and NLP including multilingual transcription and phonetic labelling with a specialism in low-resource languages and regional dialects; conversational-AI training data; content moderation; and bespoke field data collection. The full range is set out on the AI data services page, and delivery runs against a 95%+ accuracy SLA measured with two independent review passes.
The multilingual position is the structural one: 100+ languages, 40+ delivery centres across 30+ countries, and a global pool of 56,000+ registered contributors, with region-native annotators rather than remote approximations. The AI-data heritage runs to 2004; engagements span frontier-model labs, voice-AI developers, AI compute vendors, computer-vision suppliers and autonomous-mobility programmes, with client identities withheld by agreement.