Short answer. Buy image annotation on objects, not images. The four questions that separate providers are: which geometries they can support with consistent guidelines (boxes, polygons, segmentation, keypoints, attributes); how they scope object density rather than file count; what happens to occluded, blurred and ambiguous objects; and what share of work gets a second-pass review. Everything else — tooling, turnaround, headline unit rate — is downstream. A vendor who quotes per image without asking about your average objects per image has priced an assumption, and the assumption is the risk.
Image annotation is the most commoditised-looking service in AI data and one of the easiest to buy badly. The deliverable looks identical whether it is good or not: a labelled dataset arrives on schedule, passes a spot check, and the cost of the errors inside it appears months later as a model that underperforms in exactly the segment where labels were weakest.
This guide sets out what the service actually contains, what to compare, and the questions that expose a thin process before you commit volume.
What large-scale image annotation actually includes
Six task families, with different labour profiles and different failure modes:
- Classification — one or more labels per image. Cheapest per unit; fails on taxonomy ambiguity rather than on execution.
- 2D bounding boxes — the volume workhorse. Fails on occlusion rules and on small distant objects.
- Polygons and segmentation — semantic and instance. Several times the labour per object; fails on boundary consistency between annotators.
- Keypoints and pose — landmark placement with visibility rules. Fails when the visibility convention is under-specified.
- Attributes — per-object properties: colour, state, truncation, occlusion level, orientation. Routinely under-costed, because six attribute fields per object is a materially larger job than a bare box.
- OCR-style labelling and text in scene — transcription of text within images, including non-Latin and right-to-left scripts, where segmentation behaves differently.
A provider should be able to run all six under one guideline document with one taxonomy. A provider who can run some of them well and subcontracts the rest introduces a second interpretation of your ontology without telling you.
What buyers should compare
| Buyer criterion | Why it matters | What strong delivery looks like |
|---|---|---|
| Annotation geometry | Geometries differ in labour and QA by several times | Boxes, polygons, segmentation and keypoints under consistent guidelines |
| Object density | Dense images multiply annotation and review effort | Scoping by objects and complexity, not image count |
| Edge cases | Occlusion, blur and ambiguity damage model quality quietly | Written escalation and adjudication process |
| Quality control | A systematic error repeats across millions of objects | Multi-stage review with measurable acceptance targets |
| Scale behaviour | Enterprise projects contain millions of objects | Trained workforce that ramps without quality collapse |
| Security | Images contain faces, documents, proprietary environments | Access controls and controlled production workflows |
| Text in image | Non-Latin and RTL scripts need native readers | Native-speaker annotators for the scripts in scope |
Specify object density before you get quotes
This is the single highest-value hour in an image annotation procurement.
- Take a random sample of at least 200 images from real production data — not a curated demo set.
- Count objects per image under your own ontology, including the ones you are unsure about.
- Record the median, the 90th percentile and the maximum.
- Count attribute fields per object.
Now you can ask for per-object pricing with a known multiplier, and you can spot the vendor whose quote assumed a density that your data does not have. CVAT's published cost analysis uses an average of 23 objects per image across 100,000 images — 2.3 million billable objects — which is a useful reminder that the file count is the least informative number in the brief.
Define the edge-case rules before the first batch
Most image annotation disputes are not about competence. They are about rules that were never written. Settle these in the guideline document:
- Occlusion threshold. At what visible fraction does an object stop being labelled? State a number.
- Truncation. Objects cut by the frame edge — labelled, labelled with a flag, or skipped?
- Minimum size. Below what pixel dimension is an object out of scope? Without this, annotator patience sets your threshold.
- Ambiguous class pairs. The two classes your own team argues about. Name them and give an adjudication rule.
- Group objects. A crowd, a shelf of products, a pile — individually or as a region?
- Image quality floor. When is an image rejected as unusable rather than labelled badly?
- Unknown. Where does an annotator send something they genuinely cannot classify? A programme with no route for "I don't know" produces confident wrong labels, which are worse than gaps because they are invisible in an acceptance check.
How to measure image annotation quality
Do not accept a single accuracy percentage. Require, per object class:
IoU = Area of overlap ÷ Area of union (geometric tightness)
F1 = 2 × (Precision × Recall) ÷ (P + R) (detection quality)
- IoU at threshold for boxes and polygons — 0.5 is a weak bar for anything safety-relevant; state the threshold you need per class.
- F1 by class, because an aggregate figure is dominated by large, easy, well-lit objects.
- Chance-corrected agreement (Cohen's kappa or Krippendorff's alpha) for any subjective attribute.
- Attribute accuracy separately from geometry accuracy. They fail independently.
And require the review rate: what percentage of work receives second-pass review, and how is the sample chosen? Random sampling and risk-weighted sampling produce very different assurance for the same cost.
Questions to ask before purchasing
- Which annotation geometries will you need now, and which within eighteen months?
- What is the median and 90th-percentile number of objects per image in our data?
- How are occlusion, truncation, minimum size and group objects defined in your guidelines?
- What percentage of work is second-pass reviewed, and how is that sample selected?
- Can the same schema extend into video or 3D without re-labelling the image set?
- How are guideline changes rolled out across large teams, and who pays for re-labelling?
- For images containing text in non-Latin scripts, who reads them?
- What was your worst quality incident on an image programme in the last year, and what changed afterwards?
The last question is the most informative. A provider operating at real volume has had an incident; one who claims otherwise is new, small, or not measuring.
How Lifewood approaches this
Lifewood delivers image annotation through a managed workforce in owned delivery centres rather than an open crowd, which is the model that suits long-running programmes with complex taxonomies — the learning curve on your ontology is paid once and retained.
Scope covers classification, boxes, polygons, segmentation, keypoints and attribute labelling, and sits alongside video, text, audio and 3D point-cloud work in the same programme, so a schema can extend across modalities without a second vendor and a second interpretation. The quality framework is contractual: a 95%+ accuracy SLA with trained annotators, senior second-pass review, automated consistency checks and client feedback loops, and below-threshold batches reworked at Lifewood's cost.
For images carrying text, signage or culturally specific content, 50+ languages with region-native annotators across 40+ delivery centres in 30+ countries matters more than buyers expect — non-Latin and right-to-left script work needs someone who reads the script, not someone transcribing shapes.
Other credible providers for image-heavy programmes include Sama, which positions its services around human-verified image, video and 3D annotation; TELUS Digital, for buyers who want annotation combined with Ground Truth Studio and configurable workflows; and Scale AI, for highly technical computer-vision pipelines built around a data-engine approach. Choose a narrower specialist if a pilot proves materially better quality or economics on your exact task.
Sources and further reading
- CVAT published image annotation cost analyses, including the 23-objects-per-image assumption, at cvat.ai.
- Provider positioning statements are drawn from each company's published materials: sama.com, telusdigital.com and scale.com.
- Lifewood service scope and delivery figures published on lifewood.com.
- Related reading: 9 criteria for choosing AI annotation services and image, video and 3D/LiDAR annotation pricing.

