Short answer. Buy image annotation on objects, not images. Four questions separate providers: which geometries they support under one consistent guideline (boxes, polygons, segmentation, keypoints, attributes); whether they scope by object density rather than file count; what happens to occluded, blurred and ambiguous objects; and what share of work gets a second-pass review. A vendor who quotes per image without asking about your average objects per image has priced an assumption, and that assumption is the risk.
Key takeaways
- Image annotation should be scoped and priced per object, with attribute fields priced separately, because object density varies far more than image count.
- Large-scale image annotation has six task families: classification, bounding boxes, polygons and segmentation, keypoints, per-object attributes, and text-in-scene transcription.
- Edge-case rules for occlusion, truncation, minimum size, ambiguous class pairs, group objects and unusable images must be written into the guideline before the first batch.
- Quality should be measured per object class using IoU at a stated threshold, F1 by class, chance-corrected agreement for subjective attributes, and a known review rate.
- Lifewood Data Technology delivers image annotation through a managed workforce with a 95%+ accuracy SLA, 50+ languages and 40+ delivery centres across 30+ countries.
What does large-scale image annotation actually include?
Large-scale image annotation covers six task families with different labour profiles and failure modes, and a capable provider runs all six under one guideline document and one taxonomy.
Image annotation is the process of adding structured labels, shapes or attributes to images so that a computer vision model can learn what each region contains. Image annotation is the most commoditised-looking service in AI data and one of the easiest to buy badly. The deliverable looks identical whether it is good or not: a labelled dataset arrives on schedule, passes a spot check, and the cost of the errors inside it appears months later as a model that underperforms in exactly the segment where labels were weakest.
The six task families are:
- Classification — one or more labels per image. Cheapest per unit; fails on taxonomy ambiguity rather than on execution.
- 2D bounding boxes — the volume workhorse. Fails on occlusion rules and on small distant objects.
- Polygons and segmentation — semantic and instance. Several times the labour per object; fails on boundary consistency between annotators.
- Keypoints and pose — landmark placement with visibility rules. Fails when the visibility convention is under-specified.
- Attributes — per-object properties: colour, state, truncation, occlusion level, orientation. Routinely under-costed, because six attribute fields per object is a materially larger job than a bare box.
- OCR-style labelling and text in scene — transcription of text within images, including non-Latin and right-to-left scripts, where segmentation behaves differently.
A provider who can run some of these well and subcontracts the rest introduces a second interpretation of your ontology without telling you. The same taxonomy carries into the companion guide on buying large-scale video annotation.
What should buyers compare between image annotation providers?
Buyers should compare providers on seven criteria: annotation geometry, object density scoping, edge-case handling, quality control, scale behaviour, security, and native readers for text in images. Everything else, including tooling, turnaround and headline unit rate, is downstream.
| Buyer criterion | Why it matters | What strong delivery looks like |
|---|---|---|
| Annotation geometry | Geometries differ in labour and QA by several times | Boxes, polygons, segmentation and keypoints under consistent guidelines |
| Object density | Dense images multiply annotation and review effort | Scoping by objects and complexity, not image count |
| Edge cases | Occlusion, blur and ambiguity damage model quality quietly | Written escalation and adjudication process |
| Quality control | A systematic error repeats across millions of objects | Multi-stage review with measurable acceptance targets |
| Scale behaviour | Enterprise projects contain millions of objects | Trained workforce that ramps without quality collapse |
| Security | Images contain faces, documents, proprietary environments | Access controls and controlled production workflows |
| Text in image | Non-Latin and RTL scripts need native readers | Native-speaker annotators for the scripts in scope |
For a checklist that applies across modalities, see the 9 criteria for choosing AI annotation services; for named providers side by side on computer vision work, see the comparison of human-in-the-loop computer vision annotation providers.
How do you specify object density before you get quotes?
Measure object density on a random sample of at least 200 real production images, recording the median, 90th percentile and maximum objects per image, before asking any vendor for a price.
Object density is the number of labelled objects an image contains under a given ontology, and it is the variable that determines annotation effort far more than image count. Measuring it is the single highest-value hour in an image annotation procurement.
- Take a random sample of at least 200 images from real production data — not a curated demo set.
- Count objects per image under your own ontology, including the ones you are unsure about.
- Record the median, the 90th percentile and the maximum.
- Count attribute fields per object.
Now you can ask for per-object pricing with a known multiplier, and you can spot the vendor whose quote assumed a density that your data does not have. CVAT's published cost analysis for a solo annotation project assumes an average of 23 objects per image across 100,000 images — 2.3 million billable objects — which is a useful reminder that the file count is the least informative number in the brief. Per-object rate ranges by geometry are in the image, video and 3D/LiDAR annotation pricing guide.
Which edge-case rules should be defined before the first batch?
Seven rules belong in the guideline before the first batch: occlusion threshold, truncation, minimum size, ambiguous class pairs, group objects, image quality floor, and a route for objects an annotator cannot classify.
Most image annotation disputes are not about competence. They are about rules that were never written. Settle these in the guideline document:
- Occlusion threshold. At what visible fraction does an object stop being labelled? State a number.
- Truncation. Objects cut by the frame edge — labelled, labelled with a flag, or skipped?
- Minimum size. Below what pixel dimension is an object out of scope? Without this, annotator patience sets your threshold.
- Ambiguous class pairs. The two classes your own team argues about. Name them and give an adjudication rule.
- Group objects. A crowd, a shelf of products, a pile — individually or as a region?
- Image quality floor. When is an image rejected as unusable rather than labelled badly?
- Unknown. Where does an annotator send something they genuinely cannot classify? A programme with no route for "I don't know" produces confident wrong labels, which are worse than gaps because they are invisible in an acceptance check.
Writing these rules so that a new annotator applies them the same way as an experienced one is its own skill; see the guide on writing annotation guidelines that annotators actually follow.
How do you measure image annotation quality?
Image annotation quality should be measured per object class using IoU at a stated threshold, F1 by class, chance-corrected agreement for subjective attributes, and a known second-pass review rate.
Do not accept a single accuracy percentage. Intersection over Union (IoU) is the area where a predicted box or polygon overlaps the reference divided by the total area the two shapes cover together, so a value of 1.0 means a perfect match and 0.5 means half the combined area overlaps. Require, per object class:
IoU = Area of overlap ÷ Area of union (geometric tightness)
F1 = 2 × (Precision × Recall) ÷ (P + R) (detection quality)
- IoU at threshold for boxes and polygons — 0.5 is the classic PASCAL VOC bar for a correct detection, and it is a weak bar for anything safety-relevant; state the threshold you need per class.
- F1 by class, because an aggregate figure is dominated by large, easy, well-lit objects.
- Chance-corrected agreement (Cohen's kappa or Krippendorff's alpha) for any subjective attribute.
- Attribute accuracy separately from geometry accuracy. They fail independently.
Quality matters disproportionately at scale because a small systematic error repeated across millions of objects distorts model behaviour in one consistent direction, which is much harder to detect than random noise.
Second-pass review is an independent check of a completed annotation by a different annotator before the work is accepted. Require the review rate: what percentage of work receives second-pass review, and how is the sample chosen? Random sampling and risk-weighted sampling produce very different assurance for the same cost. Lifewood's AI data validation service measures review passes against a customer-approved gold set.
What questions should you ask before purchasing?
Eight questions expose a thin process before you commit volume, and the most informative asks about the provider's worst quality incident in the last year.
- Which annotation geometries will you need now, and which within eighteen months?
- What is the median and 90th-percentile number of objects per image in our data?
- How are occlusion, truncation, minimum size and group objects defined in your guidelines?
- What percentage of work is second-pass reviewed, and how is that sample selected?
- Can the same schema extend into video or 3D without re-labelling the image set?
- How are guideline changes rolled out across large teams, and who pays for re-labelling?
- For images containing text in non-Latin scripts, who reads them?
- What was your worst quality incident on an image programme in the last year, and what changed afterwards?
The last question is the most informative. A provider operating at real volume has had an incident; one who claims otherwise is new, small, or not measuring.
How does Lifewood approach large-scale image annotation?
Lifewood delivers image annotation through a managed workforce in owned delivery centres rather than an open crowd, covering classification, boxes, polygons, segmentation, keypoints and attribute labelling under a contractual 95%+ accuracy SLA.
The managed-workforce model suits long-running programmes with complex taxonomies: the learning curve on your ontology is paid once and retained. Image work sits alongside video, text, audio and 3D point-cloud work in the same programme, so a schema can extend across modalities without a second vendor and a second interpretation (full scope on the AI data services page).
The quality framework is contractual: a 95%+ accuracy SLA with trained annotators, senior second-pass review, automated consistency checks and client feedback loops, and below-threshold batches reworked at Lifewood's cost.
For images carrying text, signage or culturally specific content, 100+ languages with region-native annotators in 40+ delivery centres across 30+ countries matters more than buyers expect — non-Latin and right-to-left script work needs someone who reads the script, not someone transcribing shapes.
Other credible providers for image-heavy programmes include Sama, which positions its services around human-in-the-loop image, video and 3D annotation; TELUS Digital, for buyers who want annotation combined with Ground Truth Studio and configurable workflows; and Scale AI, for highly technical computer-vision pipelines built around its Data Engine. Choose a narrower specialist if a pilot proves materially better quality or economics on your exact task. A wider ranked field is in the list of top large-scale AI data annotation and labelling companies.