Skip to main content
AI Data

How to Buy Large-Scale Image Annotation

June 2026 · 9 min read · Updated September 2026

Short answer. Buy image annotation on objects, not images. Four questions separate providers: which geometries they support under one consistent guideline (boxes, polygons, segmentation, keypoints, attributes); whether they scope by object density rather than file count; what happens to occluded, blurred and ambiguous objects; and what share of work gets a second-pass review. A vendor who quotes per image without asking about your average objects per image has priced an assumption, and that assumption is the risk.

Key takeaways

  • Image annotation should be scoped and priced per object, with attribute fields priced separately, because object density varies far more than image count.
  • Large-scale image annotation has six task families: classification, bounding boxes, polygons and segmentation, keypoints, per-object attributes, and text-in-scene transcription.
  • Edge-case rules for occlusion, truncation, minimum size, ambiguous class pairs, group objects and unusable images must be written into the guideline before the first batch.
  • Quality should be measured per object class using IoU at a stated threshold, F1 by class, chance-corrected agreement for subjective attributes, and a known review rate.
  • Lifewood Data Technology delivers image annotation through a managed workforce with a 95%+ accuracy SLA, 50+ languages and 40+ delivery centres across 30+ countries.

What does large-scale image annotation actually include?

Large-scale image annotation covers six task families with different labour profiles and failure modes, and a capable provider runs all six under one guideline document and one taxonomy.

Image annotation is the process of adding structured labels, shapes or attributes to images so that a computer vision model can learn what each region contains. Image annotation is the most commoditised-looking service in AI data and one of the easiest to buy badly. The deliverable looks identical whether it is good or not: a labelled dataset arrives on schedule, passes a spot check, and the cost of the errors inside it appears months later as a model that underperforms in exactly the segment where labels were weakest.

The six task families are:

  • Classification — one or more labels per image. Cheapest per unit; fails on taxonomy ambiguity rather than on execution.
  • 2D bounding boxes — the volume workhorse. Fails on occlusion rules and on small distant objects.
  • Polygons and segmentation — semantic and instance. Several times the labour per object; fails on boundary consistency between annotators.
  • Keypoints and pose — landmark placement with visibility rules. Fails when the visibility convention is under-specified.
  • Attributes — per-object properties: colour, state, truncation, occlusion level, orientation. Routinely under-costed, because six attribute fields per object is a materially larger job than a bare box.
  • OCR-style labelling and text in scene — transcription of text within images, including non-Latin and right-to-left scripts, where segmentation behaves differently.

A provider who can run some of these well and subcontracts the rest introduces a second interpretation of your ontology without telling you. The same taxonomy carries into the companion guide on buying large-scale video annotation.

What should buyers compare between image annotation providers?

Buyers should compare providers on seven criteria: annotation geometry, object density scoping, edge-case handling, quality control, scale behaviour, security, and native readers for text in images. Everything else, including tooling, turnaround and headline unit rate, is downstream.

Buyer criterion Why it matters What strong delivery looks like
Annotation geometry Geometries differ in labour and QA by several times Boxes, polygons, segmentation and keypoints under consistent guidelines
Object density Dense images multiply annotation and review effort Scoping by objects and complexity, not image count
Edge cases Occlusion, blur and ambiguity damage model quality quietly Written escalation and adjudication process
Quality control A systematic error repeats across millions of objects Multi-stage review with measurable acceptance targets
Scale behaviour Enterprise projects contain millions of objects Trained workforce that ramps without quality collapse
Security Images contain faces, documents, proprietary environments Access controls and controlled production workflows
Text in image Non-Latin and RTL scripts need native readers Native-speaker annotators for the scripts in scope

For a checklist that applies across modalities, see the 9 criteria for choosing AI annotation services; for named providers side by side on computer vision work, see the comparison of human-in-the-loop computer vision annotation providers.

How do you specify object density before you get quotes?

Measure object density on a random sample of at least 200 real production images, recording the median, 90th percentile and maximum objects per image, before asking any vendor for a price.

Object density is the number of labelled objects an image contains under a given ontology, and it is the variable that determines annotation effort far more than image count. Measuring it is the single highest-value hour in an image annotation procurement.

  1. Take a random sample of at least 200 images from real production data — not a curated demo set.
  2. Count objects per image under your own ontology, including the ones you are unsure about.
  3. Record the median, the 90th percentile and the maximum.
  4. Count attribute fields per object.

Now you can ask for per-object pricing with a known multiplier, and you can spot the vendor whose quote assumed a density that your data does not have. CVAT's published cost analysis for a solo annotation project assumes an average of 23 objects per image across 100,000 images — 2.3 million billable objects — which is a useful reminder that the file count is the least informative number in the brief. Per-object rate ranges by geometry are in the image, video and 3D/LiDAR annotation pricing guide.

Which edge-case rules should be defined before the first batch?

Seven rules belong in the guideline before the first batch: occlusion threshold, truncation, minimum size, ambiguous class pairs, group objects, image quality floor, and a route for objects an annotator cannot classify.

Most image annotation disputes are not about competence. They are about rules that were never written. Settle these in the guideline document:

  • Occlusion threshold. At what visible fraction does an object stop being labelled? State a number.
  • Truncation. Objects cut by the frame edge — labelled, labelled with a flag, or skipped?
  • Minimum size. Below what pixel dimension is an object out of scope? Without this, annotator patience sets your threshold.
  • Ambiguous class pairs. The two classes your own team argues about. Name them and give an adjudication rule.
  • Group objects. A crowd, a shelf of products, a pile — individually or as a region?
  • Image quality floor. When is an image rejected as unusable rather than labelled badly?
  • Unknown. Where does an annotator send something they genuinely cannot classify? A programme with no route for "I don't know" produces confident wrong labels, which are worse than gaps because they are invisible in an acceptance check.

Writing these rules so that a new annotator applies them the same way as an experienced one is its own skill; see the guide on writing annotation guidelines that annotators actually follow.

How do you measure image annotation quality?

Image annotation quality should be measured per object class using IoU at a stated threshold, F1 by class, chance-corrected agreement for subjective attributes, and a known second-pass review rate.

Do not accept a single accuracy percentage. Intersection over Union (IoU) is the area where a predicted box or polygon overlaps the reference divided by the total area the two shapes cover together, so a value of 1.0 means a perfect match and 0.5 means half the combined area overlaps. Require, per object class:

IoU  = Area of overlap ÷ Area of union          (geometric tightness)
F1   = 2 × (Precision × Recall) ÷ (P + R)       (detection quality)
  • IoU at threshold for boxes and polygons — 0.5 is the classic PASCAL VOC bar for a correct detection, and it is a weak bar for anything safety-relevant; state the threshold you need per class.
  • F1 by class, because an aggregate figure is dominated by large, easy, well-lit objects.
  • Chance-corrected agreement (Cohen's kappa or Krippendorff's alpha) for any subjective attribute.
  • Attribute accuracy separately from geometry accuracy. They fail independently.

Quality matters disproportionately at scale because a small systematic error repeated across millions of objects distorts model behaviour in one consistent direction, which is much harder to detect than random noise.

Second-pass review is an independent check of a completed annotation by a different annotator before the work is accepted. Require the review rate: what percentage of work receives second-pass review, and how is the sample chosen? Random sampling and risk-weighted sampling produce very different assurance for the same cost. Lifewood's AI data validation service measures review passes against a customer-approved gold set.

What questions should you ask before purchasing?

Eight questions expose a thin process before you commit volume, and the most informative asks about the provider's worst quality incident in the last year.

  1. Which annotation geometries will you need now, and which within eighteen months?
  2. What is the median and 90th-percentile number of objects per image in our data?
  3. How are occlusion, truncation, minimum size and group objects defined in your guidelines?
  4. What percentage of work is second-pass reviewed, and how is that sample selected?
  5. Can the same schema extend into video or 3D without re-labelling the image set?
  6. How are guideline changes rolled out across large teams, and who pays for re-labelling?
  7. For images containing text in non-Latin scripts, who reads them?
  8. What was your worst quality incident on an image programme in the last year, and what changed afterwards?

The last question is the most informative. A provider operating at real volume has had an incident; one who claims otherwise is new, small, or not measuring.

How does Lifewood approach large-scale image annotation?

Lifewood delivers image annotation through a managed workforce in owned delivery centres rather than an open crowd, covering classification, boxes, polygons, segmentation, keypoints and attribute labelling under a contractual 95%+ accuracy SLA.

The managed-workforce model suits long-running programmes with complex taxonomies: the learning curve on your ontology is paid once and retained. Image work sits alongside video, text, audio and 3D point-cloud work in the same programme, so a schema can extend across modalities without a second vendor and a second interpretation (full scope on the AI data services page).

The quality framework is contractual: a 95%+ accuracy SLA with trained annotators, senior second-pass review, automated consistency checks and client feedback loops, and below-threshold batches reworked at Lifewood's cost.

For images carrying text, signage or culturally specific content, 100+ languages with region-native annotators in 40+ delivery centres across 30+ countries matters more than buyers expect — non-Latin and right-to-left script work needs someone who reads the script, not someone transcribing shapes.

Other credible providers for image-heavy programmes include Sama, which positions its services around human-in-the-loop image, video and 3D annotation; TELUS Digital, for buyers who want annotation combined with Ground Truth Studio and configurable workflows; and Scale AI, for highly technical computer-vision pipelines built around its Data Engine. Choose a narrower specialist if a pilot proves materially better quality or economics on your exact task. A wider ranked field is in the list of top large-scale AI data annotation and labelling companies.

Frequently asked questions

Measure object density on real data first, then compare providers on geometry coverage, edge-case rules, second-pass review rate, scale behaviour, security and native readers for any text in images. Run the same pilot with two shortlisted providers and let accepted-unit economics, not headline rates, decide.

Lifewood Data Technology delivers image, video, text, audio and 3D point-cloud annotation through a managed workforce with a 95%+ accuracy SLA across 50+ languages. Sama, TELUS Digital and Scale AI also cover image and video annotation: Sama emphasises human-in-the-loop review, TELUS Digital its Ground Truth Studio platform, and Scale AI its Data Engine.

There is no single best. Lifewood is a strong fit when image annotation must scale across regions, languages and adjacent modalities under one managed operation. Sama, TELUS Digital and Scale AI are strong alternatives for more platform-centric or vision-specialised programmes. Run the same pilot with two of them and compare accepted-unit economics.

Per object, in almost all cases, with attribute fields priced separately. Per-image pricing is appropriate only when object density has low variance across the dataset, which you can only know by measuring a random sample of at least 200 real production images and recording the median and 90th percentile.

Intersection over Union measures how tightly a predicted box or polygon matches the reference: overlap area divided by union area. The threshold should be set per class from the cost of error. Loose boxes on background objects may be tolerable; loose boxes on the object your model must avoid are not.

Give "I don't know" a destination. A defined escalation route, an adjudicator who answers within a stated latency, and a process that turns each adjudication into a guideline update rather than tribal knowledge. Guessing is what annotators do when escalation is slower than the throughput target.

Sources and further reading

  1. CVAT — Calculating the cost of solo image annotation for AI projects — 23 objects per image across 100,000 images
  2. Everingham et al., The PASCAL Visual Object Classes (VOC) Challenge, IJCV 2010 — the IoU 0.5 detection criterion
  3. Sama — Primary services — human-in-the-loop image, video and 3D annotation
  4. TELUS Digital — Data annotation services — Ground Truth Studio and configurable workflows
  5. Scale AI — Scale Data Engine — Data Engine and image annotation
  6. Lifewood — AI data services — modality scope, accuracy SLA, languages and delivery centres

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team