Short answer. Image, video and 3D/LiDAR annotation should never be budgeted from the same unit price. Cost rises with geometry and review effort, not with file count: classification is cheapest, boxes and keypoints add labour, segmentation adds boundary precision, video adds temporal-consistency work, and 3D/LiDAR adds spatial interpretation plus cross-sensor review. CVAT's published example of 23 objects per image across 100,000 images shows how file counts hide workload. Budget by geometry and QA effort per object, then validate on a real sample.
Key takeaways
- Annotation cost scales with the number of objects, the geometry of each label and the depth of review, not with the number of images or clips.
- CVAT's published outsourcing example assumes an average of 23 objects per image across 100,000 images, which produces roughly 2.3 million annotation objects.
- Video annotation costs more than image annotation because track identity, occlusion and interpolation rules create sequence-level review work that does not exist in still images.
- 3D and LiDAR annotation costs more again because annotators position cuboids in three axes plus yaw on sparse point clouds and reconcile them across cameras, LiDAR and radar.
- A defensible multimodal budget models each modality on measured objects per asset, then adds onboarding, guideline-change and rework provisions once.
Why are most annotation budgets built from the wrong number?
Most annotation budgets are built from a file count multiplied by a rate quoted for a different dataset, and that figure survives only until the first invoice. The number that actually drives cost is the object count, the geometry of each label and the review depth the quality target demands.
Object density is the mean number of labelled objects per asset, measured on production data under the buyer's own ontology, and it is the single figure that most annotation budgets assume rather than measure.
This guide sets out the relative cost structure across modalities, why each one behaves the way it does, and what to measure before committing a budget. It sits alongside Lifewood's guide to data annotation pricing models, which covers per-object, per-hour and per-project billing in general terms; this post is about how the modality itself moves the cost.
How does annotation cost differ by type of task?
Annotation effort rises in a predictable order: image classification is cheapest, bounding boxes and keypoints cost more, polygons and segmentation cost more again, video tracking and 3D point clouds sit at the high end, and sensor fusion and expert medical annotation sit at the very top. Absolute rates depend on language, geography, quality target and security requirement, so the useful comparison is relative.
| Annotation type | Typical effort level | Main cost driver |
|---|---|---|
| Image classification | Low | Number of assets and classes |
| 2D bounding boxes | Low to medium | Objects per image and occlusion |
| Polygons / segmentation | Medium to high | Boundary precision and object shape |
| Keypoints / pose | Medium to high | Number of landmarks and visibility rules |
| Video tracking | High | Frames, object persistence, occlusion, identity consistency |
| 3D cuboids / point clouds | High | Point density, object count, spatial precision |
| Sensor fusion | Very high | Cross-view and cross-sensor consistency |
| Medical / expert visual annotation | Very high | Specialist labour and validation requirements |
The "effort level" column is deliberately relative; any table that pretends to give absolute rates across every geography and quality target is inventing numbers. For what each task type involves at procurement level, see the guides on buying large-scale image annotation and buying large-scale video annotation.
Why does price per image mislead?
Price per image misleads because the workload lives in the objects inside the image, not in the image itself, and object density varies by an order of magnitude between datasets. Two projects with identical file counts can differ tenfold in billable labour.
CVAT's published cost analysis for outsourced annotation assumes an average of 23 objects per image across 100,000 images, creating approximately 2.3 million individual annotation objects. Change that ratio to 5 and the same file count becomes a 500,000-object project. Change it to 60 and it becomes a 6-million-object project. The file count did not move.
Before you accept any per-image quote, do this:
- Take a random sample of at least 200 assets from actual production data, not a curated demo set.
- Count objects per asset under your own ontology.
- Record the median, the 90th percentile and the maximum.
- If the 90th percentile exceeds twice the median, per-image pricing is transferring real risk to somebody, and it will be resolved either in your invoice or in the vendor's quality.
The same exercise applies to attributes. An image with 23 boxes is one job; an image with 23 boxes each carrying six attribute fields is a substantially larger one. Lifewood's guide to comparing data annotation vendor quotes covers how to normalise quotes that use different units.
Why does video annotation cost more than image annotation?
Video annotation costs more because the same object must keep the same identity across hundreds of frames, and errors propagate along the sequence rather than staying inside one image. That creates structural review work, not simply more frames.
Temporal consistency is the requirement that a labelled object keeps the same track identity, class and attributes across every frame in which it appears, including across occlusions and re-entries.
Video annotation is not image annotation multiplied by frame count. The additional work is structural:
- Identity persistence. The same object must carry the same track ID across the sequence. This is a review problem more than a labelling problem, and it does not parallelise cleanly: splitting a clip across two annotators is exactly how track identity breaks.
- Entry and exit. Objects appear, leave and return. The guideline must define whether a returning object resumes its original ID, and annotators must apply that consistently across hours of footage.
- Occlusion rules. How long an object may be hidden before its track terminates, and what happens to the frames in between.
- Interpolation policy. What may be interpolated between keyframes, and what must be labelled directly. Aggressive interpolation is a legitimate cost saving and a legitimate source of systematic error, depending entirely on motion characteristics.
- Sequence-level review. Errors propagate for hundreds of frames, so reviewing random individual frames catches almost nothing. Review has to run on sequences, which costs more per unit of data reviewed.
For autonomous-driving and robotics datasets, add synchronised camera and sensor labels, which multiplies all of the above by the number of sensors.
Why do 3D and LiDAR annotation cost more?
3D and LiDAR annotation cost more because a human must interpret sparse or noisy point clouds, position cuboids accurately in three axes including yaw, and hold that consistent across views and across time. Sensor fusion adds a cross-modal reconciliation step on top of that.
Sensor fusion annotation is the labelling of the same physical object consistently across camera, LiDAR and radar data so that one identity links a 2D box, a 3D cuboid and a radar return.
Distant and reflective objects may return a handful of points, at which point the annotator is inferring rather than observing. The guideline must say when inference is permitted, when the object should be excluded, and when it should be escalated.
When camera, LiDAR and radar are combined, cross-modal review increases both the labour requirement and the consequences of an error. An object with a correct camera box, a correct LiDAR cuboid and a mismatched identity between them teaches the model that two geometries describe different things, which is worse than a missing label. The task-level specification behind that cost is set out in Lifewood's guide to autonomous driving data annotation requirements.
How do you build a defensible annotation budget?
A defensible budget works from billable objects and effort per object rather than from files, then adds review, rework, onboarding and guideline-change provisions as separate lines. The two lines most budgets omit are onboarding and the guideline-change allowance.
A review multiplier is the factor applied to base labour cost to fund the second-pass and audit effort a given accuracy target requires, and it rises with the quality threshold and the safety criticality of the classes involved.
Billable objects = Assets × Mean objects per asset
Base labour cost = Billable objects × Unit rate for that geometry
QA and review cost = Base labour cost × Review multiplier for your quality target
Rework provision = Base + QA, ÷ Expected first-pass acceptance rate
Total programme cost = Rework-adjusted cost + Onboarding + Guideline change allowance
Onboarding and calibration are real costs that appear once per vendor per taxonomy. A guideline-change allowance is a real cost that appears every time your ML team learns something, which on a healthy programme is often.
| Budget line | Commonly omitted? | Typical trigger |
|---|---|---|
| Objects per asset variance | Yes | Dense scenes in production data |
| Attribute fields per object | Yes | Ontology expansion after pilot |
| Sequence-level video review | Yes | Track identity errors found late |
| Cross-sensor reconciliation | Yes | Fusion work added to a 2D programme |
| Onboarding and calibration | Sometimes | New vendor or new taxonomy |
| Guideline change and re-labelling | Almost always | Any real research programme |
| Security or residency premium | Sometimes | Client or regulatory mandate |
All quoted figures in this guide are published third-party examples for specific scenarios. They are not industry averages and they are not Lifewood prices.
How does Lifewood approach multimodal annotation pricing?
Lifewood scopes multimodal programmes per project rather than publishing a rate card, because the cost structures of image, video and 3D work differ too much to reduce to one number. Three things affect total cost rather than unit rate.
Coverage across text, image, audio, video and 3D point-cloud work means one provider can hold a consistent schema and QA definition as a programme moves from 2D into video and then into sensor data. The alternative is a second vendor, a second onboarding and a permanent reconciliation task between two interpretations of the same ontology. Lifewood's AI data services span those modalities, and its autonomous driving annotation service covers LiDAR, camera and radar labelling with fusion alignment.
The quality framework is contractual rather than described: a 95%+ accuracy SLA, measured against a customer-approved gold set, with two independent review passes with timestamped approval records, automated consistency checks and client feedback loops, with below-threshold batches reworked at Lifewood's cost. For multimodal work specifically, that matters because the failure modes differ by modality and a single aggregate figure hides them; ask for acceptance reported per modality and per object class.
Scale across 40+ delivery centres in 30+ countries supports high-volume programmes that need sustained staffing in more than one production location, which is what continuous video and sensor pipelines require, since they cannot be batched down during a quiet week. Buyers comparing providers on this basis can start from Lifewood's list of the top autonomous driving annotation companies.