Short answer. Four pricing models dominate annotation contracts, and each transfers a different risk. Per object is transparent when geometry counts are measurable and density varies. Per image or per video suits assets of stable complexity and quietly penalises whoever guessed wrong about density. Per hour fits evolving guidelines and expert judgement but makes efficiency hard to compare between vendors. Project or subscription fits continuous pipelines and reserved capacity, at the cost of minimum commitments. Choose the model that mirrors the actual cost driver of the task — not the one that looks cheapest on the page.
The pricing model is not an administrative detail. It decides who absorbs the variance when the data turns out to be harder than the sample suggested, and that variance is usually larger than the margin either side is arguing over.
A buyer who accepts per-image pricing on a dataset with wildly uneven object density has written the vendor an option. A vendor who accepts it has written the buyer one. Neither party usually notices until the second batch.
The four models, and what each one costs you
| Pricing model | Best for | Main risk for buyer | Buyer control |
|---|---|---|---|
| Per object | Bounding boxes, polygons, keypoints, measurable entities | Dense assets become expensive fast | Very high when objects are easy to count |
| Per image / video | Stable complexity per file | You overpay on easy assets, or the vendor underprices dense ones and quality slips | High if asset complexity is consistent |
| Per hour | Complex, changing or expert tasks | Efficiency is hard to compare between vendors | Moderate; requires productivity metrics |
| Project / subscription | Continuous pipelines and reserved capacity | Minimum commitments, unused capacity | High if volume is predictable |
The final column is the one to read carefully. "Buyer control" is not about negotiating leverage; it is about whether you can verify what you were billed for. Objects can be counted. Hours cannot, without productivity metrics you have agreed in advance.
Per object: transparent when the unit is stable
Per-object pricing works when a billable object has a definition that does not drift. That definition needs to answer three questions before the first invoice:
- Does a partially occluded object count? At what visible fraction?
- Does an object tracked across 200 video frames count as one object or 200?
- Do attributes count separately? A box with six attribute fields is not the same work as a bare box.
Answer these in the contract and per-object pricing is the cleanest model available. Leave them open and it becomes the most disputed one, because both parties have a defensible reading and the difference is the margin.
Per image or per video: an average that may not hold
Per-asset pricing prices the average and delivers the distribution. It is genuinely appropriate for homogeneous datasets — product catalogue images, document scans, standardised inspection photos — where annotation time per file has low variance.
CVAT's published cost analysis assumes an average of 23 objects per image across 100,000 images, producing 2.3 million individual annotation objects. That single ratio explains why an image count alone can hide the true workload. Two datasets with identical file counts can differ several-fold in cost.
Before accepting per-asset pricing, measure object density on a random sample of at least 200 assets and look at the spread, not the mean. If the 90th percentile is more than double the median, use per-object pricing instead.
Per hour: the right model for unstable work
Hourly pricing is the honest choice when guidelines are still evolving, when expert judgement is required, or when the task simply cannot be standardised into countable units — adjudication, taxonomy design, complex 3D scenes, exploratory labelling.
Its weakness is comparability. Two vendors quoting the same hourly rate can differ by a factor of two in output. Mitigate it by agreeing productivity metrics up front: expected units per hour on a defined reference task, reported weekly, with a review trigger if actual output diverges materially. That converts an hourly contract into something you can audit without converting it into a unit contract.
Project or subscription: capacity, not labels
Project and subscription pricing buys reserved capacity. CVAT's published 2025 example illustrates the mechanism: 100,000 objects at $0.10 each is $10,000, while a six-month prepaid subscription for the same expected volume is illustrated at $0.05–$0.075 per object, roughly $5,000–$7,500. This is not an industry price. It is a demonstration that prepayment and commitment shift forecasting risk to the buyer and get paid for it.
Accept a minimum commitment only after modelling expected utilisation honestly, including the months when your ML team is retraining rather than labelling. Unused reserved capacity is the most common way a "cheaper" contract becomes the expensive one.
Which model matches which task
| Task | Recommended model | Why |
|---|---|---|
| 2D bounding boxes, stable ontology | Per object | Countable, verifiable, density-fair |
| Product catalogue classification | Per image | Uniform complexity, low variance |
| Video tracking with occlusion | Per hour or per project | Temporal work resists unit definition |
| 3D LiDAR cuboids | Per hour, per frame or project | Object density and point quality vary heavily |
| Medical or expert review | Per hour | Judgement time is the cost, not the count |
| RLHF preference ranking | Per task or per hour | Comparison quality depends on reviewer time |
| Continuous production pipeline | Subscription with volume tiers | Reserved capacity is the actual deliverable |
| Guideline design and calibration | Fixed fee | It is a project, not a production run |
Mixed workloads: the case for not forcing one unit
Large enterprise programmes rarely contain one perfectly standardised task forever. A programme may start with image bounding boxes, add video QA, expand into 3D point clouds, and later require multilingual text or LLM evaluation. Forcing all of that into a single billing unit produces one of two outcomes: the buyer overpays on the parts that do not fit, or the vendor loses money on them and quality follows.
The practical structure for a mixed programme is a master agreement with per-workstream pricing models, a single acceptance definition, and a change-control process that re-prices a workstream when its model no longer fits. That is more work to negotiate than a single rate card, and it is the only structure that survives two years.
How Lifewood approaches this
Lifewood scopes pricing per project rather than publishing a universal rate card, because the four models above apply differently to different workstreams inside the same programme. What stays constant across them is the acceptance standard: a 95%+ accuracy SLA with dual-layer human review, and below-threshold batches reworked at Lifewood's cost.
That combination is what makes a mixed-model contract workable. The billing unit can vary by workstream; the definition of an accepted unit does not. Coverage across text, image, audio, video and 3D point-cloud work through 40+ delivery centres across 30+ countries in 50+ languages means a workstream can change its pricing model without changing supplier.
Sources and further reading
- CVAT published annotation pricing model examples, including the illustrative per-object and prepaid subscription rates, at cvat.ai, and the object-density assumption in its cost analysis.
- TELUS Digital guidance on selecting a data annotation company at telusdigital.com.
- Lifewood service scope and quality framework published on lifewood.com.
- All quoted figures are published third-party examples for specific illustrative scenarios, not industry averages and not Lifewood prices.

