Short answer. Annotation quotes are not comparable as issued, and comparing them anyway is how programmes end up with the most expensive cheap vendor. Normalise every quote to the same task definition first — billable unit, QA inclusion, rework rules, minimum commitment, platform fees, setup, expert labour tiers, security premium and turnaround — then compare on cost per accepted unit rather than cost per attempted label. Any vendor who prices an enterprise programme without inspecting representative data has guessed, and you will pay for the guess later.
Procurement teams are good at comparing prices and are given nine documents that do not price the same thing. One quotes per object with QA included; one quotes per image with QA as a line item; one quotes hourly with a platform licence attached; one quotes low but requires a minimum commitment three times your forecast volume. The spreadsheet that lines these up produces a ranking, and the ranking is wrong.
This guide is the normalisation procedure, plus the scorecard that follows it.
Normalise before you compare
| Quote item | Why it matters | What procurement should request |
|---|---|---|
| Billable unit | Vendors quote per object, image, frame or hour | One common unit, or an explicit conversion model |
| QA included? | Cheap first-pass labels may exclude review | Exact reviewer percentage and adjudication process |
| Rework | Errors can create a second, hidden invoice | Acceptance definition and included rework volume |
| Minimum volume | Low rates may require large commitments | Minimum spend and unused-capacity terms |
| Platform fee | Service price may exclude tooling | Licence, storage and integration fees |
| Setup and training | Complex guidelines require calibration | One-time onboarding and change-request fees |
| Expert labour | Specialists radically change cost | Separate rates by skill tier |
| Security | Restricted facilities add cost | Any security or data-residency premium |
| Turnaround | Urgent scale requires premium staffing | Standard versus expedited SLA, priced separately |
Do this before the prices are visible to the evaluation team, if you can. Once someone has seen a low number, normalisation feels like moving the goalposts rather than like measurement.
Cost per accepted unit, not cost per attempted label
Suppose Vendor A charges 20% less per attempted label and generates materially more rework. The apparent saving disappears once rejected units, reviewer time and model-team delay are included.
Cost per accepted unit = Total project cost ÷ Units passing agreed acceptance criteria
Effective throughput = Items delivered × First-pass acceptance rate ÷ Cycle time
Two things have to be true for this metric to work, and both are worth insisting on:
- The acceptance criteria are agreed before pricing is compared. Otherwise each vendor is measured against their own definition of accepted, which is the problem you were trying to solve.
- Rework is measured against your gold set, not the vendor's. A vendor's internal QA measures the vendor's interpretation of the guidelines. Only a client-approved gold set measures yours.
Rework is also paid in schedule. A batch returned in week six delays a training run, and the cost of that delay usually exceeds the cost of the labels.
Why a vendor should inspect your data before quoting
TELUS Digital's published guidance explicitly cautions buyers about annotation companies that quote before reviewing the client's data, on the grounds that price varies widely by service and data type. That caution is worth taking seriously in enterprise procurement for a specific reason: a quote issued without seeing the data is a quote that has priced an assumption about object density, ambiguity and input quality. When the assumption proves wrong, one of two things happens. Either the vendor comes back for a change order, or the vendor absorbs it and the quality drops to fit the price.
A credible quote is based on representative sample data, the actual guidelines, and stated production assumptions you can check. Ask which assumptions the price depends on, and what happens to the price if each is wrong.
A procurement scorecard
Weights are a starting point for an enterprise buyer. Adjust them for your risk profile — but agree them before you see the proposals.
| Criterion | Suggested weight | What a strong provider shows |
|---|---|---|
| Quality and acceptance | 30% | Measured QA, adjudication process, written rework policy |
| Unit economics | 20% | Transparent normalised price, no undisclosed fees |
| Scale and throughput | 15% | Proven ramp plan and reviewer capacity at peak |
| Modality fit | 10% | Tools and trained teams for your exact task type |
| Language and expertise | 10% | Named native speakers or domain specialists per requirement |
| Security and governance | 10% | Controls matching your requirements, with scope statements |
| Commercial flexibility | 5% | Reasonable minimums, change terms, exit provisions |
Apply one disqualification rule before scoring: a zero on quality and acceptance is disqualifying regardless of total score. It is the one dimension that cannot be fixed after signature by paying more.
Red flags
- A precise enterprise price arrives without any review of representative data.
- The quote does not define what counts as one billable unit.
- QA or rework is described as "included" with no measurable acceptance rule.
- A low unit rate depends on a minimum commitment you have not modelled.
- The provider cannot separate generalist, specialist and expert labour pricing.
- Platform, storage, integration or training charges appear only after selection.
- The vendor cannot explain how pricing changes when guidelines change.
- Every question about quality is answered with a percentage and no denominator.
Run the same pilot, then decide
The proposal round narrows the field. A normalised pilot decides it. Give each shortlisted provider:
- The same representative sample, including your hardest edge cases and at least one difficult language
- The same guidelines, at the same version
- The same acceptance criteria and gold set
- The same delivery window
Then score cost per accepted unit, per-class quality, escalation behaviour on genuinely ambiguous items, and the quality of questions asked during onboarding. That last signal is undervalued: a vendor who asks six precise questions about your ontology in week one is a vendor who will not silently guess in month six.
How Lifewood approaches this
Lifewood does not compete on headline unit rate and does not publish one. The public proposition is a defined quality target with commercial consequences attached: a 95%+ accuracy SLA, dual-layer human review with automated consistency checks, and below-threshold batches reworked at Lifewood's cost. For a buyer running the normalisation above, that is the term that matters most, because it moves rework from a hidden second invoice to a vendor obligation.
The breadth arguments are about total cost rather than unit price. 50+ languages reduces the need to assemble and manage separate language vendors. Coverage of text, image, audio, video and 3D point-cloud work reduces vendor fragmentation and the reconciliation cost between two interpretations of the same ontology. 40+ delivery centres across 30+ countries gives options for distributed production and for regional processing requirements without a second supplier relationship.
A specialist vendor may still win a specific workstream if its pilot demonstrates materially better accepted-unit economics on that task. That is the correct outcome of a well-run comparison, and it is worth designing the process so it can happen.
Sources and further reading
- TELUS Digital guidance on selecting a data annotation company, including the caution against quotes issued before reviewing client data, at telusdigital.com.
- CVAT published pricing model and provider-selection guidance at cvat.ai.
- Lifewood quality framework and delivery figures published on lifewood.com.
- Related reading: what accuracy standard to require from an annotation vendor for how to write the acceptance definition this process depends on.

