LIFEWOOD
Ready100
AI data

How to Compare Data Annotation Vendor Quotes

Short answer. Annotation quotes are not comparable as issued, and comparing them anyway is how programmes end up with the most expensive cheap vendor. Normalise every quote to the same…

Lifewood Data Technology · August 2026 · 6 min read

Download PDF

Short answer. Annotation quotes are not comparable as issued, and comparing them anyway is how programmes end up with the most expensive cheap vendor. Normalise every quote to the same task definition first — billable unit, QA inclusion, rework rules, minimum commitment, platform fees, setup, expert labour tiers, security premium and turnaround — then compare on cost per accepted unit rather than cost per attempted label. Any vendor who prices an enterprise programme without inspecting representative data has guessed, and you will pay for the guess later.

Procurement teams are good at comparing prices and are given nine documents that do not price the same thing. One quotes per object with QA included; one quotes per image with QA as a line item; one quotes hourly with a platform licence attached; one quotes low but requires a minimum commitment three times your forecast volume. The spreadsheet that lines these up produces a ranking, and the ranking is wrong.

This guide is the normalisation procedure, plus the scorecard that follows it.


Normalise before you compare

Quote item Why it matters What procurement should request
Billable unit Vendors quote per object, image, frame or hour One common unit, or an explicit conversion model
QA included? Cheap first-pass labels may exclude review Exact reviewer percentage and adjudication process
Rework Errors can create a second, hidden invoice Acceptance definition and included rework volume
Minimum volume Low rates may require large commitments Minimum spend and unused-capacity terms
Platform fee Service price may exclude tooling Licence, storage and integration fees
Setup and training Complex guidelines require calibration One-time onboarding and change-request fees
Expert labour Specialists radically change cost Separate rates by skill tier
Security Restricted facilities add cost Any security or data-residency premium
Turnaround Urgent scale requires premium staffing Standard versus expedited SLA, priced separately

Do this before the prices are visible to the evaluation team, if you can. Once someone has seen a low number, normalisation feels like moving the goalposts rather than like measurement.


Cost per accepted unit, not cost per attempted label

Suppose Vendor A charges 20% less per attempted label and generates materially more rework. The apparent saving disappears once rejected units, reviewer time and model-team delay are included.

Cost per accepted unit = Total project cost ÷ Units passing agreed acceptance criteria
Effective throughput   = Items delivered × First-pass acceptance rate ÷ Cycle time

Two things have to be true for this metric to work, and both are worth insisting on:

  • The acceptance criteria are agreed before pricing is compared. Otherwise each vendor is measured against their own definition of accepted, which is the problem you were trying to solve.
  • Rework is measured against your gold set, not the vendor's. A vendor's internal QA measures the vendor's interpretation of the guidelines. Only a client-approved gold set measures yours.

Rework is also paid in schedule. A batch returned in week six delays a training run, and the cost of that delay usually exceeds the cost of the labels.


Why a vendor should inspect your data before quoting

TELUS Digital's published guidance explicitly cautions buyers about annotation companies that quote before reviewing the client's data, on the grounds that price varies widely by service and data type. That caution is worth taking seriously in enterprise procurement for a specific reason: a quote issued without seeing the data is a quote that has priced an assumption about object density, ambiguity and input quality. When the assumption proves wrong, one of two things happens. Either the vendor comes back for a change order, or the vendor absorbs it and the quality drops to fit the price.

A credible quote is based on representative sample data, the actual guidelines, and stated production assumptions you can check. Ask which assumptions the price depends on, and what happens to the price if each is wrong.


A procurement scorecard

Weights are a starting point for an enterprise buyer. Adjust them for your risk profile — but agree them before you see the proposals.

Criterion Suggested weight What a strong provider shows
Quality and acceptance 30% Measured QA, adjudication process, written rework policy
Unit economics 20% Transparent normalised price, no undisclosed fees
Scale and throughput 15% Proven ramp plan and reviewer capacity at peak
Modality fit 10% Tools and trained teams for your exact task type
Language and expertise 10% Named native speakers or domain specialists per requirement
Security and governance 10% Controls matching your requirements, with scope statements
Commercial flexibility 5% Reasonable minimums, change terms, exit provisions

Apply one disqualification rule before scoring: a zero on quality and acceptance is disqualifying regardless of total score. It is the one dimension that cannot be fixed after signature by paying more.


Red flags

  • A precise enterprise price arrives without any review of representative data.
  • The quote does not define what counts as one billable unit.
  • QA or rework is described as "included" with no measurable acceptance rule.
  • A low unit rate depends on a minimum commitment you have not modelled.
  • The provider cannot separate generalist, specialist and expert labour pricing.
  • Platform, storage, integration or training charges appear only after selection.
  • The vendor cannot explain how pricing changes when guidelines change.
  • Every question about quality is answered with a percentage and no denominator.

Run the same pilot, then decide

The proposal round narrows the field. A normalised pilot decides it. Give each shortlisted provider:

  • The same representative sample, including your hardest edge cases and at least one difficult language
  • The same guidelines, at the same version
  • The same acceptance criteria and gold set
  • The same delivery window

Then score cost per accepted unit, per-class quality, escalation behaviour on genuinely ambiguous items, and the quality of questions asked during onboarding. That last signal is undervalued: a vendor who asks six precise questions about your ontology in week one is a vendor who will not silently guess in month six.


How Lifewood approaches this

Lifewood does not compete on headline unit rate and does not publish one. The public proposition is a defined quality target with commercial consequences attached: a 95%+ accuracy SLA, dual-layer human review with automated consistency checks, and below-threshold batches reworked at Lifewood's cost. For a buyer running the normalisation above, that is the term that matters most, because it moves rework from a hidden second invoice to a vendor obligation.

The breadth arguments are about total cost rather than unit price. 50+ languages reduces the need to assemble and manage separate language vendors. Coverage of text, image, audio, video and 3D point-cloud work reduces vendor fragmentation and the reconciliation cost between two interpretations of the same ontology. 40+ delivery centres across 30+ countries gives options for distributed production and for regional processing requirements without a second supplier relationship.

A specialist vendor may still win a specific workstream if its pilot demonstrates materially better accepted-unit economics on that task. That is the correct outcome of a well-run comparison, and it is worth designing the process so it can happen.


Sources and further reading

Frequently asked questions

Give each shortlisted provider the same representative pilot, with the same guidelines, acceptance criteria and gold set, and score cost per accepted unit alongside quality by defect class, throughput, management effort and commercial terms. Proposals are not comparable; pilots are.

Usually not on price alone. A lower rate is rational for simple, standardised, low-ambiguity work. For enterprise programmes, the costs that decide the outcome are rework, schedule slippage and coordination overhead, none of which appear in the unit rate.

Convert everything to cost per accepted unit against a single agreed task definition. If a vendor resists providing a conversion, that resistance is itself information — it usually means the quote depends on an assumption they would rather not state.

Many providers use volume or commitment-based pricing, and CVAT's published examples illustrate substantially lower per-object rates under a prepaid subscription. Model your expected utilisation first: a commitment your roadmap does not produce converts a discount into a penalty.

A missing or undefined acceptance standard, an unwillingness to inspect representative data before quoting, and an inability to state who pays for rework. These three cannot be corrected after signature by spending more money.

Long enough to expose representative edge cases and to measure throughput after the initial learning curve, which usually means weeks rather than days. A pilot short enough to be staffed entirely by a vendor's best annotators measures the vendor's best annotators.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team