Skip to main content
AI Data

How to Compare Data Annotation Vendor Quotes

June 2026 · 9 min read · Updated September 2026

Short answer. Annotation quotes are not comparable as issued, and comparing them anyway is how programmes end up with the most expensive cheap vendor. Normalise every quote to the same task definition first — billable unit, QA inclusion, rework rules, minimum commitment, platform fees, setup, expert labour tiers, security premium and turnaround — then compare on cost per accepted unit rather than cost per attempted label. A vendor who prices an enterprise programme without inspecting representative data has guessed.

Key takeaways

  • Data annotation vendors quote per object, per image, per frame or per hour, so a spreadsheet that lines up raw prices produces a ranking that is wrong.
  • Cost per accepted unit is total project cost divided by the units that pass an agreed acceptance standard, and it is the only price that includes rework.
  • Acceptance criteria and a client-approved gold set must be agreed before prices are compared, otherwise each vendor is measured against its own definition of accepted.
  • A precise enterprise price issued without any review of representative data is a red flag, because it has priced an assumption about object density, ambiguity and input quality.
  • Proposals narrow the field; a normalised pilot with the same sample, guidelines, gold set and delivery window is what decides it.

Why are annotation vendor quotes not comparable as issued?

Annotation vendors price different things under the same heading, so the lowest headline number rarely corresponds to the lowest total cost.

Procurement teams are good at comparing prices and are given nine documents that do not price the same thing. One quotes per object with QA included; one quotes per image with QA as a line item; one quotes hourly with a platform licence attached; one quotes low but requires a minimum commitment three times your forecast volume. The spreadsheet that lines these up produces a ranking, and the ranking is wrong.

Quote normalisation is the process of restating every vendor proposal against one shared task definition, billable unit and acceptance standard so that the prices can be compared on equal terms. The choice of billable unit is covered in the guide to data annotation pricing models; the top large-scale annotation companies listicle ranks the vendors most buyers are comparing.

How do you normalise annotation quotes before comparing them?

Request the same nine items from every vendor and restate each quote against a single task definition before anyone on the evaluation team sees a price.

Quote item Why it matters What procurement should request
Billable unit Vendors quote per object, image, frame or hour One common unit, or an explicit conversion model
QA included? Cheap first-pass labels may exclude review Exact reviewer percentage and adjudication process
Rework Errors can create a second, hidden invoice Acceptance definition and included rework volume
Minimum volume Low rates may require large commitments Minimum spend and unused-capacity terms
Platform fee Service price may exclude tooling Licence, storage and integration fees
Setup and training Complex guidelines require calibration One-time onboarding and change-request fees
Expert labour Specialists radically change cost Separate rates by skill tier
Security Restricted facilities add cost Any security or data-residency premium
Turnaround Urgent scale requires premium staffing Standard versus expedited SLA, priced separately

Do this before the prices are visible to the evaluation team, if you can. Once someone has seen a low number, normalisation feels like moving the goalposts rather than like measurement. The same nine items form the pricing section of an annotation vendor RFP.

What is cost per accepted unit and why does it matter?

Cost per accepted unit is the total project cost divided by the units that pass the agreed acceptance criteria, and it is the metric to compare vendors on because a price per attempted label ignores rework.

Cost per accepted unit is the total amount paid to a vendor, including rework, fees and setup, divided by the number of delivered units that meet a client-defined acceptance standard. Suppose Vendor A charges 20% less per attempted label and generates materially more rework. The apparent saving disappears once rejected units, reviewer time and model-team delay are included.

Cost per accepted unit = Total project cost ÷ Units passing agreed acceptance criteria
Effective throughput   = Items delivered × First-pass acceptance rate ÷ Cycle time

Two things have to be true for this metric to work, and both are worth insisting on:

  • The acceptance criteria are agreed before pricing is compared. Otherwise each vendor is measured against their own definition of accepted, which is the problem you were trying to solve. The guide to what accuracy standard to require from an annotation vendor covers how to write that acceptance definition.
  • Rework is measured against your gold set, not the vendor's. A vendor's internal QA measures the vendor's interpretation of the guidelines. Only a client-approved gold set measures yours, which is why gold sets, audit sampling and consensus are not interchangeable QA methods.

A gold set is a sample of items labelled and approved by the client that serves as the reference answer key against which a vendor's accuracy and rework are measured. Rework is also paid in schedule. A batch returned in week six delays a training run, and the cost of that delay usually exceeds the cost of the labels.

Why should a vendor inspect your data before quoting?

A quote issued without seeing the data has priced an assumption about object density, ambiguity and input quality, and when the assumption proves wrong the buyer pays for it. Either the vendor comes back for a change order, or the vendor absorbs the difference and quality drops to fit the price.

TELUS Digital's published guidance on selecting a data annotation company recommends choosing an agency that refrains from quoting a price before it has reviewed your data, because price can vary widely depending on the service or data type. In enterprise procurement, an unseen-data quote is a guess with a decimal point.

A credible quote is based on representative sample data, the actual guidelines, and stated production assumptions you can check. Ask which assumptions the price depends on, and what happens to the price if each is wrong.

How do you score annotation vendor proposals?

Score normalised proposals against weighted criteria agreed before the proposals arrived, with quality and acceptance carrying the largest weight. The weights below are a starting point for an enterprise buyer; adjust them for your risk profile, but agree them before you see the proposals.

Criterion Suggested weight What a strong provider shows
Quality and acceptance 30% Measured QA, adjudication process, written rework policy
Unit economics 20% Transparent normalised price, no undisclosed fees
Scale and throughput 15% Proven ramp plan and reviewer capacity at peak
Modality fit 10% Tools and trained teams for your exact task type
Language and expertise 10% Named native speakers or domain specialists per requirement
Security and governance 10% Controls matching your requirements, with scope statements
Commercial flexibility 5% Reasonable minimums, change terms, exit provisions

Apply one disqualification rule before scoring: a zero on quality and acceptance is disqualifying regardless of total score. It is the one dimension that cannot be fixed after signature by paying more. Independent AI data validation of a vendor's delivered batches is one way to make the quality score a measurement rather than a claim.

What are the red flags in an annotation quote?

A quote that omits the billable unit, the acceptance rule or the party who pays for rework is signalling that the price depends on something the vendor would rather not state.

  • A precise enterprise price arrives without any review of representative data.
  • The quote does not define what counts as one billable unit.
  • QA or rework is described as "included" with no measurable acceptance rule.
  • A low unit rate depends on a minimum commitment you have not modelled.
  • The provider cannot separate generalist, specialist and expert labour pricing.
  • Platform, storage, integration or training charges appear only after selection.
  • The vendor cannot explain how pricing changes when guidelines change.
  • Every question about quality is answered with a percentage and no denominator.

How do you run a pilot that makes vendors comparable?

Give every shortlisted vendor an identical pilot and score the results on cost per accepted unit, per-class quality and behaviour on ambiguous items. The proposal round narrows the field; the pilot decides it.

Give each shortlisted provider:

  • The same representative sample, including your hardest edge cases and at least one difficult language
  • The same guidelines, at the same version
  • The same acceptance criteria and gold set
  • The same delivery window

Then score cost per accepted unit, per-class quality, escalation behaviour on genuinely ambiguous items, and the quality of questions asked during onboarding. That last signal is undervalued: a vendor who asks six precise questions about your ontology in week one is a vendor who will not silently guess in month six.

Make the pilot long enough to expose representative edge cases and to measure throughput after the initial learning curve, which usually means weeks rather than days. A pilot short enough to be staffed entirely by a vendor's best annotators measures the vendor's best annotators. What happens after the pilot is won is covered in the guide to scaling annotation from pilot to production.

How does Lifewood price and guarantee annotation quality?

Lifewood does not compete on headline unit rate and does not publish one; its public proposition is a defined quality target with commercial consequences attached. Below-threshold batches are reworked at Lifewood's cost, which moves rework from a hidden second invoice to a vendor obligation.

The published terms are a 95%+ accuracy SLA, dual-layer human review with automated consistency checks, and rework of below-threshold batches at Lifewood's cost. For a buyer normalising quotes, that is the term that matters most.

The breadth arguments are about total cost rather than unit price. 100+ languages reduces the need to assemble and manage separate language vendors. Coverage of text, image, audio, video and 3D point-cloud work across Lifewood's AI data services reduces vendor fragmentation and the reconciliation cost between two interpretations of the same ontology. 40+ delivery centres across 30+ countries gives options for distributed production and for regional processing requirements without a second supplier relationship.

A specialist vendor may still win a specific workstream if its pilot demonstrates materially better accepted-unit economics on that task. That is the correct outcome of a well-run comparison, and it is worth designing the process so it can happen.

Frequently asked questions

Give each shortlisted provider the same representative pilot, with the same guidelines, acceptance criteria and gold set, and score cost per accepted unit alongside quality by defect class, throughput, management effort and commercial terms. Proposals priced on different units are not comparable; identical pilots scored on one acceptance standard are.

Normalise every quote to one billable unit and one acceptance standard, then run an identical pilot on your hardest images or frames, including occlusion and low-light edge cases. Score cost per accepted unit and per-class quality, and disqualify any vendor that priced the work without inspecting representative data.

Usually not on price alone. A lower rate is rational for simple, standardised, low-ambiguity work. For enterprise programmes, the costs that decide the outcome are rework, schedule slippage and coordination overhead, none of which appear in the unit rate, and all of which appear in cost per accepted unit.

Convert everything to cost per accepted unit against a single agreed task definition, asking each vendor for an explicit conversion model between objects, images, frames and hours. If a vendor resists providing a conversion, that resistance is itself information: it usually means the quote depends on an assumption they would rather not state.

Many providers use volume or commitment-based pricing. CVAT's published pricing example shows a one-time per-object rate of $0.10 falling to $0.05–0.075 under a prepaid subscription, a discount it describes as typically 20% to 50%. Model your expected utilisation first: a commitment your roadmap does not produce converts a discount into a penalty.

A missing or undefined acceptance standard, an unwillingness to inspect representative data before quoting, and an inability to state who pays for rework. These three cannot be corrected after signature by spending more money, so they belong before the weighted scorecard rather than inside it.

Sources and further reading

  1. TELUS Digital: How to select the right data annotation company — the recommendation to choose an agency that does not quote before reviewing your data
  2. CVAT: Annotation services pricing — per-object rate examples under one-time versus prepaid subscription pricing
  3. Lifewood QA process — 95% accuracy SLA, dual-layer review and rework of below-threshold batches at Lifewood's cost

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team