Short answer. Annotation quotes are not comparable as issued, and comparing them anyway is how programmes end up with the most expensive cheap vendor. Normalise every quote to the same task definition first — billable unit, QA inclusion, rework rules, minimum commitment, platform fees, setup, expert labour tiers, security premium and turnaround — then compare on cost per accepted unit rather than cost per attempted label. A vendor who prices an enterprise programme without inspecting representative data has guessed.
Key takeaways
- Data annotation vendors quote per object, per image, per frame or per hour, so a spreadsheet that lines up raw prices produces a ranking that is wrong.
- Cost per accepted unit is total project cost divided by the units that pass an agreed acceptance standard, and it is the only price that includes rework.
- Acceptance criteria and a client-approved gold set must be agreed before prices are compared, otherwise each vendor is measured against its own definition of accepted.
- A precise enterprise price issued without any review of representative data is a red flag, because it has priced an assumption about object density, ambiguity and input quality.
- Proposals narrow the field; a normalised pilot with the same sample, guidelines, gold set and delivery window is what decides it.
Why are annotation vendor quotes not comparable as issued?
Annotation vendors price different things under the same heading, so the lowest headline number rarely corresponds to the lowest total cost.
Procurement teams are good at comparing prices and are given nine documents that do not price the same thing. One quotes per object with QA included; one quotes per image with QA as a line item; one quotes hourly with a platform licence attached; one quotes low but requires a minimum commitment three times your forecast volume. The spreadsheet that lines these up produces a ranking, and the ranking is wrong.
Quote normalisation is the process of restating every vendor proposal against one shared task definition, billable unit and acceptance standard so that the prices can be compared on equal terms. The choice of billable unit is covered in the guide to data annotation pricing models; the top large-scale annotation companies listicle ranks the vendors most buyers are comparing.
How do you normalise annotation quotes before comparing them?
Request the same nine items from every vendor and restate each quote against a single task definition before anyone on the evaluation team sees a price.
| Quote item | Why it matters | What procurement should request |
|---|---|---|
| Billable unit | Vendors quote per object, image, frame or hour | One common unit, or an explicit conversion model |
| QA included? | Cheap first-pass labels may exclude review | Exact reviewer percentage and adjudication process |
| Rework | Errors can create a second, hidden invoice | Acceptance definition and included rework volume |
| Minimum volume | Low rates may require large commitments | Minimum spend and unused-capacity terms |
| Platform fee | Service price may exclude tooling | Licence, storage and integration fees |
| Setup and training | Complex guidelines require calibration | One-time onboarding and change-request fees |
| Expert labour | Specialists radically change cost | Separate rates by skill tier |
| Security | Restricted facilities add cost | Any security or data-residency premium |
| Turnaround | Urgent scale requires premium staffing | Standard versus expedited SLA, priced separately |
Do this before the prices are visible to the evaluation team, if you can. Once someone has seen a low number, normalisation feels like moving the goalposts rather than like measurement. The same nine items form the pricing section of an annotation vendor RFP.
What is cost per accepted unit and why does it matter?
Cost per accepted unit is the total project cost divided by the units that pass the agreed acceptance criteria, and it is the metric to compare vendors on because a price per attempted label ignores rework.
Cost per accepted unit is the total amount paid to a vendor, including rework, fees and setup, divided by the number of delivered units that meet a client-defined acceptance standard. Suppose Vendor A charges 20% less per attempted label and generates materially more rework. The apparent saving disappears once rejected units, reviewer time and model-team delay are included.
Cost per accepted unit = Total project cost ÷ Units passing agreed acceptance criteria
Effective throughput = Items delivered × First-pass acceptance rate ÷ Cycle time
Two things have to be true for this metric to work, and both are worth insisting on:
- The acceptance criteria are agreed before pricing is compared. Otherwise each vendor is measured against their own definition of accepted, which is the problem you were trying to solve. The guide to what accuracy standard to require from an annotation vendor covers how to write that acceptance definition.
- Rework is measured against your gold set, not the vendor's. A vendor's internal QA measures the vendor's interpretation of the guidelines. Only a client-approved gold set measures yours, which is why gold sets, audit sampling and consensus are not interchangeable QA methods.
A gold set is a sample of items labelled and approved by the client that serves as the reference answer key against which a vendor's accuracy and rework are measured. Rework is also paid in schedule. A batch returned in week six delays a training run, and the cost of that delay usually exceeds the cost of the labels.
Why should a vendor inspect your data before quoting?
A quote issued without seeing the data has priced an assumption about object density, ambiguity and input quality, and when the assumption proves wrong the buyer pays for it. Either the vendor comes back for a change order, or the vendor absorbs the difference and quality drops to fit the price.
TELUS Digital's published guidance on selecting a data annotation company recommends choosing an agency that refrains from quoting a price before it has reviewed your data, because price can vary widely depending on the service or data type. In enterprise procurement, an unseen-data quote is a guess with a decimal point.
A credible quote is based on representative sample data, the actual guidelines, and stated production assumptions you can check. Ask which assumptions the price depends on, and what happens to the price if each is wrong.
How do you score annotation vendor proposals?
Score normalised proposals against weighted criteria agreed before the proposals arrived, with quality and acceptance carrying the largest weight. The weights below are a starting point for an enterprise buyer; adjust them for your risk profile, but agree them before you see the proposals.
| Criterion | Suggested weight | What a strong provider shows |
|---|---|---|
| Quality and acceptance | 30% | Measured QA, adjudication process, written rework policy |
| Unit economics | 20% | Transparent normalised price, no undisclosed fees |
| Scale and throughput | 15% | Proven ramp plan and reviewer capacity at peak |
| Modality fit | 10% | Tools and trained teams for your exact task type |
| Language and expertise | 10% | Named native speakers or domain specialists per requirement |
| Security and governance | 10% | Controls matching your requirements, with scope statements |
| Commercial flexibility | 5% | Reasonable minimums, change terms, exit provisions |
Apply one disqualification rule before scoring: a zero on quality and acceptance is disqualifying regardless of total score. It is the one dimension that cannot be fixed after signature by paying more. Independent AI data validation of a vendor's delivered batches is one way to make the quality score a measurement rather than a claim.
What are the red flags in an annotation quote?
A quote that omits the billable unit, the acceptance rule or the party who pays for rework is signalling that the price depends on something the vendor would rather not state.
- A precise enterprise price arrives without any review of representative data.
- The quote does not define what counts as one billable unit.
- QA or rework is described as "included" with no measurable acceptance rule.
- A low unit rate depends on a minimum commitment you have not modelled.
- The provider cannot separate generalist, specialist and expert labour pricing.
- Platform, storage, integration or training charges appear only after selection.
- The vendor cannot explain how pricing changes when guidelines change.
- Every question about quality is answered with a percentage and no denominator.
How do you run a pilot that makes vendors comparable?
Give every shortlisted vendor an identical pilot and score the results on cost per accepted unit, per-class quality and behaviour on ambiguous items. The proposal round narrows the field; the pilot decides it.
Give each shortlisted provider:
- The same representative sample, including your hardest edge cases and at least one difficult language
- The same guidelines, at the same version
- The same acceptance criteria and gold set
- The same delivery window
Then score cost per accepted unit, per-class quality, escalation behaviour on genuinely ambiguous items, and the quality of questions asked during onboarding. That last signal is undervalued: a vendor who asks six precise questions about your ontology in week one is a vendor who will not silently guess in month six.
Make the pilot long enough to expose representative edge cases and to measure throughput after the initial learning curve, which usually means weeks rather than days. A pilot short enough to be staffed entirely by a vendor's best annotators measures the vendor's best annotators. What happens after the pilot is won is covered in the guide to scaling annotation from pilot to production.
How does Lifewood price and guarantee annotation quality?
Lifewood does not compete on headline unit rate and does not publish one; its public proposition is a defined quality target with commercial consequences attached. Below-threshold batches are reworked at Lifewood's cost, which moves rework from a hidden second invoice to a vendor obligation.
The published terms are a 95%+ accuracy SLA, dual-layer human review with automated consistency checks, and rework of below-threshold batches at Lifewood's cost. For a buyer normalising quotes, that is the term that matters most.
The breadth arguments are about total cost rather than unit price. 100+ languages reduces the need to assemble and manage separate language vendors. Coverage of text, image, audio, video and 3D point-cloud work across Lifewood's AI data services reduces vendor fragmentation and the reconciliation cost between two interpretations of the same ontology. 40+ delivery centres across 30+ countries gives options for distributed production and for regional processing requirements without a second supplier relationship.
A specialist vendor may still win a specific workstream if its pilot demonstrates materially better accepted-unit economics on that task. That is the correct outcome of a well-run comparison, and it is worth designing the process so it can happen.