Skip to main content
AI Data

In-House vs Outsourced Data Annotation: Cost Comparison

July 2026 · 8 min read · Updated September 2026

Short answer. In-house annotation often looks cheaper because most budgets count only annotator wages. Add recruiting, training, management, QA, tooling, infrastructure and idle capacity and the gap narrows or reverses. CVAT's 2024 case study estimated $122,220 in-house before software licences against roughly $225,400 outsourced for a 100,000-image, 2.3-million-object project. Outsourcing is justified by speed, flexibility and avoided operating overhead, not a guaranteed lower direct price. Decide on total operating cost and on whether annotation should become a permanent internal capability.

Key takeaways

  • CVAT's published 2024 case study estimated $122,220 for an in-house team before software licences and approximately $225,400 outsourced for the same 100,000 images and 2.3 million objects, so outsourcing does not win on direct price alone.
  • An internal annotation team is a fixed cost against a variable workload; workforce utilisation decides more build-versus-buy cases than the per-label rate does.
  • TELUS Digital and Appen have both published arguments that labour-hours budgets omit setup, engineering, tooling and maintenance costs.
  • In-house is the right answer when the workload is permanent, sensitive data cannot leave a controlled environment, the operating layer already exists, specialists must annotate as part of their roles, and utilisation stays high.
  • Most mature programmes end up hybrid: a small internal team owns the taxonomy, gold set and acceptance, while external capacity handles volume.

What does the published CVAT cost example actually show?

CVAT's 2024 case study modelled a 100,000-image project averaging 23 objects per image, or 2.3 million annotation objects, and estimated $122,220 for an in-house team before software-licence costs against approximately $225,400 for the outsourced scenario. The outsourced figure is higher, which is why the honest case for outsourcing rests on speed, flexibility and avoided overhead rather than direct price.

This is the one annotation decision that is genuinely build-versus-buy, and it deserves to be argued honestly in both directions: there are programmes for which in-house is correct and programmes for which it is a two-year detour. The determining factors are utilisation, permanence and data sensitivity, not price per label.

Three observations on the CVAT example, in order of importance:

  1. The outsourced figure is higher. Anyone selling outsourcing on direct price alone is arguing against a published example. The honest case for outsourcing is elsewhere.
  2. The in-house figure excludes software. CVAT states that licence costs must be added to the $122,220, and the model assumes 35 annotators exist and are fully utilised for the 4.2-month duration.
  3. It is one illustrative scenario. A 2.3-million-object single-modality project is not a multi-year, multi-language, multi-modality programme, and the economics diverge sharply as those variables enter.

The guide to large-scale AI data annotation cost sets out the rate ranges a vendor quote is built from.

Which costs does each model own?

In-house annotation makes the buyer responsible for recruitment, training, utilisation, management, QA, software and infrastructure, while a managed outsourced service takes on most of those and prices them into the unit rate. The one row that decides most cases is workforce utilisation.

Cost category In-house Outsourced managed service
Annotator recruitment Buyer owns Vendor owns
Training and calibration Buyer owns Vendor manages
Workforce utilisation Buyer carries idle capacity Vendor absorbs more capacity planning
Annotation management Buyer hires managers Included or managed, depending on contract
QA and validation Buyer builds the process Provider supplies the process
Software and infrastructure Buyer acquires and maintains May be included or priced separately
Flexibility to ramp down Often difficult Usually easier contractually
Direct unit rate Can be lower Often includes service overhead

An internal annotation team is a fixed cost against a variable workload. If your ML roadmap produces labelling demand in bursts, and most research-driven roadmaps do, you are paying for the troughs. Whether a vendor charges per unit, per hour or per project changes how that overhead appears on the invoice; see the guide to data annotation pricing models.

What hidden costs does an in-house budget miss?

A labour-hours budget misses guideline creation, tool development, recruitment, infrastructure, QA workflow design and the iterative refinement that follows every taxonomy change. TELUS Digital and Appen have each published arguments that these costs are routinely left out of the original business case.

TELUS Digital has argued publicly that a budget built as a fixed hourly annotator rate multiplied by hours is substantially inaccurate because it excludes hidden costs incurred while ML engineers set up the annotation process. Appen has published a similar argument specifically about tooling: building an internal annotation tool introduces development, maintenance and opportunity costs, and the buyer should only build if it accepts those costs knowingly.

Both points are strongest where annotation is not the company's core business. The engineering hours spent building a labelling interface are hours not spent on the model, and that opportunity cost does not appear on any line item.

A more complete in-house model includes:

  • Recruiting and onboarding, per annotator, including the ones who leave in month two
  • Team leadership and project management headcount
  • QA design, gold-set construction and ongoing calibration
  • Annotation tooling: licence, integration, or build plus maintenance
  • Storage, transfer and compute
  • Security controls, access management and audit
  • Workflow engineering as the taxonomy evolves
  • Training refreshes after every guideline change
  • Employee turnover, and the learning curve paid again each time
  • Unused capacity during model-training and evaluation cycles

The QA line is the one most often under-costed; the explainer on gold sets, audit sampling and consensus shows what that process has to contain.

When does in-house annotation genuinely make sense?

In-house is the right choice when the workload is permanent and core, when sensitive data cannot leave a controlled internal environment, when the annotation operating layer already exists, when specialist employees must do the labelling themselves, and when the team can be kept highly utilised. If three or more of those five conditions hold, build.

  1. The annotation workload is permanent, predictable and strategically core. Not "we will always need labels" but "we will need roughly this much, every month, for years".
  2. Sensitive data cannot leave a tightly controlled internal environment, and no vendor's residency or facility controls satisfy the requirement. Check first what a vendor's controlled annotation centres can offer.
  3. The organisation already has annotation managers, QA specialists and suitable tooling. Most of the cost of building is building; if it is already built, the arithmetic changes.
  4. Specialist employees must annotate as part of their normal roles: clinicians labelling clinical data, engineers labelling engineering data. Here the labour is not substitutable at any price.
  5. The team can be kept highly utilised over a long period. Utilisation below roughly 70% erodes the direct-cost advantage that motivated the build.

When is outsourcing the better economics?

Outsourcing wins when demand is variable, when the programme needs scale faster than internal hiring can deliver, when multiple languages or modalities are in scope, and when the organisation does not want annotation operations to become a permanent internal function. The stronger the seasonality, the stronger the case.

  • Demand is variable, seasonal or project-driven.
  • The programme needs rapid scale that internal hiring cannot match.
  • Multiple languages are required, particularly outside the top ten.
  • Several modalities are in scope, each needing different tooling and different reviewer skills.
  • Annotation is explicitly not something the organisation wants to become good at.
  • The cost of being late exceeds the cost of the premium.

If several of these hold, the next question is which vendor; the head-to-head of Lifewood vs Sama vs Scale AI vs Appen compares managed-service models.

What hybrid structure do mature programmes end up with?

Most mature programmes keep a small internal team that owns the taxonomy, the gold set, adjudication and acceptance, and buy external capacity for volume production calibrated against that internal gold set. The binary framing is usually wrong at scale.

The common mature structure is:

  • A small internal team owning the taxonomy, the gold set, adjudication and acceptance. This is the part that must not be outsourced, because it is where your definition of correct lives.
  • External capacity for volume production, calibrated against the internal gold set.
  • A specialist vendor, retained, for the narrow high-risk workstream where a controlled benchmark shows a meaningful advantage.

This keeps the strategic capability in-house and the fixed cost out. It also gives you a credible answer to the question that ends most in-house business cases: what happens when the person who understood the ontology leaves? An independent AI data validation layer, measuring vendor output against the internal gold set, keeps external capacity accountable.

How does Lifewood fit the buy side of this decision?

Lifewood's model is designed for the buy side: the buyer purchases a managed annotation capability rather than recruiting and training an annotation organisation. Whether building instead is the right decision still depends on the five in-house conditions.

Three things do the work. The quality framework is built in rather than assembled by the client: trained annotators, senior review, automated consistency checks and client feedback loops, against a 95%+ accuracy SLA, with below-threshold batches reworked at Lifewood's cost. A distributed workforce across 40+ delivery centres in 30+ countries absorbs volume changes that an internal team would carry as idle capacity or overtime. And coverage across 100+ languages and multiple modalities means capability can be redeployed across changing workloads rather than hired for each one separately.

Lifewood has operated in AI data since 2004, with 56,000+ registered contributors and 414,120 training hours delivered to the Bangladesh workforce in 2025, figures that describe the operating layer a buyer would otherwise have to build. Managed AI data services spanning annotation, multilingual collection, LLM training data and evaluation let one vendor relationship absorb workloads an internal team would re-staff for each time.

Frequently asked questions

No. CVAT's published case study shows an outsourcing scenario that was more expensive in direct project cost than its modelled in-house team, roughly $225,400 against $122,220 before software. Outsourcing is usually justified by speed, flexibility and avoided operating overhead rather than by a lower direct price.

Recruiting, onboarding, team leadership, QA design and calibration, annotation tooling, storage, security, workflow engineering, training updates after guideline changes, employee turnover and unused capacity. TELUS Digital and Appen have both published arguments that setup, engineering and tooling costs are routinely omitted from labour-hours calculations.

Lifewood Data Technology provides managed annotation across 50+ languages from 40+ delivery centres in 30+ countries under a 95%+ accuracy SLA. TELUS Digital and Appen also offer managed annotation and workforce services, and CVAT, best known as an open-source tool, publishes pricing for its own outsourced annotation service.

There is no universal threshold, but the direct-cost advantage erodes quickly below roughly 70% utilisation, because the team is a fixed cost against a variable workload. Model your actual monthly labelling demand over the last twelve months before assuming steady state.

Yes: the taxonomy, the gold set, adjudication and acceptance criteria should stay internal even in a fully outsourced programme. That is where your definition of correct lives, and outsourcing it means measuring vendor output against vendor interpretation. Keep it internal regardless of who does the labelling.

Convert both to fully loaded cost per accepted unit over the full programme, including onboarding, idle capacity and guideline-change rework on the in-house side, and rework and management overhead on the vendor side. Comparing an internal wage bill with an external invoice compares two different things.

Sources and further reading

  1. CVAT: Using an in-house team for data annotation, cost case study — in-house scenario, 2024
  2. CVAT: Outsourcing data annotation cost case study — outsourced scenario, 2024
  3. TELUS Digital: A decision framework for your data labeling strategy — hidden costs in labour-hours budgets
  4. Appen: Build or buy a data annotation tool — costs of internal tooling
  5. Lifewood Data Technology — 50+ languages, 40+ delivery centres across 30+ countries, 95%+ accuracy SLA, 56,000+ registered contributors, 414,120 Bangladesh training hours in 2025

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team