Skip to main content
AI Data

Annotation for Frontier-Model Labs vs Enterprise Teams

September 2026 · 13 min read · Updated September 2026

Short answer. Almost everything except the word "annotation". A frontier lab is buying judgement it cannot generate internally — PhD-level reasoning tasks, preference rankings, adversarial probes — and pays accordingly, with domain expert evaluators commanding $130 to $1,000 an hour. An enterprise team is usually buying a defensible, auditable dataset that must satisfy EU AI Act Article 10's documented data governance over collection, origin, annotation and labelling. One optimises for the ceiling of quality, the other for the floor of risk.

Key takeaways

  • Each frontier AI lab spends on the order of $1 billion a year on human training data, against a total data collection and labelling market of $3.77 billion in 2024.
  • Pay spans six roles, from $15–25 an hour for data annotators to $130–1,000 an hour for domain expert evaluators, a gap of 50× or more.
  • Enterprise tasks have a right answer before the task exists; frontier tasks often do not, which changes recruiting, guidelines and QA together.
  • Noise in RLHF preference annotations commonly exceeds 20% in real datasets and significantly degrades alignment, a problem inspection alone cannot solve.
  • EU AI Act Article 10 names annotation and labelling in its data-governance duties; after the 2026 Digital Omnibus, stand-alone high-risk systems must comply from 2 December 2027, with fines of up to €15 million or 3% of global annual turnover.

How big is the gap between frontier labs and enterprise buyers?

Wider than most people outside the industry assume, and it starts with the budgets. A handful of frontier labs each spend on the order of a billion dollars a year on human training data, in a market worth $3.77 billion in total in 2024.

Labelbox CEO Manu Sharma has said that each frontier lab is probably spending over $1 billion a year on data, a figure also reported in Time's 2025 investigation of the data industry. Set that against the whole data collection and labelling market, which Grand View Research values at $3.77 billion in 2024 and forecasts to reach $17.10 billion by 2030 at a 28.4% CAGR, and the shape of the market becomes clear: a handful of buyers account for an outsized share of the spend, and they are buying something quite specific. The vendors serving that spend differ from the providers most enterprises shortlist, which is why our list of the best data annotation companies for LLM training separates expert-evaluation specialists from volume annotators.

That spending shows up as an extraordinary spread in what different work is worth. The taxonomy in circulation among recruiters puts six distinct roles in the same industry — data annotators at $15–25 an hour, AI tutors and trainers at $20–55, RLHF specialists at $50–65, prompt engineers at $40–65, red teamers at $100–200 and domain expert evaluators at $130–1,000 — and the gap between the bottom and top tiers runs to 50× or more.

What the top tier actually produces looks less like labelling and more like publishing. Turing sells expert-authored, PhD-level STEM reasoning data packs — including a set of 201 expert-authored, verifier-graded agentic tasks — written and adjudicated by PhD and PhD-candidate researchers in physics, chemistry, biology, mathematics and materials science. That is the sort of artefact only a credentialed specialist who also thinks like a test designer can create. As one recruiting analysis put it, at that tier you are not hiring a labeller at all, you are hiring a curriculum designer with a graduate degree.

Meanwhile the labs keep the strategy internal and rent the execution. Anthropic has advertised a Data Operations Manager, Human Data role at $270,000 to $365,000 to own data strategy across RLHF, safety, tool use and agentic workflows while driving strategic vendor partnerships — a pattern one analyst summarised neatly as "keep the brain in-house, rent the hands". Outsourced providers now deliver 69% of all data labelling work, according to Mordor Intelligence figures cited in the same recruiting guide.

How does task design change between the two buyers?

Enterprise work asks annotators to apply a schema. Frontier work asks them to have an opinion, and then defends the process that produced it.

An enterprise annotation task usually has a right answer that exists before the task does. Is this invoice line a freight charge? Does this scan show the defect? The schema is knowable, the guideline can be written in advance, and the job is consistent application at volume. That is a genuinely hard operational problem, but it is a bounded one.

Frontier tasks frequently have no pre-existing right answer. Which of these two model responses is better, and why? Where does this reasoning chain go wrong? Can you construct a prompt that breaks the safety policy? The output is a judgement, and its value comes from the credibility of the person making it. This is why the market for this work sits with expert marketplaces rather than general annotation platforms, and why the assessment, sourcing and pay all change together. Writing a rubric that raters can apply consistently is its own discipline, covered in our guide to preference rubrics raters agree on.

We see this split cleanly in our own pipeline at Lifewood. A multilingual enterprise programme and a frontier evaluation programme might both be described in a brief as "text annotation", but almost nothing transfers between them — not the recruiting profile, not the guideline structure, not the tooling, not the unit economics. Treating them as one service is the fastest way to disappoint both clients.

The same word, two different operations

Dimension Enterprise team Frontier-model lab
What they buy Consistent application of a known schema at volume Credentialed judgement that does not exist inside the company
Task shape Bounded: a right answer exists before the task is written Open: preference, critique, adversarial construction, curriculum design
Workforce Trained generalists plus domain specialists where needed MS and PhD specialists; oncology, aerospace, law, chemistry
Success metric Agreement, throughput, cost per unit, audit readiness Whether the resulting model behaves better on held-out evaluations
Dominant risk Regulatory exposure and downstream business error Noisy preference data silently degrading alignment
Contract focus Provenance, indemnity, audit rights, DPA terms Exclusivity, non-compete across labs, IP in the tasks themselves
Typical price signal $6–12 an hour for managed annotation services; $0.02–0.09 per bounding box $50–100 per example for expert RLHF or medical annotation; $130–1,000 an hour for domain expert evaluators

Both are legitimate, sophisticated buyers. They are simply solving different problems with the same vocabulary.

How does QA depth change?

Enterprise QA proves consistency. Frontier QA fights noise that consistency checks cannot see.

On an enterprise programme, agreement metrics do most of the work. Two annotators disagreeing signals guideline ambiguity, and the fix is a clearer rule. Lifewood runs enterprise programmes against a 95%+ accuracy SLA, with a 95%+ inter-annotator agreement threshold measured against a customer-approved gold set and two independent review passes with timestamped approval records — the instruments described in our explainer on gold sets, audit sampling and consensus. On a frontier preference task, two annotators disagreeing may simply mean the question is genuinely contested, and averaging their answers can destroy exactly the signal the lab is paying for.

The stakes are quantified: noise in RLHF preference annotations commonly exceeds 20% in real datasets, and that noise significantly degrades alignment performance. You cannot inspect your way out of that with a spot-check rate. It has to be handled through who does the task, how the task is framed, and how disagreement is adjudicated rather than smoothed — the drift problem we describe in RLHF preference ratings at scale.

The market has responded by splitting into layers. Beneath the expert marketplaces — where one 2026 benchmark puts average pay at Mercor around $85 an hour and expert pay at Handshake AI at $100 to $125 — sits the bulk annotation workforce that handles the repetitive data the frontier players no longer touch, and that layer still matters to labs, particularly in computer vision and multilingual data. Programmatic approaches occupy a third position: Snorkel AI, spun out of the Stanford AI Lab in 2019, uses code and expert-written heuristics to generate and de-noise labels at scale rather than brute-force human clicking, raised a $100 million Series D at a $1.3 billion valuation in May 2025, and has layered an Expert Data-as-a-Service network of subject-matter specialists on top. Independent AI data validation sits across all three layers, checking delivered labels against a held-out gold standard before they reach training.

Enterprise QA and frontier QA compared

QA element Enterprise QA Frontier QA
Headline metric Inter-annotator agreement Held-out evaluation of the trained model as ground truth
Screening Gold sets, honeypots, defined spot-check rates Annotator credibility screened before the work starts
Disagreement Treated as a guideline defect Adjudicated by seniority, not averaged away
Evidence Documented, repeatable, evidence-producing; the output has to survive an audit two years from now Adversarial review: can the label be defended
Optimised for Cost per unit at a quality floor Signal quality, largely regardless of cost
Failure mode Inconsistent labels reaching a regulated system RLHF preference noise above 20% degrading alignment

How do data-rights requirements change?

This is where enterprise buyers become the demanding ones, because the regulatory obligation sits with the provider of the high-risk system and cannot be delegated to an annotation vendor.

EU AI Act Article 10 is unusually specific about annotation. It requires data governance practices covering data collection processes and the origin of data, and data-preparation operations such as annotation, labelling, cleaning, updating, enrichment and aggregation; training, validation and testing datasets that are relevant, sufficiently representative and, to the best extent possible, free of errors and complete; and examination for biases likely to affect health, safety or fundamental rights or lead to discrimination. Article 11 and Annex IV require technical documentation with detailed information about the data used for training. Non-compliance with high-risk obligations can reach €15 million or 3% of total worldwide annual turnover, whichever is higher.

The regulatory clock has moved, but it has not stopped. The Digital Omnibus on AI, agreed in May 2026, deferred the high-risk deadline from 2 August 2026 to 2 December 2027 for stand-alone Annex III systems and to 2 August 2028 for AI embedded in Annex I regulated products. Article 50 transparency obligations kept their original August 2026 date. For an enterprise buying annotation now for a system that will still be in production in 2028, the documentation requirement is unchanged; only the date it is tested has shifted.

The line that matters most for anyone buying annotation: you cannot outsource this obligation to your vendor. The obligation sits with the provider of the high-risk AI system, which in most enterprise deployments is the enterprise itself. Guidance to in-house counsel now recommends requiring vendors to contractually warrant AI Act compliance and indemnify deployers for vendor compliance failures, alongside disclosure of training-data source categories, third-party audit rights for high-stakes use cases, 30–60 days' notice of material model updates and DPA terms. That advice exists because AI vendors have been observed carving AI-generated outputs out of IP indemnity, reserving rights to train on customer prompts, and capping liability at amounts that do not begin to cover a copyright suit. The security and residency controls that back those warranties are set out in our guide to enterprise annotation security and compliance.

And provenance is a live legal risk rather than a theoretical one, with active multidistrict litigation consolidating copyright claims against AI developers. Which is why, on the enterprise side, the questions we field most often at Lifewood are not about throughput at all. They are about where the data came from, who touched it, in which jurisdiction, under what consent, and whether we can produce that record on demand three years later. Our enterprise LLM training data programmes are built to answer those questions as a matter of course. For a frontier lab those questions matter too, but they arrive after the conversation about whether we can find forty native-speaking specialists with the right credentials by the end of the month.

What should you do differently for each buyer type?

Name your buyer type before you write the brief, because nearly every downstream decision follows from that one classification.

  • Name your buyer type before you write the brief. Recruiting, guidelines, QA and pricing all follow from whether the work is bounded schema application or open expert judgement.
  • For frontier work, screen credibility before task design. The output is only as good as the person's standing to make the judgement; sourcing MS and PhD specialists is a recruiting problem, not a tooling one, as our guide to domain-expert SFT datasets explains.
  • Do not average away disagreement on subjective tasks. Adjudicate it. Averaging destroys the signal you paid a premium to obtain.
  • For enterprise work, build the audit trail as you go. Article 10 wants collection, origin, annotation, labelling and quality processes documented; reconstructing that later is far more expensive.
  • Get provenance and indemnity into the contract, not the SOW. Ask specifically whether AI-generated outputs are carved out of IP indemnity.
  • Match the QA instrument to the task type. Agreement metrics for bounded schemas; expert adjudication and held-out model evaluation for open judgement.
  • Keep the layers separate operationally. Bulk multilingual volume and expert evaluation need different teams, tools and unit economics even inside one vendor.
  • Assume both buyers will eventually want both. Labs need multilingual bulk data; enterprises are starting to need expert evaluation. Build for the overlap.

Frequently asked questions

Yes, but only by running them as separate operations. The recruiting profiles, guideline structures, QA instruments and unit economics have almost nothing in common, and blending them tends to produce work that is over-engineered for one client and under-specified for the other. Lifewood runs frontier evaluation and multilingual enterprise programmes as distinct pipelines.

No. Foundation models now handle much routine pre-labelling, which pushes human effort toward edge cases, subjective judgement and regulated domains, but substantial volume work remains, particularly in computer vision and multilingual data. Outsourced providers still deliver 69% of all data labelling work, and commodity rates have roughly halved since 2022.

After the Digital Omnibus on AI agreed in May 2026, stand-alone high-risk systems under Annex III must comply from 2 December 2027, and AI embedded in Annex I regulated products from 2 August 2028. Article 10 explicitly lists annotation and labelling among the data-preparation operations that must be governed and documented.

It applies based on where the system is placed on the market or used, not where you are incorporated. Enterprises serving EU users generally fall in scope, and deployer obligations sit alongside GDPR rather than replacing them; sending EU personal data to a non-EU provider triggers GDPR Chapter V transfer rules as well.

Three layers serve different buyers. Expert marketplaces such as Mercor and Handshake AI supply credentialed evaluators for frontier labs; programmatic platforms such as Snorkel AI combine code-generated labels with an expert network; and managed workforces such as Lifewood, with 56,788 registered contributors across 40+ delivery centres in 30+ countries, handle enterprise volume in 50+ languages.

Which of the two operations you are being sold. Then, depending on the answer: for frontier work, how they source and verify credentials and adjudicate disagreement; for enterprise work, what provenance, consent and QA documentation they produce as a matter of course rather than on request, and whether they will warrant it contractually.

Sources and further reading

  1. Pin, How AI Labs Are Hiring People to Train Models 2026 — six-role pay bands, the 69% outsourced share and the Time-reported per-lab spend
  2. Cognitive Revolution, The Data Factory: Inside the $100B Race for Post-Training Supremacy, with Labelbox CEO Manu Sharma — each frontier lab probably spending over $1 billion a year on data
  3. Grand View Research, Data Collection And Labeling Market To Reach $17.10Bn By 2030 — $3.77 billion (2024) to $17.10 billion (2030) at a 28.4% CAGR
  4. HeroHunt.ai, Data Annotation for AI Labs: Recruiting Guide 2026 — the "keep the brain in-house, rent the hands" pattern and the "curriculum designer with a graduate degree" observation
  5. Anthropic, Data Operations Manager, Human Data (job listing) — $270,000–365,000 range; strategy across RLHF and safety; vendor partnerships
  6. Turing, Frontier STEM Data Packs — PhD-level STEM reasoning data packs, including 201 expert-authored, verifier-graded agentic tasks
  7. HeroHunt.ai, Top 10 Data Annotators for AI Labs (2026 Benchmark) — Mercor and Handshake AI expert pay, halved commodity rates, and Snorkel AI's Stanford AI Lab origin
  8. BigDATAwire, Snorkel AI Announces $100M Series D and Expanded Platform to Power Next Phase of AI with Expert Data — $100 million Series D at a $1.3 billion valuation, May 2025
  9. Lightly AI, 5 Best Data Annotation Companies in 2026 — pricing benchmarks and foundation models pushing human effort toward edge cases
  10. Kili Technology, Data Annotation Guide: How to Achieve High Quality in Complex AI Data Operations — RLHF preference noise commonly exceeding 20%
  11. EU Artificial Intelligence Act, Article 10: Data and Data Governance — data-governance duties and the 2 December 2027 / 2 August 2028 application dates
  12. EU Artificial Intelligence Act, Article 99: Penalties — fines of up to EUR 15 million or 3% of worldwide annual turnover
  13. Gibson Dunn, EU AI Act Omnibus Agreement — Postponed High-Risk Deadlines and Other Key Changes — deferral of Annex III obligations to 2 December 2027 and Annex I to 2 August 2028
  14. Cloud Security Alliance, EU AI Act High-Risk Deadline Pushed to December 2027 — Article 50 transparency obligations remaining on the 2 August 2026 timeline
  15. NeuralTrust, Data Sovereignty Requirements under the EU AI Act — the Article 10 obligation sitting with the provider rather than the vendor, and GDPR Chapter V interaction
  16. Ertas AI, EU AI Act Training Data Compliance: The Complete Guide (2026) — Article 11 / Annex IV technical documentation detailing the data used for training
  17. Promise Legal, AI Vendor Contract Requirements: A 2026 Checklist — indemnity carve-outs, disclosure, audit rights, model-update notice and multidistrict copyright litigation
  18. The Data Governor, EU AI Act Data Governance: 2026 Compliance Guide — requiring vendors to contractually warrant AI Act compliance and indemnify deployers for vendor compliance failures

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team