Short answer. Almost everything except the word "annotation". A frontier lab is buying judgement it cannot generate internally — PhD-level reasoning tasks, preference rankings, adversarial probes — and will pay accordingly, with domain expert evaluators commanding $130 to $1,000 an hour. An enterprise team is usually buying a defensible, auditable dataset for a system that has to pass a regulator, which since 2 August 2026 means documented governance over collection, origin, preparation, labelling and quality assurance under EU AI Act Article 10. One buyer optimises for the ceiling of quality. The other optimises for the floor of risk. Serving both well means running two different operations under one roof.
How big is the gap between these two buyers?
Wider than most people outside the industry assume — and it starts with the budgets.
Each frontier AI lab now spends on the order of a billion dollars a year on human-generated training data, a figure attributed to Labelbox CEO Manu Sharma and reported in Time's 2025 investigation. Set that against the whole data collection and labelling market, valued at $3.77 billion in 2024 and forecast to reach $17.10 billion by 2030 at a 28.4% CAGR, and the shape of the market becomes clear: a handful of buyers account for an outsized share of the spend, and they are buying something quite specific.
That spending shows up as an extraordinary spread in what different work is worth. The taxonomy in circulation among recruiters puts six distinct roles in the same industry, and the gap between the tiers runs to 50× or more.
What the top tier actually produces looks less like labelling and more like publishing. Turing shipped a pack of 1,106 expert-authored PhD-level reasoning tasks across computer science, data science and chemistry — the sort of artefact only a credentialed specialist who also thinks like a test designer can create. As one recruiting analysis put it, at that tier you are not hiring a labeller at all, you are hiring a curriculum designer with a graduate degree.
Meanwhile the labs keep the strategy internal and rent the execution. Anthropic has advertised a Data Operations Manager, Human Data at $270,000 to $365,000 to own strategy across RLHF and safety while orchestrating outside vendors — a pattern one analyst summarised neatly as "keep the brain in-house, rent the hands". Roughly 69% of labelling work still runs through outsourced annotation platforms.
How does task design change?
Enterprise work asks annotators to apply a schema. Frontier work asks them to have an opinion, and then defends the process that produced it.
An enterprise annotation task usually has a right answer that exists before the task does. Is this invoice line a freight charge? Does this scan show the defect? The schema is knowable, the guideline can be written in advance, and the job is consistent application at volume. That is a genuinely hard operational problem, but it is a bounded one.
Frontier tasks frequently have no pre-existing right answer. Which of these two model responses is better, and why? Where does this reasoning chain go wrong? Can you construct a prompt that breaks the safety policy? The output is a judgement, and its value comes from the credibility of the person making it. This is why the market for this work sits with expert marketplaces rather than general annotation platforms, and why the assessment, sourcing and pay all change together.
We see this split cleanly in our own pipeline at Lifewood. A multilingual enterprise programme and a frontier evaluation programme might both be described in a brief as "text annotation", but almost nothing transfers between them — not the recruiting profile, not the guideline structure, not the tooling, not the unit economics. Treating them as one service is the fastest way to disappoint both clients.
The same word, two different operations DIMENSION ENTERPRISE TEAM FRONTIER-MODEL LAB WHAT THEY BUY Consistent application of a known schema at volume Credentialed judgement that does not exist inside the company TASK SHAPE Bounded: a right answer exists before the task is written Open: preference, critique, adversarial construction, curriculum design WORKFORCE Trained generalists plus domain specialists where needed MS and PhD specialists; oncology, aerospace, law, chemistry SUCCESS METRIC Agreement, throughput, cost per unit, audit readiness Does the resulting model behave better on held-out evaluations DOMINANT RISK Regulatory exposure and downstream business error Noisy preference data silently degrading alignment CONTRACT FOCUS Provenance, indemnity, audit rights, DPA terms Exclusivity, non-compete across labs, IP in the tasks themselves Both are legitimate, sophisticated buyers. They are simply solving different problems with the same vocabulary.
How does QA depth change?
Enterprise QA proves consistency. Frontier QA fights noise that consistency checks cannot see.
On an enterprise programme, agreement metrics do most of the work. Two annotators disagreeing signals guideline ambiguity, and the fix is a clearer rule. On a frontier preference task, two annotators disagreeing may simply mean the question is genuinely contested — and averaging their answers can destroy exactly the signal the lab is paying for.
The stakes are quantified: noise in RLHF preference annotations commonly exceeds 20% in real datasets, and that noise significantly degrades alignment performance. You cannot inspect your way out of that with a spot-check rate. It has to be handled through who does the task, how the task is framed, and how disagreement is adjudicated rather than smoothed.
The market has responded by splitting into layers. Beneath the expert marketplaces sits the bulk annotation workforce that handles the repetitive data the frontier players no longer touch — and that layer still matters to labs, particularly in computer vision and multilingual data. Programmatic approaches occupy a third position: Snorkel AI, spun out of the Stanford AI Lab in 2019, uses code and expert-written heuristics to generate and de-noise labels at scale rather than bruteforce human clicking, raised $100 million at a $1.3 billion valuation in May 2025, and has layered an expert network of MS and PhD specialists on top.
ENTERPRISE QA FRONTIER QA
Inter-annotator agreement as the headline metric
Annotator credibility screened before the work starts
Gold sets, honeypots, defined spot-check rates
Disagreement treated as a guideline defect
Disagreement adjudicated by seniority, not averaged away
Documented, repeatable, evidence-producing
Held-out evaluation of the trained model as ground truth
Optimised for cost per unit at a quality floor
Adversarial review: can the label be defended The output has to survive an audit two years from now.
Optimised for signal quality, largely regardless of cost RLHF preference noise commonly exceeds 20% and degrades alignment.
How do data-rights requirements change?
This is where enterprise buyers become the demanding ones — and where the regulatory clock has already run out.
EU AI Act obligations for high-risk systems became enforceable on 2 August 2026, and Article 10 is unusually specific about annotation. It requires documented data governance covering collection, origin, preparation, labelling and quality assurance; datasets that are relevant, sufficiently representative and as free of errors as possible; and examination for biases that could lead to discriminatory outcomes. Article 30 requires technical documentation detailing the data used for training. Non-compliance can reach €15 million or 3% of global annual turnover.
The line that matters most for anyone buying annotation: you cannot outsource this obligation to your vendor. Guidance to in-house counsel now recommends requiring vendors to contractually warrant AI Act compliance and indemnify deployers for vendor compliance failures, alongside training-data source disclosure, third-party audit rights, model-update notification and DPA terms. That advice exists because AI vendors have been observed carving AI-generated outputs out of IP indemnity, reserving rights to train on customer prompts, and capping liability well below the cost of a copyright suit.
And provenance is a live legal risk rather than a theoretical one, with active multidistrict litigation consolidating copyright claims against AI developers. Which is why, on the enterprise side, the questions we field most often at Lifewood are not about throughput at all. They are about where the data came from, who touched it, in which jurisdiction, under what consent, and whether we can produce that record on demand three years later. For a frontier lab those questions matter too — but they arrive after the conversation about whether we can find forty native-speaking specialists with the right credentials by the end of the month.
Name your buyer type before you write the brief. Nearly every downstream decision — recruiting, guidelines, QA, pricing — follows from that one classification.
For frontier work, screen credibility before task design. The output is only as good as the person's standing to make the judgement.
Do not average away disagreement on subjective tasks. Adjudicate it. Averaging destroys the signal you paid a premium to obtain.
For enterprise work, build the audit trail as you go. Article 10 wants collection, origin, preparation, labelling and QA documented — reconstructing that later is far more expensive.
Get provenance and indemnity into the contract, not the SOW. Ask specifically whether AI-generated outputs are carved out of IP indemnity.
Match the QA instrument to the task type. Agreement metrics for bounded schemas; expert adjudication and held-out model evaluation for open judgement.
Keep the layers separate operationally. Bulk multilingual volume and expert evaluation need different teams, tools and unit economics even inside one vendor.
Assume both buyers will eventually want both. Labs need multilingual bulk data; enterprises are starting to need expert evaluation. Build for the overlap.
Key takeaways
- Each frontier lab spends on the order of $1 billion a year on human training data, against a total 2024 market of $3.77 billion.
- Pay spans six roles from $15–25/hr for data annotators to $130–1,000/hr for domain expert evaluators — a gap of 50× or more.
- Frontier output looks like curriculum design: Turing shipped 1,106 expert-authored PhD-level reasoning tasks across three disciplines.
- Labs keep strategy in-house and rent execution; roughly 69% of labelling work runs through outsourced platforms.
- Enterprise tasks have a right answer before the task exists; frontier tasks often do not, which changes recruiting, guidelines and QA together.
- RLHF preference noise commonly exceeds 20% in real datasets and significantly degrades alignment — a problem inspection alone cannot solve.
- EU AI Act Article 10 became enforceable on 2 August 2026 and explicitly covers labelling within its data-governance requirements.
- You cannot outsource the Article 10 obligation to a vendor; contracts should warrant compliance, disclose provenance and grant audit rights.
- Penalties for high-risk non-compliance reach €15 million or 3% of global annual turnover.
Sources and further reading
- Pin, "How AI Labs Are Hiring People to Train Models 2026" — the six-role taxonomy and pay bands (citing HireArt's 2025 AI compensation survey and Built In), the ~$1B per lab annual human-data spend (citing Time, 2025), the 69% outsourced-platform share, and Grand View Research market figures of $3.77B (2024) to $17.10B (2030) at 28.4% CAGR (data for Figures 1 and 2)
- HeroHunt.ai, "Data Annotation for AI Labs: Recruiting Guide 2026" — on the ~$1B figure attributed to Labelbox CEO Manu Sharma via Cognitive Revolution, Anthropic's Data Operations Manager, Human Data role at $270,000–365,000, the "keep the brain in-house, rent the hands" pattern, and Turing's 1,106 expert-authored PhD-level reasoning tasks (via TechCrunch). herohunt.ai/blog/data-annotation-ai-labs-the-recruiting-guide-2026 HeroHunt.ai, "Top 10 Data Annotators for AI Labs (2026 Benchmark)" — on expert marketplace rates above $100, the bulk annotation layer beneath them, and Snorkel AI's Stanford AI Lab origins, programmatic labelling approach, $100M raise at a $1.3B valuation in May 2025, and Expert Data-as-a-Service network. herohunt.ai/blog/top-10-data-annotators-for-ai-labs-2026-benchmark Lightly AI, "5 Best Data Annotation Companies in 2026" — on 2026 pricing benchmarks ($0.02–0.09 per bounding box, $6–12/hr managed services, $50–100 per example for expert RLHF or medical annotation) and foundation models pushing human effort toward edge cases and regulated domains
- Kili Technology, "Data Annotation Guide: How to Achieve High Quality in Complex AI Data Operations" (2026) — on RLHF preference-annotation noise commonly exceeding 20% in real datasets and degrading alignment performance. kili-technology.com/blog/data-annotation-guide-how-to-achieve-high-quality-data-in-complex-ai-data-operations Ertas AI, "EU AI Act Training Data Compliance: The Complete Guide (2026)" — on Article 10 data governance covering collection, origin, preparation, labeling and quality assurance; data quality and bias-examination criteria; Article 30 technical documentation; and the phased enforcement timeline to August 2026
- NeuralTrust, "Data Sovereignty Requirements under the EU AI Act (2026)" — on Article 10 obligations becoming enforceable 2 August 2026, the point that the obligation cannot be outsourced to a vendor, GDPR Chapter V interaction, and penalties of €15 million or 3% of global annual turnover
- Promise Legal, "AI Vendor Contract Requirements: A 2026 Checklist" — on IP indemnity carve-outs for AI-generated outputs, training-data source disclosure, third-party audit rights, model-update notification, and active multidistrict copyright litigation making provenance a live legal risk. blog.promise.legal/ai-vendor-contract-requirements-due-diligence The Data Governor, "EU AI Act Data Governance: 2026 Compliance Guide" — on requiring vendors to contractually warrant AI Act compliance and indemnify deployers for vendor compliance failures. thedatagovernor.com/eu-ai-act-data-governance-requirements Daeryun Law, "AI Licensing: Training Data Rights and Commercial Use Agreements" — on GPAI obligations under Article 53 including trainingdata summary disclosure, and the EU AI Act's four-tier risk classification. daeryunlaw.com/us/practices/detail/ai-licensing Lifewood, AI data, annotation and evaluation services
- Charts in Figures 1 and 2 were produced by Lifewood from the figures reported in the sources cited beneath each chart.