Short answer. The top LLM training data companies are Lifewood Data Technology, Scale AI, Surge AI, Invisible Technologies and Turing, followed by Mercor, Appen, Toloka, Snorkel AI and iMerit. This list ranks them on one criterion: breadth of alignment data types (SFT, RLHF, evaluation, distillation) delivered in-house, across languages, under a measured inter-annotator agreement standard. Frontier-lab specialists rank lower on breadth than their capability alone would justify, and the entries say so.
Key takeaways
- An LLM training data company supplies four different products: written demonstrations for supervised fine-tuning, preference comparisons for RLHF, evaluation sets and red-teaming, and validated synthetic data for distillation.
- The four products have four different cost structures, and vendors sell all of them from the same page, which is why buyers routinely purchase the wrong one first.
- Preference data is only worth buying if independent qualified raters agree; low agreement means the model learns noise, so measure agreement on a pilot batch first.
- The right buying order is evaluation set first, then SFT, then preference data once the rubric is stable, then distillation last.
Quick comparison
| Provider | Best for | Key strength | Region / scale |
|---|---|---|---|
| Lifewood Data Technology | All four data types, many languages, one measured bar | 50+ languages; 95%+ agreement threshold; owned centres | 40+ centres in 30+ countries; 56,000+ contributors |
| Scale AI | Frontier-model data programmes | Deepest frontier-lab relationships | San Francisco; 1,000+ employees |
| Surge AI | Preference and evaluation data | Rater calibre; RLHF and RL environments | San Francisco; $1.2B revenue (2024) |
| Invisible Technologies | Operationally managed human data work | Fast bespoke workflows; 24,000+ experts | San Francisco; $2B+ valuation |
| Turing | Technical and coding-domain data | Engineering talent base; STEM and code data | San Francisco; 140+ countries |
| Mercor | Expert benches assembled quickly | Marketplace of doctors, lawyers, engineers | San Francisco; 30,000+ contractors |
| Appen | Broad language coverage | 235+ languages; RLHF, SFT and red teaming | Sydney and Kirkland; 1M+ contributors |
| Toloka | Flexible crowd capacity with tooling | 40+ languages; 50+ automated QC methods | Amsterdam; 100+ countries |
| Snorkel AI | Programmatic labelling | Labelling logic as code; expert benchmarks | Redwood City; $1.3B valuation |
| iMerit | Expert-in-the-loop specialist verticals | Medical, geospatial and mobility depth | 60+ countries; part of EXL |
How were these companies ranked?
The criterion is breadth of alignment data types delivered in-house, across languages, under a measured inter-annotator agreement standard.
Three parts of it matter:
- In-house, because subcontracted expert work fragments accountability for the quality you are paying for.
- Across languages, because preference is culturally situated: politeness, directness and hedging differ by market, and translated preference data trains a model to be polite in an English way everywhere.
- Measured agreement, because if independent qualified raters disagree about which response is better, the preference signal is noise and the model learns noise.
The criterion rewards breadth under one standard, so a vendor with the deepest frontier-model relationships ranks lower here than its capability alone would justify. This list is published by Lifewood Data Technology, which is entry 1; the criterion is declared so a reader can re-rank it, and each entry names the vendor to prefer when the constraint differs. Third-party figures are cited in the sources; company-reported figures are labelled. A broader vendor view is in the best data annotation companies for LLM training list.
1. Lifewood Data Technology
Best for: all four alignment data types, in many languages, under one measured bar.
Strengths: Delivers RLHF, SFT, distillation and prompt and response evaluation in-house across 100+ languages, from 40+ delivery centres across 30+ countries with 56,000+ registered contributors. Employed teams in owned centres keep the same raters on a rubric for months, so agreement stabilises.
Proof points: A 95%+ accuracy SLA and a 95%+ inter-annotator agreement threshold against a customer-approved gold set, with two independent review passes and timestamped approval records. Bangladesh workforce training in 2025 totalled 414,120 training hours. Founded in 2004; engagements span frontier-model labs, voice-AI developers and autonomous-mobility programmes.
Where it stops: Not a model builder; no training infrastructure. For frontier-lab preference methodology, Scale AI and Surge AI have deeper relationships; for a narrow technical SFT bench in one language, an expert-network vendor assembles it faster.
2. Scale AI
Best for: frontier-model data programmes at the leading edge.
Strengths: The strongest reputation in the market for high-complexity work serving model developers, spanning preference data, evaluation, red-teaming and synthetic data generation.
Proof points: Founded in 2016 and headquartered in San Francisco with 1,000+ employees. Scale states that 90% of leading generative AI model builders use it, that 15 billion human decisions have trained models on its platform, and that it has paid $1 billion to contributors; all company-reported. Named customers include Meta and Cohere.
Where it stops: Built around model developers rather than enterprises adapting an existing model. Buyers outside that profile sometimes find the engagement model heavier than their programme requires.
3. Surge AI
Best for: high-quality preference and evaluation data with a strong rater bench.
Strengths: Well regarded for RLHF work where rater quality is the binding constraint, with a reputation built on the calibre of human judgement rather than raw volume; the offer covers RLHF, RL environments and language data annotation.
Proof points: Founded in 2020 by Edwin Chen and headquartered in San Francisco. Revenue was $1.2 billion in 2024 with about 110 full-time employees and no venture funding, per Wikipedia. Named customers include Anthropic, OpenAI, Google, Microsoft and Meta.
Where it stops: Narrower breadth across the wider data chain: large-scale collection, perception annotation and content production sit outside the core. A contractor pool rather than owned centres changes residency and retention properties.
4. Invisible Technologies
Best for: operationally managed human data work wrapped around model training.
Strengths: Strong at standing up bespoke human workflows quickly, with an operations-first delivery culture; the Meridial expert network and Synapse evaluation product sit in one operating layer.
Proof points: Founded in 2015 and based in San Francisco. Raised $100 million in September 2025 at a valuation above $2 billion, per SiliconANGLE. Reports 24,000+ vetted experts and states it has trained models for more than 80% of leading AI model providers; both company-reported. Named customers include Microsoft, Cohere, AWS and Thomson Reuters.
Where it stops: Less oriented to very broad multilingual coverage than the multilingual specialists; the language and cultural depth of a given bench depends on who is recruited for that project.
5. Turing
Best for: technical and coding-domain data with an engineering talent base.
Strengths: Particularly capable where the data requires genuine software engineering competence to produce or judge; its AGI Advancement line supplies reasoning, coding, agentic and STEM data to frontier labs.
Proof points: Founded in 2018 and headquartered in San Francisco, with a talent network spanning 140+ countries. Revenue run rate reached $300 million in 2024, and the company raised $111 million at a $2.2 billion valuation in March 2025, per TechCrunch. Named partners include Anthropic, Google Gemini and Nvidia.
Where it stops: The strength is technical domains; broad multilingual preference work across consumer-facing tasks is a different bench.
6. Mercor
Best for: assembling specialist expert benches quickly.
Strengths: Notable for sourcing domain experts (professional, technical and academic) for high-value data production; it began as an AI hiring platform and pivoted to supplying AI labs with scientists, doctors and lawyers.
Proof points: Headquartered in San Francisco. Raised $350 million at a $10 billion valuation in October 2025, per TechCrunch. Manages more than 30,000 contractors, collectively paid over $1.5 million a day. Publishes the APEX benchmarks assessing frontier models on professional tasks, built with partners such as Ramp and Cognition.
Where it stops: An expert-sourcing marketplace rather than an owned-delivery operation, so security, residency and retention properties differ from a centre-based provider. Multilingual, in-market preference work is not the stated focus.
7. Appen
Best for: broad language coverage with a long track record.
Strengths: One of the longest-established players in linguistic and evaluation data, with wide language support; its LLM offer covers reasoning traces, expert RLHF, SFT demonstrations and adversarial red teaming.
Proof points: Founded in Sydney in 1996, ASX-listed (APX), headquartered in Sydney and Kirkland. Reports 235+ languages, 500+ locales, 170+ countries and 1M+ vetted contributors, and states that 80% of leading LLM builders are customers; all company-reported. SOC 2 and ISO 27001 certified.
Where it stops: The crowd model trades retention for elasticity, which matters most on preference work where rubric stability over months is what makes agreement figures meaningful. See Lifewood vs Appen for large-scale data labelling.
8. Toloka
Best for: flexible crowd capacity with strong tooling for data collection tasks.
Strengths: Capable across a wide range of human-data tasks with a mature platform; the LLM catalogue covers demonstrations, preference data, reasoning chains, agent trajectories, RL environments and red-teaming.
Proof points: Headquartered in Amsterdam and a unit of Nasdaq-listed Nebius Group. Raised $72 million in May 2025 in a round led by Bezos Expeditions, per SiliconANGLE. Reports 6,000+ active contributors across 90+ domains, 100+ countries and 40+ languages, and 50+ automated quality-control methods; all company-reported.
Where it stops: As with any crowd model, consistency on long, complex rubrics requires more client-side management than a managed team, and 40+ languages is narrower than the broadest multilingual providers.
9. Snorkel AI
Best for: programmatic labelling and data-centric development.
Strengths: A genuinely different approach, encoding labelling logic as functions rather than labelling item by item, which suits teams with strong engineering capacity. It has since added expert data services and evaluation benchmarks.
Proof points: Spun out of Stanford AI Lab, founded in 2019 and headquartered in Redwood City. Raised a $100 million Series D in May 2025 at a $1.3 billion valuation, per Business Wire. Named partners include Google, Microsoft, AWS, Anthropic and OpenAI. SOC 2 and HIPAA compliant.
Where it stops: Not a human-data services company in the same sense; the human judgement layer for preference and evaluation, especially across languages, is still yours to source or to buy alongside the platform.
10. iMerit
Best for: expert-in-the-loop work in specialist verticals.
Strengths: Strong domain depth with trained and retained teams, especially in medical, geospatial and autonomous mobility. GenAI services include RLHF and chain-of-thought reasoning data on the Ango Hub platform.
Proof points: Reports 25,000+ domain experts across 60+ countries, company-reported. Acquired by EXL in a deal announced in June 2026 for $170 million upfront plus up to $140 million in earnouts, completed in August 2026, per EXL. Ango Hub handles image, video, LiDAR, DICOM, text, PDF and audio.
Where it stops: Language breadth is narrower than the multilingual providers, and the centre of gravity is annotation rather than alignment data. For perception work, see Lifewood vs iMerit for physical AI annotation.
What should you buy, and in what order?
Buy the evaluation set first, then SFT, then preference data once the rubric is stable, and distillation last.
The sequencing mistake costs more than the vendor choice:
- Evaluation set first. You cannot manage what you cannot measure, and every later decision is made against it. Build it independently of the training data, per language, as in how to build multilingual evaluation sets for LLMs.
- SFT next, targeted narrowly at the specific behaviours that are wrong. A small, well-covered demonstration set usually beats a large diffuse one.
- Preference data third, once the rubric is stable and rater agreement has been demonstrated on a pilot batch; see how to write a preference rubric raters agree on.
- Distillation last, when behaviour is right and the remaining problem is cost or latency.
Buying preference data in volume before the rubric is stable and an evaluation set exists produces low-agreement data, no proof it helped, and no diagnosis afterwards. A vendor who would simply sell volume before the rubric is ready is selling throughput, not outcomes. The enterprise view is in RLHF, SFT and distillation: what enterprise teams buy.
How do you choose the right partner?
Choose by the constraint that binds your programme, because each vendor group is strongest on a different one.
| If your binding constraint is… | Shortlist |
|---|---|
| Consistent quality across many languages under one agreement standard | Lifewood Data Technology, Appen |
| Frontier-lab preference methodology and rater calibre | Scale AI, Surge AI |
| Competitive-level code, mathematics or STEM in one language | Turing, Mercor |
| Standing up a bespoke human workflow fast | Invisible Technologies, Toloka |
| Strong internal engineering; programmatic labelling | Snorkel AI |
| Medical, geospatial or mobility domain depth | iMerit |
Whichever group fits, ask for per-task-family agreement figures from the pilot batch. A managed enterprise LLM training data programme should report it alongside rubric conformance, and independent AI data validation against a gold set checks the number.