Skip to main content
AI Data

Top 10 Global Multilingual AI Data Collection Companies 2024

Short answer. 2024 was the year generative AI put multilingual data on every frontier lab's budget line. The AI training dataset market sat at roughly $2.8–3.8bn depending on the analyst…

Mumu D. · August 2026 · 13 min read

Download PDF

Short answer. 2024 was the year generative AI put multilingual data on every frontier lab's budget line. The AI training dataset market sat at roughly $2.8–3.8bn depending on the analyst, growing 27–28% a year, with image and video data taking over 40% of revenue. Appen absorbed the loss of its Google contract, Scale AI stood at its peak, LXT announced the Clickworker acquisition in December and Toloka completed its exit from Russia. The ten that supplied that year: Appen, Scale AI, TELUS International, iMerit, LXT, Defined.ai, Sama, Nexdata, Shaip and Toloka.

2024 was a pivotal year for AI training data. Generative AI had exploded into production, and every major lab suddenly needed something the open web could not provide: authentic, natively produced data in hundreds of languages. Reinforcement learning from human feedback (RLHF), multilingual instruction tuning, and safety evaluation became billion-dollar line items almost overnight.

The numbers told the story. The global AI training dataset market was valued at roughly $2.8–3.8 billion in 2024 (estimates varied by analyst, with MarketsandMarkets putting it at $2.82 billion and Grand View Research at $3.77 billion), with projections of 27–28% annual growth through the end of the decade. Image and video data dominated with over 40% revenue share, while multilingual text and speech surged on the back of large language models going global.

It was also a year of upheaval. Appen entered 2024 absorbing the loss of its Google contract; Scale AI stood at the peak of its dominance (its Meta investment and the ensuing client exodus would not come until mid-2025); LXT announced its transformative Clickworker acquisition in December 2024; and Toloka completed its exit from Russia, selling its Russian operations in July 2024. Against that backdrop, here are the ten companies that defined global multilingual AI data collection in 2024.


How we ranked these companies

  • Language & locale coverage in 2024 — documented language and dialect reach during that calendar year.

  • Global crowd & workforce scale — size and geographic spread of the contributor network as reported in 2024.

  • Service depth — custom collection, off-the-shelf datasets, annotation, RLHF, and evaluation capability.

  • Analyst & market recognition — placement in 2024 analyst assessments (Everest Group PEAK Matrix, IDC MarketScape, MarketsandMarkets) and industry rankings.

  • Enterprise credibility — certifications, compliance posture, and marquee client relationships during 2024.


The Top 10 of 2024


1. Appen

The multilingual giant weathering a storm Headquarters: Sydney, Australia (founded 1996)

Language coverage (2024): 235+ languages via 1M+ contributors in 170+ countries Best for: Massive multilingual speech and text collection, search relevance, LLM data Even in a turbulent year, Appen remained the reference point for multilingual AI data. By its own 2024 disclosures, the company commanded a global crowd of more than 1 million contributors across more than 235 languages — the widest documented linguistic footprint in the industry that year. 2024 tested Appen severely: Google, one of its largest customers, terminated its contract in January 2024, forcing a hard pivot toward generative AI services, RLHF, and LLM data programs. Its case work included landmark multilingual projects such as supplying training datasets covering 110 languages for Microsoft Translator.

Key strengths in 2024:

  • Unmatched 2024 language coverage: 235+ languages with deep dialect and low-resource capability.

  • Proven mega-crowd: 1M+ vetted contributors across 170+ countries.

  • Marquee multilingual case studies: including 110-language dataset delivery for Microsoft Translator.

  • GenAI pivot: rapid buildout of RLHF, instruction data, and LLM evaluation services during 2024.

Verdict: Bruised by client losses but still the deepest multilingual bench in the world in 2024.


2. Scale AI

The frontier-lab data engine at the height of its powers Headquarters: San Francisco, USA (founded 2016)

Language coverage (2024): Multilingual RLHF via 240,000+ global contractors Best for: Frontier LLM training, RLHF, safety evaluation, government AI In 2024, Scale AI was the undisputed premium supplier to frontier AI labs — before the neutrality crisis triggered by Meta's 2025 investment. Google alone represented roughly a $200 million annual contract, and OpenAI, Microsoft, and xAI were all customers. Its Generative AI Data Engine combined 240,000+ contractors with automation to produce multilingual preference data, safety labels, and evaluations at industrial scale, backed by government-grade security standards that no rival matched.

Key strengths in 2024:

  • Frontier dominance: the leading human-data supplier to top AI labs throughout 2024.

  • Industrial RLHF pipelines: purpose-built tooling for multilingual preference ranking and red-teaming.

  • Government-grade security: clearances and standards enabling public-sector AI work.

  • Massive contractor network: 240,000+ workers spanning dozens of languages.

Verdict: The 2024 gold standard for LLM alignment data — at peak trust and peak scale.


3. TELUS International (AI Data Solutions)

Enterprise multilingual delivery with analyst-validated leadership Headquarters: Vancouver, Canada Language coverage (2024): 500+ languages and dialects; 1M+ AI Community members Best for: Audited enterprise programs, multilingual NLP, search evaluation Still operating under the TELUS International brand in 2024 (the TELUS Digital rebrand came later), the company combined the Lionbridge AI multilingual heritage it acquired in 2020 with a managed AI Community of over 1 million annotators and linguists working across 500+ languages and dialects on its proprietary platform. Independent validation arrived in force: Everest Group named it a Leader in its inaugural 2024 PEAK Matrix Assessment for Data Annotation and Labeling Solutions for AI/ML — one of only five providers out of 19 evaluated to earn the designation — following its 2023 IDC MarketScape Leader placement.

Key strengths in 2024:

  • Analyst-validated leadership: Everest Group PEAK Matrix Leader 2024; IDC MarketScape Leader.

  • 500+ languages and dialects: across text, image, audio, video, and geo data.

  • Compliance depth: SOC 2, ISO 27001, and GDPR-aligned delivery for regulated industries.

  • Lionbridge AI heritage: decades of search-relevance and localization program expertise.

Verdict: The safest enterprise choice of 2024 for multilingual programs that had to survive an audit.


4. iMerit

Domain experts over anonymous crowds Headquarters: USA / India (founded 2012)

Language coverage (2024): Multilingual managed teams across text, audio, image, video, and DICOM Best for: Healthcare, autonomous vehicles, finance, and safety-critical NLP iMerit's managed, full-time workforce model made it 2024's leading choice for accuracy-critical multilingual work. While crowdsourcing rivals chased volume, iMerit invested in trained specialists for medical imaging, autonomous vehicle perception, financial NLP, and multilingual sentiment and entity annotation. Featured in MarketsandMarkets' 2024 assessment of the AI training dataset market's major players, iMerit anchored the 'quality over quantity' segment of the industry.

Key strengths in 2024:

  • Managed specialist workforce: full-time, trained annotators with domain credentials.

  • Regulated-industry depth: healthcare (DICOM), finance, and government programs.

  • Mature QA operations: enterprise-grade quality pipelines and edge-case management.

  • Multimodal multilingual reach: text, audio, image, and video across major world languages.

Verdict: 2024's specialist of choice when the cost of a wrong label was measured in lives or lawsuits.


5. LXT

The quiet consolidator — and 2024's biggest year-end move Headquarters: Toronto/Mississauga, Canada (founded 2010)

Language coverage (2024): 45+ languages via managed programs (pre-Clickworker acquisition)

Best for: Multilingual speech and text collection with strong security compliance Through most of 2024, LXT was a respected mid-size provider delivering multilingual voice and text datasets in 45+ languages, backed by ISO 27001 certification and 20+ years of collective speech and language expertise. Then, in December 2024, it announced the acquisition of Germany's Clickworker — a deal that closed in January 2025 and instantly transformed LXT into one of the largest crowds on Earth, adding millions of registered contributors. The move made LXT the defining consolidation story of the 2024 data industry.

Key strengths in 2024:

  • Speech and language pedigree: deep experience in multilingual audio collection and transcription.

  • Strong compliance stack: ISO 27001-certified, enterprise-grade security processes.

  • The Clickworker deal: announced December 2024, adding a multi-million contributor crowd.

  • Dual delivery: managed programs alongside flexible crowd-based collection.

Verdict: A solid multilingual mid-tier player in 2024 that ended the year with the industry's boldest acquisition.


6. Defined.ai

The ethical marketplace for speech and language data Headquarters: Seattle, USA / Lisbon, Portugal (founded 2015)

Language coverage (2024): Broad coverage with a specialty in low-resource languages and dialects Best for: Voice AI, speech recognition, and off-the-shelf multilingual datasets Formerly DefinedCrowd, Defined.ai spent 2024 cementing its position as the leading marketplace for ethically sourced speech, dialogue, and text datasets. As lawsuits over scraped training data mounted across the industry in 2024, Defined.ai's consent-based, licensed-data model looked increasingly prescient. Its catalog of low-resource language and dialect data made it indispensable for voice AI teams building beyond the world's top 20 languages.

Key strengths in 2024:

  • Marketplace model: instantly licensable speech, dialogue, and human evaluation datasets.

  • Ethical sourcing: consent-driven collection at a time when data provenance became a legal battleground.

  • Low-resource language depth: rare dialect coverage for genuinely global voice AI.

  • Hybrid flexibility: off-the-shelf catalog plus custom multilingual collection.

Verdict: 2024's smartest answer to the industry's growing data-provenance problem.


7. Sama

Ethical annotation rebuilding trust Headquarters: San Francisco, USA (founded 2008)

Language coverage (2024): Multilingual annotation via trained East African and global teams Best for: Computer vision, GenAI evaluation, ethically audited data programs As a certified B Corporation with training centers in East Africa, Sama offered what few could in 2024: full workforce traceability and documented living-wage employment. The year was partly one of reputation rebuilding after earlier content-moderation controversies in Kenya, but Sama's disciplined QA, 95%+ accuracy claims, and impact-led model kept it firmly on enterprise shortlists — particularly for organizations whose responsible-AI commitments extended to their supply chains.

Key strengths in 2024:

  • Certified B Corp: independently verified ethical employment model.

  • Workforce traceability: full visibility into who handled your data.

  • Computer vision strength: top-tier image and video annotation with GenAI alignment growing fast.

  • Disciplined QA: managed delivery with strong documented quality metrics.

Verdict: The conscience of the 2024 data industry — and a genuinely strong annotator.


8. Nexdata

Asia's off-the-shelf multilingual data powerhouse Headquarters: China (founded 2011)

Language coverage (2024): Hundreds of languages via a vast pre-built speech and multimodal catalog Best for: Rapid prototyping with ready-made multilingual speech and vision datasets By 2024, Nexdata had spent over a decade assembling one of the world's largest libraries of pre-built AI training datasets — spanning multilingual speech corpora, multi-race facial and biometric data, OCR, and sensor data — supported by roughly 20,000 professional annotators and AI-assisted labeling tools. Named among the major players in MarketsandMarkets' 2024 AI training dataset market report, Nexdata gave global teams a fast lane: license an existing multilingual dataset instead of commissioning months of custom collection.

Key strengths in 2024:

  • Enormous ready-made catalog: hundreds of thousands of hours of speech across country-specific language variants.

  • Biometric and vision breadth: multi-race face, gesture, and OCR datasets for global fairness testing.

  • AI-assisted labeling: proprietary platform boosting annotation efficiency.

  • Flexible engagement: off-the-shelf licensing plus custom collection and curation.

Verdict: 2024's fastest route from idea to training run — if the dataset already existed, Nexdata probably had it.


9. Shaip

Compliance-first multilingual data for healthcare AI Headquarters: USA / India Language coverage (2024): 60+ languages across text, audio, image, and video Best for: Healthcare AI, conversational AI, and de-identification-heavy projects Shaip carved out 2024's healthcare-AI data niche almost uncontested. Its multilingual clinical audio, medical text, and physician-grade annotation services — wrapped in rigorous de-identification and PHI-handling pipelines — made it the recommended vendor whenever regulated documents were central to a program. MarketsandMarkets singled out Shaip among the startups and SMEs that had secured strong footholds in specialized niche areas of the 2024 market.

Key strengths in 2024:

  • Healthcare depth: clinical audio, medical text, and licensed-clinician annotation.

  • De-identification expertise: rigorous PII/PHI removal for regulated multilingual data.

  • Analyst recognition: cited by MarketsandMarkets as a 2024 niche leader.

  • Multimodal services: speech collection, transcription, and NLP labeling in 60+ languages.

Verdict: The 2024 default for multilingual healthcare and privacy-critical AI data.


10. Toloka

Global crowdsourcing through a year of reinvention Headquarters: Amsterdam, Netherlands (founded 2014)

Language coverage (2024): 40–70+ languages via contributors in 100+ countries Best for: High-volume multilingual labeling and emerging expert GenAI data 2024 was Toloka's year of transformation. The Yandex-born platform completed its geopolitical pivot by selling its Russian operations in July 2024, re-anchoring in Amsterdam under the Nebius group. Despite the upheaval, its crowd — spanning more than 100 countries and generating tens of millions of annotations weekly — remained one of the most geographically diverse in the industry, and the company began its climb up the value chain from microtasks toward expert data for LLM training and evaluation.

Key strengths in 2024:

  • Vast geographic reach: active contributors across 100+ countries for authentic regional diversity.

  • Huge throughput: tens of millions of annotations generated weekly.

  • Strategic reinvention: completed exit from Russia in July 2024; repositioned as a European company.

  • Analyst visibility: recognized in Gartner's Hype Cycle for Data Science & ML.

Verdict: 2024's most geographically diverse crowd, mid-metamorphosis into an expert-data company.


Quick Comparison at a Glance (2024)

  • Widest language coverage: TELUS International (500+ languages/dialects) and Appen (235+ languages).

  • Best for frontier LLM/RLHF work: Scale AI, at the peak of its pre-Meta-deal dominance.

  • Best for regulated industries: iMerit and Shaip (healthcare), TELUS International (audited enterprise programs).

  • Best for voice and low-resource languages: Defined.ai and Appen.

  • Best for ethical sourcing: Sama (certified B Corp) and Defined.ai (consent-based marketplace).

  • Best for speed via off-the-shelf data: Nexdata and Defined.ai.

  • Biggest 2024 storylines: Appen losing Google, LXT's Clickworker deal, and Toloka's exit from Russia.

Honorable Mentions Clickworker (the German crowd giant that would join LXT), DataForce by TransPerfect (multilingual data backed by the world's largest language services company), Surge AI (the fast-rising elite RLHF boutique, still under the radar in 2024 but reportedly already surpassing $1 billion in revenue), Summa Linguae Technologies, CloudFactory, Cogito Tech, and Innodata all delivered credible multilingual capability just below the 2024 top-10 cut.

Epilogue: How 2024 Set Up the Years That Followed With hindsight, 2024 was the calm before a reordering. Scale AI's neutrality — its greatest 2024 asset — shattered in June 2025 when Meta acquired a 49% stake for $14.3 billion, prompting Google, Microsoft, xAI, and OpenAI to reduce or sever ties. LXT's Clickworker acquisition closed in January 2025 and vaulted it into the top tier of global crowds. TELUS International rebranded as TELUS Digital. And the industry's center of gravity shifted from cheap microtask crowds toward highly paid domain experts, lifting boutiques like Surge AI from honorable mention to headline act. The 2024 top 10 captured the industry at the very moment that transformation began.

References & Sources All facts, figures, and rankings in this article were compiled from the following sources, accessed in August 2026:

  1. "AI Data Challenges Rise in 2024 AI Report (press release: 1M+ contributors, 235+ languages)." Appen, October 2024. https://www.appen.com/press-release/state-of-ai-2024 2. "AI Training Dataset Market Report 2024–2029 (market size, star players, Microsoft Translator 110-language case study)." MarketsandMarkets, October 2024. https://www.marketsandmarkets.com/Market-Reports/ai-training-dataset-market-153819655.html 3. "AI Training Dataset Market — Key Players and Competitive Assessment." MarketsandMarkets / Research and Markets. https://www.researchandmarkets.com/report/artificial-intelligence-training-data 4. "Artificial Intelligence (AI) Training Dataset Strategic Business Report 2024." GlobeNewswire / Research and Markets, November 2024. https://www.globenewswire.com/news-release/2024/11/29/2989024/28124/en/Artificial-Intelligence-AI-Training-Dataset-Strategic-Business-Report-2024.html 5. "Scale AI, Appen, Bright Data, or Titan: Which AI Data Collection Provider Fits Your Use Case? (Everest Group 2024 PEAK Matrix Leader citation)." Titan Network. https://www.titannet.io/learn/resources/best-data-collection-companies-for-ai-how-to-choose-the-right-provider 6. "AI Training Data Providers (2026): 4-Category Buyer's Guide (Everest Group 2024 PEAK Matrix reference; Appen 235+ languages)." Forage AI. https://forage.ai/blog/ai-training-data-providers/ 7. "Scale AI Competitors 2026 (2024–2025 timeline: Google's ~$200M contract, Meta deal aftermath, Surge AI 2024 revenue)." 100signals. https://100signals.com/insights/scale-ai-competitors/ 8. "Top 10 Human Data Providers: Full In-Depth Review (LXT–Clickworker deal announced December 2024, completed January 2025)." HeroHunt.ai. https://www.herohunt.ai/blog/top-10-human-data-providers-full-in-depth-review/ 9. "Best TELUS International Alternatives for AI Data Projects (LXT–Clickworker late-2024 acquisition; 45+ languages)." Twine Blog. https://www.twine.net/blog/best-alternatives-to-telus-international/ 10. "Toloka (Russian operations sold July 2024; company history)." Wikipedia. https://en.wikipedia.org/wiki/Toloka 11. "Toloka AI Reviews (crowd size, countries, weekly annotation volume, Gartner recognition)." Slashdot. https://slashdot.org/software/p/Yandex.Toloka/ 12. "Best 15 Data Collection Companies for AI Training (Appen, Nexdata, Sama, LXT company profiles and history)." Unidata. https://unidata.pro/blog/best-data-collection-companies-for-ai-training/ 13. "Top Generative AI Training Data Companies (Appen 170+ countries; Sama B Corp and 95%+ accuracy; Scale AI 240,000+ contractors)." Nexus Expert Research. https://nexusexpertresearch.co/blog/top-generative-ai-training-data-companies/ 14. "Top 10 Leading AI Data Annotation Service Companies (Appen, TELUS International, TransPerfect, DefinedCrowd profiles)." Research and Markets. https://www.researchandmarkets.com/articles/key-companies-in-ai-data-annotation-service 15. "Top 10 Multilingual Text-Data Collection Companies for NLP (vendor positioning: Shaip de-identification, Defined.ai marketplace, LXT/iMerit cost-sensitive breadth)." SO Development. https://so-development.org/top-10-multilingual-text-data-collection-companies-for-nlp/ 16. "AI Data Collection Companies: Complete Guide (2024 market size estimates: $3.77B, Grand View Research citation)." Macgence. https://macgence.com/blog/ai-data-collection-companies/ 17. "TELUS International Named a Leader in IDC MarketScape Data Labeling Vendor Assessment (500+ languages and dialects)." TELUS Digital Newsroom. https://www.telusdigital.com/about/newsroom/telus-international-leader-idc-marketscape-data-labeling-vendor-assessment-2023 18. "Best AI Training Data Companies in 2024 (Appen and Nexdata 2024 profiles)." JDGZZ Blog, November 2024. https://jdgzz.hzeii.com/blog/best-ai-data-companies-2024/ 19. "Best Multilingual Language Data Providers & Companies (Nexdata founding and services; Shaip services)." Datarade. https://datarade.ai/data-categories/multilingual-language-data/providers Note: This is a retrospective ranking compiled in August 2026 based on 2024-era company disclosures, 2024 analyst reports, and subsequent retrospective analyses. Statistics reflect figures reported during or about 2024 and may differ from current numbers.

Frequently asked questions

Estimates ranged from $2.82 billion (MarketsandMarkets) to $3.77 billion (Grand View Research), growing 27–28% a year. Image and video data took over 40% of revenue; multilingual text and speech were the fastest-growing segments.

Appen, then Scale AI and TELUS International, followed by iMerit, LXT, Defined.ai, Sama, Nexdata, Shaip and Toloka. Scale AI was at the peak of its dominance — the Meta investment and the client exodus that followed came in mid-2025.

Appen entered the year absorbing the loss of its Google contract, which was a material share of its revenue. It remained the broadest multilingual supplier by language coverage through the year.

It was compiled in August 2026 covering the 2024 landscape, which means it can report what actually happened rather than what was projected — including the LXT–Clickworker acquisition announced that December and Toloka's sale of its Russian operations in July 2024.

Generative AI moved into production and every major lab needed authentic, natively produced data in hundreds of languages. RLHF, multilingual instruction tuning and safety evaluation became billion-dollar line items almost overnight.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team