Skip to main content
AI Data

Top 10 Multilingual AI Data Collection Companies in Asia 2025

Short answer. The 2025 Asian top ten — Nexdata, iMerit, DataoceanAI, Datumo, Shaip, Karya, FutureBeeAI, Macgence, Indika AI and Pixta AI — ranked on Asian roots, documented language and…

Mumu D. · August 2026 · 15 min read

Download PDF

Short answer. The 2025 Asian top ten — Nexdata, iMerit, DataoceanAI, Datumo, Shaip, Karya, FutureBeeAI, Macgence, Indika AI and Pixta AI — ranked on Asian roots, documented language and locale coverage, dataset and workforce scale, innovation and independent recognition. This is the year Asian data companies stopped being the global industry's workforce and started publishing its benchmarks.

The year Asia's data companies stopped supplying the AI revolution — and started leading it Compiled August 2026, covering the 2025 landscape Introduction: 2025, Asia's Breakout Year If 2024 established Asia as the engine room of multilingual AI data, 2025 was the year Asian companies moved into the driver's seat. Beijing's DataoceanAI didn't just sell speech corpora — it co-trained and open-sourced Dolphin, a multilingual ASR model with Tsinghua University covering 40 Eastern languages and 22 Chinese dialects on 210,000+ hours of data. Seoul's Datumo raised $15.5 million backed by Salesforce to challenge Scale AI in LLM evaluation. Bengaluru's Karya — 'the world's first ethical data company' — became the template for inclusive data collection, celebrated by India's NITI Aayog and selling Indian-language data to Microsoft and Google.

The macro forces were equally powerful. The Meta–Scale AI deal in June 2025 shattered Western vendor neutrality and sent AI labs hunting for alternatives worldwide. India's IndiaAI Mission poured resources into sovereign language data — with initiatives like BharatGen targeting 15,000+ hours of annotated voice data across all 22 scheduled Indian languages by Q4 2025 and government plans for 500 data labs announced in September. Chinese vendors expanded into ASEAN markets and pivoted into embodied-AI and multimodal data. Asia's data industry was no longer just scaling — it was innovating.

This listicle ranks the top 10 multilingual AI data collection companies headquartered or operationally rooted in Asia during 2025, judged on language coverage, dataset scale, service depth, innovation, and independent recognition. Global players like Appen, TELUS Digital, and LXT ran large Asian operations in 2025, but this list focuses on companies whose identity and core operations are genuinely Asian.


How we ranked these companies

  • Asian roots — headquartered in Asia or with core delivery workforce and identity anchored in Asia.

  • Language & locale coverage in 2025 — documented language, dialect, and data-type reach during that year.

  • Dataset & workforce scale — size of pre-built libraries, contributor networks, and annotation teams.

  • Innovation — 2025 moves into LLM evaluation, embodied AI, speech foundation models, and ethical data models.

  • Recognition — funding, government partnerships, open-source contributions, and market-report placement in 2025.


The Top 10 in Asia, 2025


1. Nexdata (the global brand of Datatang)

China's data factory goes multimodal and embodied Headquarters / Asian base: Beijing, China (Datatang founded 2011; NEEQ-listed 2014)

Language & data coverage (2025): Hundreds of languages; 1M+ hours of speech, 800TB of image/video Best for: Off-the-shelf multilingual datasets, embodied-AI data, and custom collection at industrial scale Operating internationally as Nexdata, Beijing's Datatang solidified its position as Asia's largest pure-play data vendor in 2025 — and dramatically widened its scope. Its publicly documented 2025 case work ranged from an Indonesian language data collection project (October) and a British native lip-reading multimodal project to embodied-AI data collection and COT-VLA robotic arm annotation, while it built out an 'Embodied Intelligence Data Factory' that reached full operation soon after year-end. Alongside a catalog exceeding 1 million hours of multilingual speech and 800TB of vision data, and participation in China's ASEAN 'going global' initiatives to expand into Southeast Asian markets, Nexdata's 2025 showed China's data industry moving decisively beyond labeling into physical AI.

Key strengths in 2025:

  • Colossal multilingual catalog: 1M+ hours of speech across country-specific variants; 800TB of vision data.

  • 2025 multimodal casework: Indonesian speech collection, lip-reading multimodal data, and VLA robotics annotation.

  • Embodied-AI buildout: a dedicated Embodied Intelligence Data Factory constructed through 2025.

  • ASEAN expansion: participation in Beijing's AI 'going global' programs into Southeast Asia.

Verdict: Asia's data superpower in 2025 — now supplying the robots as well as the language models.


2. iMerit

India's expert-workforce leader, vindicated by the expert-data era Headquarters / Asian base: Kolkata, India (founded 2012; US offices)

Language & data coverage (2025): Multilingual managed teams across text, audio, image, video, and medical DICOM Best for: Healthcare, autonomous vehicles, finance, and safety-critical LLM data 2025's industry-wide pivot from commodity crowds to expert data played directly to iMerit's strengths. As frontier labs fled Scale AI after the Meta deal and demand surged for credentialed, auditable annotation, iMerit's model — thousands of full-time, trained, domain-skilled employees across Indian delivery centers — kept it at the top of independent rankings for regulated and detail-sensitive AI programs. Its Scholars program and deepening LLM data services (evaluation, RLHF support, domain-expert annotation) extended India's claim to the quality end of the global data market.

Key strengths in 2025:

  • Managed specialist workforce: full-time, credentialed annotators — the model 2025 validated.

  • Regulated-industry depth: healthcare (DICOM), finance, autonomous vehicles, government.

  • LLM-era services: expert evaluation and domain-heavy generative AI data programs.

  • Consistent recognition: a fixture across 2025 independent vendor analyses.

Verdict: The steady Asian anchor of the global expert-data tier in 2025.


3. DataoceanAI (formerly Speechocean)

From selling speech data to shipping speech models Headquarters / Asian base: Beijing, China (founded 2005)

Language & data coverage (2025): ~200 primary languages and dialects; Dolphin ASR covering 40 Eastern languages + 22 Chinese dialects Best for: Multilingual speech corpora, ASR/TTS data, and speech foundation model development DataoceanAI delivered Asia's most striking data-industry innovation of 2025: Dolphin, a multilingual, multitask ASR model jointly trained with Tsinghua University and released open-source. Dolphin supports 40 Eastern languages across East, South, and Southeast Asia and the Middle East plus 22 Chinese dialects, trained on over 210,000 hours of data combining DataoceanAI's proprietary corpora with open datasets — a landmark for Eastern-language speech AI, where Western models chronically underperform. The company also ran the ICME 2025 Audio Encoder Capability Challenge and rolled out frontier datasets through the year, including multilingual emotional TTS corpora and a 9,000-hour Chinese full-duplex speech corpus for real-time conversational AI.

Key strengths in 2025:

  • Dolphin ASR: open-source model for 40 Eastern languages and 22 Chinese dialects, trained on 210,000+ hours.

  • Two-decade corpus: ~200 primary languages and dialects across speech, text, image, and video.

  • Frontier datasets: full-duplex speech and emotional TTS corpora for next-generation voice AI in 2025.

  • Academic gravity: Tsinghua collaboration and the ICME 2025 audio challenge.

Verdict: 2025's boldest move by any Asian data company: proving the data vendor can build the model too.


4. Datumo (formerly SelectStar)

Seoul's Salesforce-backed challenger to Scale AI Headquarters / Asian base: Seoul, South Korea (founded 2018 by KAIST alumni)

Language & data coverage (2025): Korean-anchored multilingual collection, LLM datasets, and evaluation Best for: LLM evaluation, AI trust and safety data, and crowdsourced collection Datumo seized 2025's neutrality vacuum. In August 2025 — two months after the Meta–Scale deal upended the market — the Seoul startup raised $15.5 million backed by Salesforce Ventures to expand its LLM evaluation business and explicitly challenge Scale AI. Built on Korea's leading crowdsourcing platform (200M+ data cases processed; 300+ clients including Samsung, LG, Naver, Hyundai, and SK Telecom; ~$6M revenue in 2024), Datumo differentiated with licensed pretraining datasets sourced from published literature, automated red-teaming tools, and Korea's first AI trust-and-safety benchmark — making it Northeast Asia's flagship for the evaluation era.

Key strengths in 2025:

  • Salesforce-backed raise: $15.5M in August 2025 to scale LLM evaluation globally.

  • Evaluation-first pivot: benchmark generation, model scoring, and automated red-teaming products.

  • Licensed-data edge: pretraining data from published literature with clean provenance.

  • Blue-chip Korean base: 300+ clients spanning Korea's largest conglomerates.

Verdict: 2025's fastest-rising Asian data company — riding the trust-and-evaluation wave with global ambitions.


5. Shaip

India-powered compliance leader for the regulated AI boom Headquarters / Asian base: US-headquartered with core delivery operations in India Language & data coverage (2025): 100+ languages including rare dialects, across text, audio, image, and video Best for: Healthcare AI, conversational AI, and de-identification-heavy multilingual projects Shaip's India-anchored delivery engine kept it the region's regulated-data standout through 2025. Its HIPAA-compliant workflows, certified medical coders, clinical NLP experts, and physician-grade annotation served diagnostic AI and clinical decision-support builders, while its multilingual voice repository — spanning 100+ languages including rare dialects and low-resource languages — supplied conversational AI programs worldwide. As data-provenance and compliance scrutiny intensified across the industry in 2025, Shaip's GDPR, HIPAA, and SOC 2-aligned governance framework was precisely what regulated enterprises were shopping for.

Key strengths in 2025:

  • Healthcare depth: HIPAA-compliant workflows with certified medical coders and clinical NLP experts.

  • 100+ languages: one of the most diverse multilingual voice repositories, including rare dialects.

  • De-identification expertise: rigorous PII/PHI pipelines for regulated multilingual data.

  • Indian delivery backbone: skilled, scaled workforce powering global programs.

Verdict: The Asia-powered safe pair of hands for regulated multilingual AI data in 2025.


6. Karya

The world's first ethical data company becomes India's model Headquarters / Asian base: Bengaluru, India (nonprofit, founded 2021)

Language & data coverage (2025): Large-scale conversational datasets across all 22 official Indian languages Best for: Ethically sourced Indian-language speech, text, and evaluation data 2025 was the year Karya's radical model went mainstream. The Bengaluru nonprofit — which pays rural and marginalized workers a $5 hourly minimum (roughly 20x India's prevailing data-work wage), grants them ownership of the data they create with royalties on resale, and sells to clients including Microsoft and Google — was celebrated by India's NITI Aayog in 2025 as a blueprint for linguistic inclusivity and decentralized economic growth. Its portfolio grew to large-scale conversational datasets across all 22 official Indian languages, egocentric datasets for embodied AI, and Samiksha, a national-scale multilingual evaluation framework — work that would culminate in a formal IndiaAI Mission partnership to strengthen the country's AIKosh data infrastructure.

Key strengths in 2025:

  • 22 Indian languages: large-scale conversational and multimodal datasets across every scheduled language.

  • Ethical model: $5/hour minimum wage, worker data ownership, and resale royalties — unique in the industry.

  • Evaluation leadership: Samiksha, a major multilingual benchmark across Indian languages, models, and domains.

  • Government embrace: NITI Aayog recognition in 2025, paving the way to the IndiaAI Mission partnership.

Verdict: 2025's most important idea in AI data — proof that ethical sourcing and world-class Indian-language data can scale together.


7. FutureBeeAI

India's dataset marketplace scales up Headquarters / Asian base: Rajasthan, India (founded 2020)

Language & data coverage (2025): Multilingual speech and text; 2,000+ pre-labeled licensable datasets Best for: Off-the-shelf multilingual datasets and crowd-powered speech collection FutureBeeAI continued its rapid climb through 2025, growing its marketplace beyond 2,000 pre-labeled, licensable datasets — heavy on Indian and other under-represented languages — while its Yugo platform enabled scripted and spontaneous conversational speech collection from contributors worldwide, with dual-channel recording for natural dialogue data. Recognized on the Government of India's IndiaAI portal and cited among leading players in global market reports, it rode 2025's surging demand for licensed, provenance-clean multilingual data as legal scrutiny of scraped corpora intensified.

Key strengths in 2025:

  • 2,000+ ready datasets: pre-labeled multilingual speech and text for instant licensing.

  • Yugo platform: global crowd speech collection with dual-channel conversational recording.

  • Under-represented languages: deep Indian-language and low-resource coverage.

  • Provenance advantage: licensed datasets at the moment the industry demanded clean sourcing.

Verdict: India's nimblest dataset marketplace, perfectly positioned for 2025's licensed-data wave.


8. Macgence

India's custom multilingual collection specialist Headquarters / Asian base: India (with US presence)

Language & data coverage (2025): Multilingual speech, text, image, and video collection across global languages Best for: Custom multilingual data collection and annotation for AI/ML pipelines Macgence deepened its position through 2025 as a dependable India-based partner for bespoke multilingual programs — recruiting native speakers across Asia, Europe, and beyond for speech, text, and multimodal collection built to client specification. Its published analysis tracked the market it was riding: India's AI data sector growing at 32.6% annually toward $1.5 billion by 2030. With full-service scope spanning collection, annotation, validation, and licensing, Macgence embodied the maturing middle tier of India's data industry.

Key strengths in 2025:

  • Custom multilingual collection: native-speaker sourcing across Asian, European, and global languages.

  • Full-service scope: collection, annotation, validation, and data licensing.

  • India cost-quality advantage: competitive delivery with skilled linguistic teams.

  • Market insight: well-documented positioning in India's fastest-growing data segment.

Verdict: A reliable Indian workhorse for tailor-made multilingual data through 2025.


9. Indika AI

Mumbai's LLM-era data company comes of age Headquarters / Asian base: Mumbai, India (founded 2021)

Language & data coverage (2025): Multilingual data collection, annotation, RLHF, and fine-tuning across 15+ sectors Best for: Programmatic labeling, foundation-model fine-tuning, and Indian-language AI Indika AI matured rapidly through 2025 into a full data-centric AI company, offering programmatic data labeling, RLHF, and fine-tuning services for large language and foundation models. Its Indian-language focus — exemplified by the Nyaay AI legal platform automating transcription and document processing for India's courts — aligned squarely with 2025's sovereign-AI momentum, as the IndiaAI Mission funded local LLM development and the demand for high-quality Indic-language data exploded. Serving 15+ sectors within four years of founding, Indika represented the ambition of India's new data generation.

Key strengths in 2025:

  • LLM-era services: programmatic labeling, RLHF, and foundation-model fine-tuning data.

  • Indian-language depth: legal, healthcare, and government AI in local languages via Nyaay AI.

  • Sovereign-AI tailwind: positioned inside India's 2025 national AI data push.

  • Data governance: encryption, masking, and synthetic-data anonymization built in.

Verdict: One of Asia's most promising young data companies, riding India's sovereign-AI wave in 2025.


10. Pixta AI

Japan–Vietnam's compliant visual data engine Headquarters / Asian base: Tokyo, Japan / Hanoi, Vietnam Language & data coverage (2025): Multilingual annotation teams; 100M+ licensed visual assets via PixtaStock Best for: Compliant visual datasets, ADAS annotation, and Southeast Asian delivery Pixta AI's core asset — a fully licensed library of over 100 million visual items via PixtaStock — only grew more valuable through 2025 as data-provenance lawsuits and licensing deals reshaped the industry. Its managed annotation service, accelerated 3–4x by pre-annotation and semi-automated labeling, continued serving ADAS, smart-home, and face-recognition clients across Asia, while its Japan–Vietnam operating model exemplified Southeast Asia's expanding role in the global AI data supply chain.

Key strengths in 2025:

  • 100M+ licensed visuals: full-compliance image data at the peak of the provenance era.

  • Speed through automation: pre-annotation and semi-auto labeling at 3–4x traditional pace.

  • Japan–Vietnam model: Japanese enterprise standards with Vietnamese delivery scale.

  • ADAS and vision focus: ground-truth visual data for automotive and smart-device AI.

Verdict: Southeast Asia's standard-bearer for compliant visual data in 2025.


Quick Comparison at a Glance (Asia, 2025)

  • Largest dataset libraries: Nexdata/Datatang (1M+ hours speech, 800TB vision) and DataoceanAI (~200 languages).

  • Biggest 2025 innovations: DataoceanAI's open-source Dolphin ASR (40 Eastern languages), Datumo's Salesforce-backed evaluation pivot, Karya's ethical-data model going national.

  • Best for expert-grade managed annotation: iMerit (India) and Shaip (regulated/healthcare, 100+ languages).

  • Best for Indian-language data: Karya (all 22 scheduled languages), FutureBeeAI, and Indika AI.

  • Best for compliant/licensed data: Pixta AI (visuals), FutureBeeAI, and Datumo (licensed pretraining data).

  • Regional spread: China (2), India (6), South Korea (1), Japan/Vietnam (1) — with India's share growing on sovereign-AI momentum.

Honorable Mentions iFLYTEK (China's speech-AI giant with vast dialectal corpora), TaskUs (Philippines-anchored, expanding aggressively into AI data services), Innodata (major delivery centers in India, Sri Lanka, and the Philippines), AIMMO (South Korean annotation platform), BharatGen and Project EKA (India's sovereign dataset initiatives, building 15,000+ hours of voice data across 22 languages and multi-billion-token Indic corpora through 2025), and DIGI-TEXX, LTS Global Digital Services, and SunTec Data all strengthened Asia's 2025 data ecosystem just below the top-10 cut.

Epilogue: What Asia's 2025 Signaled Three shifts defined Asia's 2025. First, from data to models: DataoceanAI's Dolphin showed Asian data houses climbing the stack into foundation-model development for the languages Western AI neglects. Second, from labor to trust: Datumo's evaluation pivot and Karya's ethical-ownership model repositioned Asian vendors around the industry's new scarcities — safety, provenance, and fairness. Third, from market to mission: with the IndiaAI Mission funding sovereign LLMs, 500 planned data labs, and national corpora like BharatGen and EKA, multilingual data in Asia became state strategy, not just business. The companies on this list weren't just serving the global AI boom in 2025 — they were beginning to redirect it.

References & Sources All facts, figures, and rankings in this article were compiled from the following sources, accessed in August 2026:

  1. "Dolphin: a multilingual, multitask ASR model jointly trained by DataoceanAI and Tsinghua University (40 Eastern languages, 22 Chinese dialects, 210,000+ hours)." GitHub — DataoceanAI. https://github.com/DataoceanAI/Dolphin 2. "DataOcean AI official site (ICME 2025 Audio Encoder Capability Challenge; multilingual emotional TTS; 9,000-hour Chinese full-duplex corpus)." DataoceanAI. https://dataoceanai.com/ 3. "Dataocean AI organisation profile (~200 primary languages and dialects)." InCabin. https://incabin.com/organisation/dataocean/ 4. "Nexdata news archive (2025 case studies: Indonesian language collection, British lip-reading multimodal, embodied AI collection, COT-VLA robotic arm annotation; Embodied Intelligence Data Factory)." Nexdata. https://www.nexdata.ai/company/news 5. "Nexdata official site (embodied AI and egocentric datasets; multilingual catalog)." Nexdata. https://www.nexdata.ai/ 6. "Nexdata repository entry (formerly Datatang — brand relationship)." re3data.org Registry of Research Data Repositories. https://www.re3data.org/repository/r3d100011157 7. "Datatang company entry (founding, NEEQ listing, ASEAN 'going global' participation in 2025)." Baidu Baike (English). https://baike.baidu.com/en/item/Datatang/923869 8. "Seoul-based Datumo raises $15.5M to take on Scale AI, backed by Salesforce (August 2025; ~$6M 2024 revenue; 300+ clients; licensed literature datasets)." TechCrunch. https://techcrunch.com/2025/08/11/seoul-based-datumo-raises-15-5m-to-expand-llm-evaluation-challenging-scale-ai/ 9. "From Data Bottlenecks to AI Trust: How S. Korea's Datumo is Shaping Reliable Generative AI (200M+ data cases; Cash Mission platform; Korea's first AI trust benchmark)." KoreaTechDesk. https://www.koreatechdesk.com/from-data-bottlenecks-to-ai-trust-how-s-koreas-datumo-is-shaping-the-future-of-reliable-generative-ai/ 10. "Karya official site (conversational datasets across 22 official Indian languages; Samiksha evaluation framework; egocentric embodied-AI datasets)." Karya. https://www.karya.in/ 11. "Leveraging AI for Linguistic Inclusivity: Karya's Impact on Rural India (Microsoft and Google as data clients; micro-task model)." NITI Aayog Frontier Tech Hub, July 2025. https://frontiertech.niti.gov.in/story/leveraging-ai-for-linguistic-inclusivity-karyas-impact-on-rural-india/ 12. "The Indian Startup Making AI Fairer — While Helping the Poor (Karya's $5/hour minimum, worker data ownership and royalties)." TIME. https://time.com/6297403/the-workers-behind-ai-rarely-see-its-rewards-this-indian-startup-wants-to-fix-that/ 13. "IndiaAI Mission Partners with Karya to Build AI Ecosystem Through Diverse Language Data (AIKosh infrastructure cooperation)." Devdiscourse / IndiaAI. https://devdiscourse.com/article/law-order/3907190-indiaai-mission-partners-with-karya-to-build-ai-ecosystem-through-diverse-language-data-and-capacity-building 14. "BharatGen: India's First Sovereign AI Initiative (15,000+ hours of annotated voice data across 22 Indian languages by Q4 2025)." BharatGen. https://bharatgen.com/ 15. "India to Set Up 500 Data Labs, Boost AI Capabilities (September 2025 IndiaAI Mission announcement; sovereign LLM funding)." News On Air (Government of India). https://www.newsonair.gov.in/india-to-set-up-500-data-labs-boost-ai-capabilities-with-988-crore-investment 16. "EKA Pretraining Indic Corpus v1 (multi-billion-token Indic corpus aggregated through October 2025)." AIKosh / IndiaAI. https://aikosh.indiaai.gov.in/home/datasets/details/eka_pretraining_indic_corpus_v1_1.html 17. "Top AI Training Data Providers (Shaip: 100+ languages, HIPAA workflows, certified medical coders, rare dialect voice repository)." Technologyspell. https://technologyspell.com/top-ai-training-data-providers-2026/ 18. "Best 15 Data Collection Companies for AI Training (Nexdata catalog scale; iMerit and Shaip profiles)." Unidata. https://unidata.pro/blog/best-data-collection-companies-for-ai-training/ 19. "FutureBeeAI startup profile and official site (founded 2020; Yugo platform; 2,000+ pre-labeled datasets)." IndiaAI portal / FutureBeeAI. https://indiaai.gov.in/startup/futurebeeai and https://www.futurebeeai.com/ 20. "AI Data Collection Companies: Complete Guide (India market growing 32.6% CAGR toward $1.5B by 2030; Macgence multilingual casework)." Macgence. https://macgence.com/blog/ai-data-collection-companies/ 21. "Indika AI company profile (programmatic labeling, RLHF, fine-tuning; Nyaay AI legal platform)." CB Insights. https://www.cbinsights.com/company/indika-ai 22. "PIXTA AI service page (100M+ visual library via PixtaStock; 3–4x faster annotation; ADAS and smart home casework)." Pixta Vietnam. https://pixta.vn/pixta-ai 23. "Scale AI Competitors 2026 (June 2025 Meta–Scale deal and market reordering context)." 100signals. https://100signals.com/insights/scale-ai-competitors/ 24. "Top 9 Data Annotation Companies in Asia-Pacific Region (AIMMO, DIGI-TEXX, LTS Global Digital Services, SunTec Data)." GDS Online. https://www.gdsonline.tech/top-9-data-annotation-companies/ Note: This is a retrospective ranking compiled in August 2026 based on 2025-era company disclosures, contemporaneous reporting, and subsequent analyses. Statistics reflect figures reported during or about 2025 and may differ from current numbers. Nexdata is the international brand of Beijing-based Datatang, listed here as a single entity; Shaip and Pixta AI are included on the basis of Asia-anchored operations despite non-Asian corporate registrations.

Frequently asked questions

Nexdata, followed by iMerit and DataoceanAI, then Datumo, Shaip, Karya, FutureBeeAI, Macgence, Indika AI and Pixta AI.

DataoceanAI's open-source Dolphin ASR family covering 40 Eastern languages, Datumo's Salesforce-backed pivot into model evaluation, and Karya's ethical-data model scaling to national infrastructure.

Nexdata and Datatang, with 1M+ hours of speech and 800TB of vision data, and DataoceanAI at roughly 200 languages.

China (2), India (6), South Korea (1), and Japan/Vietnam (1). India's share grew through 2025 on sovereign-AI momentum.

iMerit in India and Shaip for regulated and healthcare work across 100+ languages. Both run credentialed managed workforces rather than open crowds.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team