Short answer. Six forces are reshaping it, and four of them arrived inside twelve months. Provenance became legally enforceable in the EU on 2 August 2026. Peer-reviewed research has confined synthetic data to a supporting role, which raises rather than lowers the value of verified human data. Data labour is moving from unregulated to regulated, with Kenya drafting fair-pay benchmarks in 2026. Governments are funding language coverage as public infrastructure. Buyer demand is shifting from annotation volume toward evaluation and preference data. And speech is becoming the primary modality, because only about 3,000 of the world's 7,000-plus languages have an established writing system.
The common thread is that the parts of this work which can be automated are being automated, and what remains requires people who speak the language, in the place it is spoken, with their contribution documented and fairly paid. That is a narrower business than bulk annotation and a more durable one.
What changed in 2026 that makes this a turning point?
Provenance stopped being good practice and became enforceable law. The sequence is worth setting out precisely, because the dates are frequently misreported.
| Date | What applied |
|---|---|
| 1 August 2024 | EU AI Act entered into force |
| 2 February 2025 | Prohibited practices and AI literacy obligations |
| 2 August 2025 | Governance rules and obligations for general-purpose AI models |
| 2 August 2026 | Act generally applicable; transparency rules in effect; AI Office gains enforcement powers over GPAI models |
| 2 December 2027 | High-risk systems in sensitive areas (per the AI Omnibus, in force 27 July 2026) |
| 2 August 2028 | AI embedded in regulated products |
One instrument matters more than any other for data suppliers. Alongside the GPAI Code of Practice, the Commission published a template for the public summary of training content, requiring providers to give an overview of the data used to train their models, including the sources it came from. "Where did this data come from" is now a published answer rather than an internal one.
A dataset is therefore judged on two axes rather than one: is it good, and can you prove where it came from? Suppliers who captured consent, contributor metadata and quality decisions as the work happened are in a different position from those planning to reconstruct it.
Will synthetic data replace human multilingual collection?
No, and the peer-reviewed evidence is unusually clear — but it is overstated in both directions, so the caveats matter.
The reference point is Shumailov and colleagues' 2024 Nature paper showing that models trained recursively on their own outputs degrade, with measurable collapse within roughly five to ten generations when trained on synthetic data alone. The failure has a specific shape: the tails of the distribution go first. Rare, unusual and low-frequency cases disappear while average performance still looks acceptable, and only later does output become bland and repetitive.
Two clarifications:
- It is not an argument against synthetic data. The experiments used purely synthetic data with the original human data discarded and no verification, which is not how serious labs operate. Later work has shown collapse is avoidable through verification and by accumulating real data alongside synthetic rather than replacing it.
- But it is a strong argument for human anchoring, and that argument is sharpest exactly where multilingual work sits. Distribution tails are the whole point of low-resource language collection: regional variants, rare constructions, dialect forms — the things that appear infrequently and are lost first under recursive training.
The economic consequence is already visible in how the field talks about data: verified human provenance is becoming a priced attribute rather than an assumed one. The broader treatment is in is it safe to train AI models on AI-generated data.
What happens as data labour becomes regulated?
Costs become explicit, supply chains become auditable, and the gap between compliant and non-compliant suppliers widens. This is the most underpriced shift in the sector.
Kenya's draft Artificial Intelligence and Other Emerging Technologies Policy, published for consultation in 2026, is the clearest signal so far. It targets data annotation, content moderation and AI quality evaluation roles, and proposes a fair-pay reference framework benchmarked against international rates rather than domestic minimums. The gap it addresses is stark: reported earnings of roughly $1.46 to $3.74 an hour for Kenyan data workers against $21 to $27 for equivalent United States roles.
The draft also proposes mandatory psychosocial support, written contracts, transparent pay reporting and grievance mechanisms, and Kenya's AI Bill 2026 would add a risk-based framework and a dedicated AI commissioner. Academic work converges: a 2026 CHI paper drawing on interviews with Kenyan data workers describes a "regime of entrapment" produced by precarious contracts, weak institutional protection and global labour arbitrage.
Three consequences follow for anyone commissioning multilingual data.
- Labour practices become a procurement question, not a values statement. Buyers subject to EU documentation obligations will increasingly be asked how contributors were treated, not only what was delivered.
- Price expectations reset. A quote built on suppressed wages is not a cheaper version of the same service; it is a different risk profile.
- Retention becomes strategic. In rare languages, trained contributors are the scarce asset. Fair terms are how a supplier retains the ability to deliver in that language next year.
Why are governments now funding language coverage?
Because language capability has been reclassified as national infrastructure, and that changes who pays for the underlying data.
In June 2026 the European Commission selected the EUROPA consortium as winner of the Frontier AI Grand Challenge — a project to build a European open-source frontier AI model in all 24 EU official languages. That is a publicly funded commitment to language coverage no commercial business case would have produced on its own, sitting alongside the EU's AI Gigafactories programme and national sovereign AI efforts elsewhere.
The shift matters for three reasons. New buyers: public programmes and national institutions are becoming significant commissioners, with different requirements from commercial buyers — openness, documentation, auditability, explicit coverage mandates. Different economics: when coverage is a policy objective rather than a revenue calculation, languages that failed a commercial test can still be funded. Open outputs: publicly funded datasets tend to be released openly, which raises the floor for everyone working in those languages.
The corollary is that languages without a state sponsor or a commercial case remain exposed. Public funding is redrawing the map, not flattening it.
How is buyer demand changing?
From volume toward judgement. The market backdrop is expansion — Grand View Research valued data collection and labelling at $3.8 billion in 2024 and projects $17.1 billion by 2030, a compound annual growth rate of 28.4%, with Asia Pacific the fastest-growing region and audio among the fastest-growing data types. Underneath that headline, the mix is changing.
- Evaluation is becoming a product. As models get capable enough that ordinary accuracy checks stop discriminating between them, the valuable work is designing tests that reveal failure — per language, written by speakers rather than translated.
- Preference and alignment data is scarce. Teaching a model to behave appropriately in a culture cannot be scraped or translated, and it is the stage where in-language authorship is least substitutable.
- Red-teaming is multilingual by necessity. Safety behaviour has to be probed in each shipped language, since a guardrail holds only where it was trained.
- Bulk annotation is automating. Routine labelling is increasingly machine-assisted with human verification, compressing margins on volume work while raising the premium on judgement.
For suppliers, the defensible position is moving from throughput to expertise. For buyers, the cheapest line item is rarely the one that determines model quality.
Why is speech becoming the centre of gravity?
Because most of the world's remaining language data was never written down. Roughly 3,000 of the world's 7,000-plus languages have an established writing system, so for predominantly oral languages speech is not one modality among several — it is the only route in.
Three developments push the same way. Recent open speech systems have shown that deliberate field collection can extend recognition to well over a thousand languages, including hundreds never previously served. Audio is among the fastest-growing segments in data collection forecasts. And in many markets voice remains the dominant interface, particularly where literacy in the written standard is lower than fluency in the spoken language.
Speech collection is also structurally harder: physical presence, field conditions rather than studios, biometric-grade consent handling, and transcription decisions that presuppose an orthography the language may not have settled. Those are logistics and governance problems, not model problems, and they favour organisations with people already in the relevant regions.
What should organisations do about all this?
- Capture provenance from day one. Consent, contributor metadata, compensation records and quality decisions recorded as work happens. Nothing in the regulatory direction of travel suggests this softens.
- Anchor synthetic pipelines in verified human data. Accumulate rather than replace, and verify before it enters training. This matters most where real data is thinnest.
- Budget for fair labour and expect it to be audited. The direction is toward benchmarked pay, written contracts and disclosure.
- Invest in evaluation ahead of volume. Per-language evaluation sets written by speakers are becoming the scarce asset.
- Plan speech capability now. For predominantly spoken languages the constraint is field presence and consent infrastructure, and both take longer to build than a data purchase.
How Lifewood approaches this
Lifewood's delivery is built around distributed centres and regional voice operations rather than one central facility, for the reason above: collecting the world's spoken languages cannot be done remotely. 40+ delivery centres across 30+ countries, 50+ languages including underrepresented dialects, and 56,788 registered contributors are what make field presence a starting condition rather than a mobilisation project.
Provenance is captured as the work happens — consent, contributor metadata, compensation records, task assignment and quality decisions — rather than assembled at delivery, because the EU timeline above turns that record into an artefact a buyer may have to publish. Contributors are trained and retained rather than sourced per project; 414,120 training hours were delivered in 2025, which is the mechanism behind being able to deliver in a rare language a second time. Quality is verified against a customer-approved gold set at a 95%+ accuracy SLA, reported per language.
Nothing here is legal advice, and obligations differ by jurisdiction and by how data is produced. See AI data services, multilingual data collection and low-resource language speech data collection.
Sources and further reading
- European Commission, AI Act policy page — application timeline, GPAI obligations, training content summary template, AI Omnibus.
- European Commission, Commission selects EUROPA consortium as the winner of the Frontier AI Grand Challenge, 19 June 2026.
- Shumailov et al., AI models collapse when trained on recursively generated data, Nature, July 2024.
- Transformer, Synthetic data is more useful than you think — the limits of the collapse result.
- ITWeb Africa, Kenya sets standards for AI workers — the draft fair-pay reference framework.
- "The plan is just survival": Data Work in Kenya and the Regime of Entrapment, CHI 2026.
- Grand View Research, Data Collection and Labeling Market Size Report, 2025–2030.
- Cost Analysis of Human-corrected Transcription for Predominately Oral Languages, arXiv.

