Skip to main content
AI Data

How Data Contributors Should Be Consented and Paid

July 2026 · 9 min read · Updated September 2026

Short answer. There is no single fair rate. NaijaS2ST paid $15 for roughly 250 recorded sentences plus $0.50 per translated sentence; IndicVoices benchmarked against local district daily wages; academic crowdsourcing studies report £15 and $8 effective hourly rates against platform fair-pay floors. Three models — local wage benchmarking, task-rate pricing, and effective hourly rate against a published standard — answer different questions about what "fair" means, and the choice should be deliberate.

Key takeaways

  • Published projects pay very differently: NaijaS2ST paid $15 for about 250 recorded sentences and $0.50 per translated sentence (over $11,000 in total translation costs); IndicVoices paid prevailing district daily wages; crowdsourcing studies reported £15 and $8 effective hourly rates.
  • Three compensation models are in use — local wage benchmarking, task-rate pricing, and effective hourly rate against a published standard — and each answers "fair compared to what?" differently.
  • Good consent practice separates consent to participate from consent to publish, delivers instructions in the contributor's own language, and keeps payment identity separate from the dataset.
  • A 2026 position paper argues that paying data contributors even at conservative rates would make LLM pretraining at current scales unaffordable for most organisations, which is the unresolved tension underneath this whole topic.
  • In small-pool languages, fair terms and prompt payment are also an operational necessity: attrition among scarce contributors is a capacity loss, not just an ethics problem.

What do published data-collection projects actually pay?

Published rates vary widely rather than converging on a norm, because different projects are benchmarking against different things.

NaijaS2ST, a Nigerian speech-to-speech translation project, collected recordings from more than 250 participants. Each recorded approximately 250 sentences and received $15, described by the authors as a fair rate in Nigeria; translators were paid $0.50 per sentence, producing a total translation cost exceeding $11,000. The team stated that given the scale, they did not rely on volunteer recordings, and that they allowed room for negotiation where necessary.

IndicVoices, building an inclusive multilingual dataset for Indian languages, took a different approach: participants were compensated in line with the prevailing daily wages in their respective districts. Academic crowdsourcing sits somewhere else again — one 2026 study paid £1.25 for a survey with a 4.5-minute median completion time, an effective rate of £15 per hour, above both UK minimum wage and the platform's own £9 fair-pay floor; another paid $8.00 per hour against the same platform's standards. Those are four defensible answers to the same question, and they are not close to each other.

What compensation models are used, and what does each assume?

Three models recur in published practice, and each is optimising for a different definition of fairness rather than competing to be the single correct one.

Local wage benchmarking ties payment to prevailing wages in the contributor's own district or country, as IndicVoices did. It is defensible locally and administratively simple, but it institutionalises the wage gap between where data work is done and where its value is captured — which is why Kenya's 2026 draft AI policy proposes benchmarking against international rates rather than domestic minimums instead.

Task-rate pricing, NaijaS2ST's $0.50 per translated sentence, is transparent, scales predictably, and lets contributors calculate their own effective hourly rate. The risk is that it rewards speed over care unless quality gates are in place, and that a rate set from an outside estimate of task duration can be badly wrong.

Effective hourly rate against a published standard is the Prolific approach, where the platform sets a floor and studies report their effective rate against it. It is the most auditable model and the easiest to defend publicly, but it requires knowing actual completion times, which means piloting before setting the rate. The honest position is that these are different answers to "compared to what?", and that answer is a policy decision that should be made deliberately rather than inherited from whatever the last project did.

Why is fair compensation at scale still unresolved?

Fair compensation at the scale modern pretraining requires is not obviously affordable, and the industry's current position rests on that gap rather than a resolution of it.

A 2026 position paper puts it directly: AI systems rely on billions of dollars worth of uncompensated labour, and paying data contributors even at conservatively low rates would make LLM training at current scales infeasible for all but the most resource-rich organisations. Two structural problems follow. First, for commissioned collection the person to compensate is clear — whoever recorded the sentence — but for web-scraped training data, attribution across a trillion tokens is unresolved. Second, the robots.txt convention (the file that tells web crawlers which pages they may access) assumes implicit consent for scraping unless a site owner explicitly opts out, and evidence suggests opt-out requests are routinely ignored by major LLM providers; an actual paid transaction would require active, enforceable consent from all parties instead. The precedent case cited is the discovery that a large number of images in LAION were copied from DeviantArt and used to train Stable Diffusion, with contributors' work learned and reproduced for profit without their knowledge or permission. Opt-in counterexamples exist — OpenAssistant Conversations, WildChat and Mozilla Common Voice all operate on explicit user consent — which is also why commissioned data has a defensibility that scraped data does not: when a client asks whether the people who produced the data agreed and were paid, a commissioned corpus has an answer.

Where does Lifewood stand on this?

Lifewood employs and contracts contributors for data collection across delivery centres in more than 30 countries, so this is an operating reality rather than an abstract question, and two positions follow from it.

Consent and compensation quality is becoming a procurement question rather than a values statement: buyers subject to EU documentation obligations increasingly ask how contributors were treated, not only what was delivered, and a supplier that cannot produce consent records, payment terms and a retention policy is carrying a risk that transfers to the client. The operational argument runs alongside the ethical one: in languages where the qualified contributor pool is small, contributors are the scarce asset in the entire supply chain, and fair terms, prompt payment and a decent experience are how a delivery operation retains the ability to deliver in that language next quarter — attrition in a rare-language programme is a capacity loss that takes months to rebuild through relationship-based recruitment, the kind of work described in recruiting native contributors for African language data. Those two arguments point the same way, which is convenient, and also why practice is improving faster than the discourse suggests.

Frequently asked questions

NaijaS2ST paid $15 for approximately 250 recorded sentences in Nigeria and $0.50 per sentence to translators. IndicVoices paid prevailing district daily wages. Academic crowdsourcing studies reported £15 and $8 effective hourly rates against platform standards.

This is contested. Local benchmarking is administratively simple and defensible in context. Kenya's 2026 draft AI policy proposes international benchmarking on the argument that local rates institutionalise the gap between where data work happens and where its value is captured.

Because someone may be willing to help build a model and unwilling to have their voice released publicly. RedVox obtained these independently, which avoids coercive bundling in a single all-or-nothing form.

The contributor's own. IndicVoices implemented this on ethics committee recommendation. Consent delivered in a national lingua franca to speakers of a regional language weakens the informed element substantially.

Collect payment details separately and do not disclose them with the dataset. AfriVoices-KE used phone numbers and IDs solely for payment processing.

A 2026 position paper argues it is not for most organisations at current scales, and that the industry's economics currently rest on uncompensated labour. That tension is unresolved, and it is a large part of why commissioned data carries a defensibility scraped data does not.

Sources and further reading

  1. "NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages", arXiv, on participant numbers, the $15 per 250 sentences rate, $0.50 per translated sentence, total translation cost, and the decision not to rely on volunteers
  2. "IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages", arXiv, on ethics committee approval, native-language instructions, district daily wage benchmarking and participant hospitality
  3. "AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages", arXiv, on review board and national permit approval, onboarding briefings and separation of payment identity from data
  4. "RedVox: Safety and Fairness Gaps in Speech Models Across Languages", arXiv, on obtaining consent to release independently from consent to participate
  5. "Position: The Most Expensive Part of an LLM should be its Training Data", arXiv, on uncompensated labour, affordability at scale, the robots.txt implicit consent paradigm and opt-in initiatives
  6. "A Pathway Towards Responsible AI Generated Content", arXiv, on the LAION and DeviantArt precedent for training without consent, credit or compensation
  7. Way With Words, "Ethical Speech Data: Navigating Voice Data Collection", on the four core principles and applicable frameworks including GDPR, POPIA and HIPAA

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team