Short answer. There is no single fair rate. NaijaS2ST paid $15 for roughly 250 recorded sentences plus $0.50 per translated sentence; IndicVoices benchmarked against local district daily wages; academic crowdsourcing studies report £15 and $8 effective hourly rates against platform fair-pay floors. Three models — local wage benchmarking, task-rate pricing, and effective hourly rate against a published standard — answer different questions about what "fair" means, and the choice should be deliberate.
Key takeaways
- Published projects pay very differently: NaijaS2ST paid $15 for about 250 recorded sentences and $0.50 per translated sentence (over $11,000 in total translation costs); IndicVoices paid prevailing district daily wages; crowdsourcing studies reported £15 and $8 effective hourly rates.
- Three compensation models are in use — local wage benchmarking, task-rate pricing, and effective hourly rate against a published standard — and each answers "fair compared to what?" differently.
- Good consent practice separates consent to participate from consent to publish, delivers instructions in the contributor's own language, and keeps payment identity separate from the dataset.
- A 2026 position paper argues that paying data contributors even at conservative rates would make LLM pretraining at current scales unaffordable for most organisations, which is the unresolved tension underneath this whole topic.
- In small-pool languages, fair terms and prompt payment are also an operational necessity: attrition among scarce contributors is a capacity loss, not just an ethics problem.
What do published data-collection projects actually pay?
Published rates vary widely rather than converging on a norm, because different projects are benchmarking against different things.
NaijaS2ST, a Nigerian speech-to-speech translation project, collected recordings from more than 250 participants. Each recorded approximately 250 sentences and received $15, described by the authors as a fair rate in Nigeria; translators were paid $0.50 per sentence, producing a total translation cost exceeding $11,000. The team stated that given the scale, they did not rely on volunteer recordings, and that they allowed room for negotiation where necessary.
IndicVoices, building an inclusive multilingual dataset for Indian languages, took a different approach: participants were compensated in line with the prevailing daily wages in their respective districts. Academic crowdsourcing sits somewhere else again — one 2026 study paid £1.25 for a survey with a 4.5-minute median completion time, an effective rate of £15 per hour, above both UK minimum wage and the platform's own £9 fair-pay floor; another paid $8.00 per hour against the same platform's standards. Those are four defensible answers to the same question, and they are not close to each other.
What compensation models are used, and what does each assume?
Three models recur in published practice, and each is optimising for a different definition of fairness rather than competing to be the single correct one.
Local wage benchmarking ties payment to prevailing wages in the contributor's own district or country, as IndicVoices did. It is defensible locally and administratively simple, but it institutionalises the wage gap between where data work is done and where its value is captured — which is why Kenya's 2026 draft AI policy proposes benchmarking against international rates rather than domestic minimums instead.
Task-rate pricing, NaijaS2ST's $0.50 per translated sentence, is transparent, scales predictably, and lets contributors calculate their own effective hourly rate. The risk is that it rewards speed over care unless quality gates are in place, and that a rate set from an outside estimate of task duration can be badly wrong.
Effective hourly rate against a published standard is the Prolific approach, where the platform sets a floor and studies report their effective rate against it. It is the most auditable model and the easiest to defend publicly, but it requires knowing actual completion times, which means piloting before setting the rate. The honest position is that these are different answers to "compared to what?", and that answer is a policy decision that should be made deliberately rather than inherited from whatever the last project did.
What does informed consent actually require?
Informed consent means a contributor understands, in their own language and before recording begins, what data is collected, how it will be used, and that they can withdraw without penalty — not simply that a form was signed.
Four principles recur across ethical frameworks: informed consent, transparency, fairness and accountability, reflected in instruments including the EU's GDPR, South Africa's POPIA and the US HIPAA regime. The operational detail is where projects differ. IndicVoices delivered instructions in participants' native language on its ethics committee's recommendation, because consent delivered in a national lingua franca to speakers of a regional language is consent in a second language and weakens the informed part considerably. RedVox obtained consent to data release independently from consent to participate, letting someone help build a model without agreeing to have their voice published — avoiding the coercive bundling that makes a single consent form all-or-nothing. AfriVoices-KE briefed participants during onboarding with time for questions before profile creation, and collected phone numbers and ID solely for payment processing without disclosing them with the dataset, separating payment identity from data. IndicVoices went through an Institute Ethics Committee and AfriVoices-KE obtained both host-institution review board approval and a national research permit; commercial projects rarely have access to an IRB, but an outside reviewer checking the protocol before it runs can be replicated without one.
Why is fair compensation at scale still unresolved?
Fair compensation at the scale modern pretraining requires is not obviously affordable, and the industry's current position rests on that gap rather than a resolution of it.
A 2026 position paper puts it directly: AI systems rely on billions of dollars worth of uncompensated labour, and paying data contributors even at conservatively low rates would make LLM training at current scales infeasible for all but the most resource-rich organisations. Two structural problems follow. First, for commissioned collection the person to compensate is clear — whoever recorded the sentence — but for web-scraped training data, attribution across a trillion tokens is unresolved. Second, the robots.txt convention (the file that tells web crawlers which pages they may access) assumes implicit consent for scraping unless a site owner explicitly opts out, and evidence suggests opt-out requests are routinely ignored by major LLM providers; an actual paid transaction would require active, enforceable consent from all parties instead. The precedent case cited is the discovery that a large number of images in LAION were copied from DeviantArt and used to train Stable Diffusion, with contributors' work learned and reproduced for profit without their knowledge or permission. Opt-in counterexamples exist — OpenAssistant Conversations, WildChat and Mozilla Common Voice all operate on explicit user consent — which is also why commissioned data has a defensibility that scraped data does not: when a client asks whether the people who produced the data agreed and were paid, a commissioned corpus has an answer.
What consent and payment details are easy to skip?
Several small operational choices determine whether consent and payment hold up in practice, and they are the details most protocols leave silent.
Say what happens if the project changes: consent obtained for "speech recognition research" does not obviously cover a later commercial licence to a third party, so either scope the consent broadly and explain that plainly, or build a re-consent path. Explain retention — how long a recording is kept and what happens at the end — since most consent forms are silent on this. Handle sensitive content by not collecting it: NaijaS2ST stated they did not collect private, personal or sensitive content, nor ask participants to read such material, which is easier than governing sensitive data afterwards. Treat hospitality as part of the protocol, as IndicVoices did by providing tea, coffee, water and biscuits — a small detail that is the difference between a session someone endures and one they would return for, and a quality mechanism in low-resource language work. Pay promptly: delayed payment because a system could not reconcile is a design failure, and in field collection it damages the community relationship the next project depends on.
Where does Lifewood stand on this?
Lifewood employs and contracts contributors for data collection across delivery centres in more than 30 countries, so this is an operating reality rather than an abstract question, and two positions follow from it.
Consent and compensation quality is becoming a procurement question rather than a values statement: buyers subject to EU documentation obligations increasingly ask how contributors were treated, not only what was delivered, and a supplier that cannot produce consent records, payment terms and a retention policy is carrying a risk that transfers to the client. The operational argument runs alongside the ethical one: in languages where the qualified contributor pool is small, contributors are the scarce asset in the entire supply chain, and fair terms, prompt payment and a decent experience are how a delivery operation retains the ability to deliver in that language next quarter — attrition in a rare-language programme is a capacity loss that takes months to rebuild through relationship-based recruitment, the kind of work described in recruiting native contributors for African language data. Those two arguments point the same way, which is convenient, and also why practice is improving faster than the discourse suggests.
What should a practical consent and pay checklist include?
A working checklist covers the compensation benchmark, the consent mechanics, and the operational details, in that order.
Decide the benchmark deliberately — local prevailing wage, task rate, or effective hourly rate against a published standard — and write down which and why. Pilot before setting a task rate, so the effective hourly rate is known rather than assumed, an approach covered in more depth in how a speech data collection programme actually runs. Deliver consent in the contributor's own language, with a briefing and room for questions, and separate consent to participate from consent to publish and payment identity from the dataset. State retention, scope and withdrawal rights explicitly, including what happens if the project's purpose changes, and design prompts to avoid sensitive disclosure rather than governing it afterwards — a discipline that also shows up in building an enterprise multilingual data collection program. Get an external review of the protocol, formal where available and informal where not, and pay promptly, treating the session experience as part of the deliverable rather than an afterthought — the same operating discipline behind one playbook across a global data operation and Lifewood's own AI data services, which also underpins low-resource speech data collection.