Skip to main content
AI Data

How Data Contributors Should Be Consented and Paid

Short answer. There is no single fair rate, and published practice varies widely: NaijaS2ST paid $15 for roughly 250 recorded sentences and $0.50 per translated sentence, totalling over…

Mumu D. · July 2026 · 10 min read

Download PDF

Short answer. There is no single fair rate, and published practice varies widely: NaijaS2ST paid $15 for roughly 250 recorded sentences and $0.50 per translated sentence, totalling over $11,000; IndicVoices compensated participants against prevailing daily wages in their own districts; academic crowdsourcing studies report effective rates of £15 and $8 per hour against platform fair-pay standards. Three models are in use — local wage benchmarking, task-rate pricing, and effective hourly rate against a published standard — and each answers a different question about what "fair" means.


Consented and Paid?

Most writing on this subject stays at the level of principle: consent should be informed, compensation should be fair. Nobody disagrees, and nobody can act on it.

So this piece leads with numbers from published projects, because the interesting question is not whether to pay fairly but what fairly has actually meant in practice.

NaijaS2ST, a Nigerian speech-to-speech translation project, collected from more than 250 participants. Each recorded approximately 250 sentences and received $15, described by the authors as a fair rate in Nigeria. Translators were paid $0.50 per sentence, producing a total translation cost exceeding $11,000. The team stated explicitly that given the scale, they did not rely on volunteer recordings, and that they allowed room for negotiation where necessary.

IndicVoices, building an inclusive multilingual dataset for Indian languages, took a different approach: participants were compensated in line with the prevailing daily wages in their respective districts.

Academic crowdsourcing sits somewhere else again. One 2026 study paid £1.25 for a survey with a 4.5 minute median completion time, an effective rate of £15 per hour, exceeding both the UK minimum wage and the platform's own £9 fair-pay floor. Another paid $8.00 per hour against the same platform's standards.

Those are four defensible answers to the same question, and they are not close to each other. Understanding why is the useful part.


The three compensation models, and what each assumes

Local wage benchmarking. IndicVoices tied payment to prevailing district daily wages. The logic is that compensation should be meaningful in the contributor's own economic context, and the practical advantage is that it is defensible locally and administratively simple.

The criticism, which I have covered elsewhere in this series, is that Kenya's 2026 draft AI policy proposes benchmarking against international rates rather than domestic minimums, precisely because local benchmarking institutionalises the wage gap between where data work is done and where its value is captured.

Task-rate pricing. NaijaS2ST's $0.50 per translated sentence. Transparent, scales predictably, and lets contributors calculate their own effective hourly rate. The risk is that it rewards speed over care unless quality gates are in place, and that a rate set from an outside estimate of task duration can be badly wrong.

Effective hourly rate against a published standard. The Prolific approach, where the platform sets a floor and studies report their effective rate against it. This is the most auditable model and the easiest to defend publicly, and it requires knowing actual completion times, which means piloting before setting the rate.

The honest position is that these are not competing philosophies so much as different answers to "compared to what?" And the answer to that question is a policy decision that should be made deliberately rather than inherited from whatever the last project did.


What informed consent actually requires

The four principles that recur across ethical frameworks are informed consent, transparency, fairness and accountability, reflected in instruments including the EU's GDPR, South Africa's POPIA and the US HIPAA regime.

The operational detail is where projects differ, and several published practices are worth copying.

Instructions in the participant's native language. The IndicVoices ethics committee specifically recommended this and the team implemented it. It sounds obvious and it is routinely skipped: consent delivered in a national lingua franca to speakers of a regional language is consent in a second language, which weakens the "informed" part considerably.

Consent to participate and consent to release, obtained separately. The RedVox project obtained consent to data release independently from consent to participate, allowing participants to contribute without agreeing to public release. This is a genuinely good design. It respects that someone may be willing to help build a model and unwilling to have their voice published, and it avoids the coercive bundling that makes a single consent form all-or-nothing.

A pre-recording briefing with room for questions. AfriVoices-KE briefed participants during onboarding and gave them the opportunity to ask questions and seek clarification before creating a profile. Consent obtained by scrolling past a wall of text is weaker than consent obtained after a conversation.

Explicit right to withdraw without consequence. Named in multiple project ethics statements, and worth stating in the contributor's own language rather than in a terms document.

Separation of payment identity from data. AfriVoices-KE collected phone numbers and ID solely for payment processing, not disclosed with the dataset. This is the practical resolution of a real tension: you cannot pay people anonymously, but their payment identity does not need to travel with their voice.

Institutional review where available. IndicVoices went through an Institute Ethics Committee; AfriVoices-KE obtained both host institution Review Board approval and a national-level research permit. Commercial projects rarely have access to an IRB, but the underlying function, someone outside the delivery team reviewing the protocol before it runs, can be replicated.


The tension nobody resolves

Here is the argument that sits underneath this whole topic, and I think it is more honest to state it than to write around it.

A 2026 position paper puts it directly: AI systems rely on billions of dollars worth of uncompensated labour, and paying data contributors even at conservatively low rates would make LLM training at current scales infeasible for all but the most resource-rich organisations.

That is the tension. Fair compensation at the scale modern pretraining requires is not obviously affordable, and the industry's current position rests on that gap.

Two structural problems follow, and the same paper names both.

Who do you compensate? For commissioned collection the answer is clear: the person who recorded the sentence. For web-scraped training data it is not. Attribution across a trillion tokens is unresolved.

The consent mechanism assumes the wrong default. The robots.txt system assumes implicit consent for scraping unless a website owner explicitly opts out, and while an increasing number of domains have used that mechanism to block major LLM providers, evidence suggests these requests are routinely ignored. As the paper observes, if providers were to compensate contributors, this paradigm would have to shift, because any transaction requires active and enforceable consent from all parties.

The precedent case people cite is the discovery that a large number of images in LAION were copied from DeviantArt and used to train Stable Diffusion, with contributors' work learned and reproduced for profit without their knowledge or permission.

Opt-in models exist and are worth knowing as counterexamples: OpenAssistant Conversations, WildChat and Mozilla Common Voice all operate on explicit user consent to data collection.

Why this matters for a commissioned collection programme: it is the reason commissioned data has a defensibility that scraped data does not. When a client asks where the data came from and whether the people who produced it agreed and were paid, a commissioned corpus has an answer.


The details that are easy to skip and shouldn't be

Say what happens if the project changes. Consent obtained for "speech recognition research" does not obviously cover a later commercial licence to a third party. Either scope the consent broadly and explain that plainly, or build a re-consent path.

Explain retention. How long is the recording kept, and what happens at the end. Most consent forms are silent on this.

Handle sensitive content by not collecting it. NaijaS2ST stated they did not collect private, personal or sensitive content, nor ask participants to read such material. Designing the prompts to avoid sensitive disclosure is easier than governing sensitive data afterwards.

Treat hospitality as part of the protocol. IndicVoices made efforts to provide tea, coffee, water and biscuits to ensure a hospitable environment. It reads as a small detail in an ethics section and it is the difference between a session someone endures and one they would return for. Retention, as I have written elsewhere, is a quality mechanism in lowresource language work.

Pay promptly. Delayed payment because a system could not reconcile is a design failure, and in field collection it damages the community relationship the next project depends on.


Where we stand on this

Declaring the interest: Lifewood employs and contracts contributors for data collection across delivery centres in more than 30 countries, so this is our operating reality rather than an abstract question.

Two positions I would defend.

The first is that consent and compensation quality is becoming a procurement question rather than a values statement. Buyers subject to EU documentation obligations are increasingly asked how contributors were treated, not only what was delivered. A supplier that cannot produce consent records, payment terms and a retention policy is carrying a risk that transfers to the client.

The second is an operational argument that runs alongside the ethical one. In languages where the qualified contributor pool is small, contributors are the scarce asset in the entire supply chain. Fair terms, prompt payment and a decent experience are how a delivery operation retains the ability to deliver in that language next quarter. Attrition in a rare-language programme is not an HR metric, it is a capacity loss that takes months to rebuild through relationship-based recruitment channels.

Those two arguments point the same way, which is convenient but also, I think, genuinely why the practice is improving faster than the discourse suggests.


A practical checklist

Decide your benchmark deliberately. Local prevailing wage, task rate, or effective hourly against a published standard.

Write down which and why.

Pilot before setting a task rate, so the effective hourly rate is known rather than assumed.

Deliver consent in the contributor's own language, with a briefing and room for questions.

Separate consent to participate from consent to publish.

Separate payment identity from the dataset.

State retention, scope and withdrawal rights explicitly, including what happens if the project's purpose changes.

Design prompts to avoid sensitive disclosure rather than governing it afterwards.

Get an external review of the protocol, formal where available and informal where not.

Pay promptly and treat the session experience as part of the deliverable.


Key takeaways

  • Published practice varies widely: NaijaS2ST paid $15 for roughly 250 recorded sentences and $0.50 per translated sentence, totalling over $11,000 in translation costs across 250-plus participants in Nigeria.
  • IndicVoices compensated participants in line with prevailing daily wages in their respective districts.
  • Academic crowdsourcing studies reported effective rates of £15 per hour and $8 per hour against platform fair-pay standards.
  • Three compensation models: local wage benchmarking, task-rate pricing, and effective hourly rate against a published standard. Each answers "fair compared to what?" differently.
  • Kenya's 2026 draft AI policy proposes benchmarking against international rather than domestic rates, on the argument that local benchmarking institutionalises the wage gap.
  • The four recurring ethical principles are informed consent, transparency, fairness and accountability, reflected in GDPR, POPIA and HIPAA.
  • IndicVoices delivered all instructions in participants' native language on ethics committee recommendation.
  • RedVox obtained consent to data release separately from consent to participate, letting people contribute without agreeing to publication.
  • AfriVoices-KE briefed participants with time for questions before profile creation, and used phone numbers and ID solely for payment processing without disclosing them.
  • AfriVoices-KE obtained both institutional review board approval and a national research permit; IndicVoices went through an Institute Ethics Committee.
  • A 2026 position paper states that AI systems rely on billions of dollars of uncompensated labour, and that paying contributors even at conservative rates would make training at current scales infeasible for most organisations.
  • The robots.txt paradigm assumes implicit consent unless owners opt out, opt-out requests are reportedly routinely ignored, and any actual transaction would require active enforceable consent from all parties.
  • LAION images copied from DeviantArt and used to train Stable Diffusion is the precedent case for training on work without consent, credit or compensation.
  • Opt-in counterexamples include OpenAssistant Conversations, WildChat and Mozilla Common Voice.
  • Practical details that matter: state retention and scope change, design prompts to avoid sensitive disclosure, provide hospitality during sessions, and pay promptly.
  • Consent and compensation quality is becoming a procurement question, and in small-pool languages fair terms are also the mechanism that retains delivery capacity.

Sources and further reading

Frequently asked questions

NaijaS2ST paid $15 for approximately 250 recorded sentences in Nigeria and $0.50 per sentence to translators. IndicVoices paid prevailing district daily wages. Academic crowdsourcing studies reported £15 and $8 effective hourly rates against platform standards.

This is contested. Local benchmarking is administratively simple and defensible in context. Kenya's 2026 draft AI policy proposes international benchmarking on the argument that local rates institutionalise the gap between where data work happens and where its value is captured.

Because someone may be willing to help build a model and unwilling to have their voice released publicly. RedVox obtained these independently, which avoids coercive bundling in a single all-or-nothing form.

The contributor's own. IndicVoices implemented this on ethics committee recommendation. Consent delivered in a national lingua franca to speakers of a regional language weakens the informed element substantially.

Collect payment details separately and do not disclose them with the dataset. AfriVoices-KE used phone numbers and IDs solely for payment processing.

A 2026 position paper argues it is not for most organisations at current scales, and that the industry's economics currently rest on uncompensated labour. That tension is unresolved, and it is a large part of why commissioned data carries a defensibility scraped data does not.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team