Skip to main content
AI Data

Consent, Privacy and Pay: How AI Data Contributors Should Be Treated

Short answer. As professionals whose consent is informed and revocable, whose personal data is protected as carefully as the client's, and whose pay is fair, hourly and stable — the…

Mumu D. · July 2026 · 8 min read

Download PDF

Short answer. As professionals whose consent is informed and revocable, whose personal data is protected as carefully as the client's, and whose pay is fair, hourly and stable — the standards we run our delivery centres on. This is not only ethics: research links better pay and stable work directly to annotation accuracy, while the documented alternative — median crowdwork wages near $2 an hour, unpaid invisible labour and precarious piecework — produces exactly the turnover and inconsistency that degrade datasets.


What do the investigations actually document?

A large, essential, mostly invisible workforce working under conditions that would embarrass any other supply chain — and an accountability gap the industry can no longer claim not to see.

The numbers are stark. Surveys of major crowdwork platforms place average earnings between $1 and $5.50 an hour with a median around $2, and only about 4% of workers clearing the US minimum wage of $7.25; for every paid hour, workers spend roughly 18 more minutes on unpaid labour — searching for tasks, qualifying, disputing rejections — and accounting for that invisible work drops measured median wages further still. SOMO's 2026 investigation traced at least 30 intermediary companies through which the largest technology firms source data work, with several accused of paying below minimum wage, dismissing workers unfairly, blocking collective organising and providing no social protections — while pricing pressure and shifting contracts from the top of the chain set the conditions below. Workers are responding: Kenya's Data Labelers Association, the Data Workers Inquiry, Turkopticon — organising documented by Brookings alongside the retaliation some of it meets.

The research community's own house is telling: a systematic review of machine-learning papers using crowdworkers found zero that reported what the workers were paid. The workforce that produces the ground truth of modern AI is, in most of the literature, not even a line item. That silence is the baseline against which any claim to responsible AI data work should be tested — and the reason the standards below are worth writing down in public.

The documented baseline — and its cost US federal minimum wage (reference line)

$7.25/hr Typical surveyed crowdwork range $1–$5.50/hr Median surveyed crowdwork wage And the structural findings +18 min of unpaid work per paid hour — task-hunting, qualifying, disputing — before invisible labour is counted ~$2/hr Workers earning above the US minimum 30+ intermediary companies through which the largest tech firms source data work, per SOMO's 2026 investigation ~4% 0 papers in a systematic ML-research review that reported crowdworker compensation at all Figures from the CrowdWorkSheets compilation of platform wage surveys, SOMO's investigation and the training-data reporting review, as cited below.


What does real consent look like in data work?

Informed, specific, documented and revocable — held to research-ethics standards, because data work is research on human contribution.

The ethics literature grounds this properly: human computation tasks should meet the same standards as behavioural-science research on human subjects, anchored in the Belmont principles — respect for persons, beneficence, justice. Translated into operations, respect for persons means consent that is genuinely informed: before contributing, a person knows what is being collected (their voice, their judgments, their demographic details), what it will be used for, who will receive it, how long it is kept, and what they are paid — in their own language, at reading level, not in a click-through. Specific means new uses need new consent: a voice recorded for speech recognition is not thereby licensed for voice cloning. Documented means the consent record travels with the dataset as part of its provenance — the practice our collection programs treat as part of the deliverable. And revocable means a working withdrawal path, with its practical limits (what has already shipped, what can still be excluded) stated honestly upfront rather than discovered later.

Consent also covers the work itself, not just the data. Contributors are told the nature of content before they opt in — which matters enormously for tasks involving disturbing material — and consent to sensitive work is opt-in with support, never a condition of employment. The Belmont framing's third principle, justice, points at the same place the next two sections do: the people bearing the work should share fairly in its benefits.


How is contributor privacy protected?

Contributors are data subjects, not just data producers — their identities, demographics and voices deserve the same protection discipline as any client dataset.

Separate identity from contribution. The operational rule is separation: identity and verification records live apart from the work product, and what ships to a client is the contribution plus the metadata the dataset legitimately claims — age band, region, dialect — never identity records. Demographic information is collected under consent, used for the balance reporting it exists for, and minimised everywhere else; frameworks like CrowdWorkSheets exist precisely because who annotated a dataset shapes it, and documenting that responsibly means aggregates and safeguards, not exposure.

Voice and content need special handling. Speech is personal data everywhere and treated as biometricgrade in several jurisdictions, so voice datasets carry the strictest tier: explicit consent naming the uses, no repurposing without re-consent, and secure handling through the pipeline. And where contributors process other people's data — the PII inside documents, recordings and screenshots — the protection points both ways: de-identification workflows (the two-pass, human-reviewed pattern we described in our clinical-data work, prompted by automated tools reaching only 95–98% recall) protect the subjects in the data, while access controls, clean-room environments and confidentiality training protect contributors from carrying risk they never chose. Privacy in data work is one discipline with three beneficiaries: the client, the data subject and the contributor.


Why is fair pay a quality decision — and how do we structure it?

Because the evidence says paying people properly is how you get accurate data — retention builds the expertise that guidelines alone cannot.

The evidence runs one direction. Oxford Internet Institute research shows clearer guidance and better pay directly improve annotation accuracy; industry analyses document that stable roles let contributors build the task-specific skill complex guidelines demand, and that fair treatment cuts the turnover that quietly injects errors as replacements relearn every edge case; the Fairwork reporting shows the same principles applied in practice. Even research teams publishing datasets now advertise their labour standards — one alignment dataset documents full-time annotators paid roughly $8–$9 an hour against a local minimum near $3.69, on regulated eight-hour days — because reviewers have started to ask. The practitioner QA literature adds a blunt line to the same ledger: paying per task instead of per hour is listed under what fails.

How we structure it. Our answer is the delivery-centre model itself: contributors work in employmentshaped roles through centres in 30+ countries — trained, hourly-oriented rather than piece-rated, with local labour law as the floor rather than the ceiling, progression tied to the quality record, and the same core standards from Cebu to Benin, because a value that varies by geography is a policy, not a value. It is the model our "AI for good" positioning has to cash out as: the GPT centres we have opened in places like Benin exist to create durable skilled work, not to arbitrage its absence. And it is self-interested in the best way — the 56,788-contributor network that makes our quality system possible only exists because people stay, and people stay because the work is worth staying for.

A caution on the numbers. Wage surveys describe specific platforms and periods, investigation findings are as published by the cited organisations, and our own practices are first-party descriptions at the level we publish them — we deliberately cite principles rather than internal rates here. Verify figures at the original sources, and treat none of this as legal advice.

The contributor standard 1 2 3 4 INFORMED CONSENT Specific, documented, revocable — in the contributor's language, with new uses needing new consent PRIVACY BOTH WAYS FAIR, STABLE PAY Identity separated from contribution; voice treated as biometric-grade; PII protection for subjects and workers alike Hourly-oriented, law as the floor, progression on the quality record — because retention is the quality system ONE STANDARD GLOBALLY The same consent, privacy and treatment rules in every centre — documented, auditable, on the record Ethics and accuracy point the same direction: the conditions that respect contributors are the conditions that produce reliable data.


Key takeaways

    • The documented baseline is grim: surveyed crowdwork medians near $2/hour, ~4% above the US minimum wage, 18 minutes of unpaid labour per paid hour, and 30+ intermediaries through which Big Tech sources data work — several accused of sub-minimum pay and anti-organising practices.
    • The research community mirrors the invisibility: a systematic review found zero ML papers reporting crowdworker compensation.
    • Real consent is research-grade: informed and specific (new uses need new consent — speech recognition is not voice cloning), documented as dataset provenance, revocable with honest limits, and covering the nature of the work itself.
    • Contributor privacy means separation of identity from contribution, demographic data used only for the balance reporting it exists for, biometric-grade handling of voice, and PII protection that shields data subjects and contributors alike.
    • Fair pay is a quality decision on the evidence: better pay and stability measurably improve accuracy, turnover injects errors, and per-task piece rates are documented failure modes.
    • Our structure is the delivery-centre model: employment-shaped, trained, hourly-oriented work with local law as the floor and one standard across 30+ countries — the operating reality behind a 56,788contributor network that only exists because people stay.
    • Ethics and dataset quality are the same investment here, which is the least romantic and most durable argument for doing this right.
    • Wage figures are platform- and period-specific; verify at the cited sources, and treat none of this as legal advice.

Sources and further reading

Frequently asked questions

Three reasons on the record: quality (pay and stability measurably improve accuracy), risk (labour and consent failures in the supply chain are now investigated and reported), and provenance (regulators and customers increasingly ask how training data was made). The cheapest label is rarely the cheapest dataset.

How contributors are engaged (employment-shaped or piece-rated), how pay relates to local wage floors, what consent covers and how it is documented, how identity is protected, and whether quality records — not churn — drive who does the work. Specific answers exist or they do not.

Within honest limits, yes: withdrawal stops future use and removes what can still be removed, while the consent document states upfront what already-shipped data means. Pretending unlimited revocability would be as dishonest as offering none.

By disclosure before opt-in, never as a condition of employment, with rotation and support for those who choose it — the area where the content-moderation record shows the industry at its worst, and where explicit policy matters most.

It changes where the cost sits: experienced, retained contributors produce higher first-pass quality with less rework, and the documented alternative's hidden costs — turnover, inconsistency, reputational and legal exposure — land later and larger. The evidence says fair conditions are how the quality is produced at all.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team