Skip to main content
AI Data

Consent, Privacy and Pay: How AI Data Contributors Should Be Treated

July 2026 · 7 min read · Updated September 2026

Short answer. As professionals: consent that is informed and revocable, personal data protected as carefully as a client's, and pay that is fair, hourly and stable — the standards a delivery centre should run on. Research links better pay and stable work directly to annotation accuracy, while the documented alternative — median crowdwork wages near $2 an hour and precarious piecework — produces the turnover that degrades datasets.

Key takeaways

  • The documented baseline is grim: surveyed crowdwork medians near $2/hour, only about 4% of workers above the US minimum wage, roughly 18 minutes of unpaid labour per paid hour, and 30+ intermediaries through which large technology firms source data work, several accused of sub-minimum pay and anti-organising practices.
  • A systematic review of machine-learning papers using crowdworkers found zero that reported what those workers were paid.
  • Real consent is research-grade: informed and specific (a new use needs new consent — speech recognition is not voice cloning), documented as dataset provenance, and revocable with honest limits stated upfront.
  • Contributor privacy means separating identity from contribution, using demographic data only for the balance reporting it exists for, treating voice as biometric-grade, and protecting both the data subject and the contributor from PII exposure.
  • Fair pay is a quality decision on the evidence: better pay and stability measurably improve accuracy, turnover injects errors, and per-task piece rates are a documented failure mode.
  • Wage and investigation figures are platform- and period-specific; treat them as reported by the cited sources, not as legal advice.

What do the investigations actually document?

A large, essential, mostly invisible workforce working under conditions that would embarrass any other supply chain, and an accountability gap the industry can no longer claim not to see.

Surveys of major crowdwork platforms place average earnings between $1 and $5.50 an hour, with a median around $2, and only about 4% of workers clearing the US minimum wage of $7.25. For every paid hour, workers spend roughly 18 more minutes on unpaid labour — searching for tasks, qualifying, disputing rejections — and accounting for that invisible work drops measured median wages further still. SOMO's 2026 investigation traced at least 30 intermediary companies through which the largest technology firms source data work, several accused of paying below minimum wage, dismissing workers unfairly, blocking collective organising and providing no social protections, while pricing pressure from the top of the chain sets the conditions below. Workers are responding through groups such as Kenya's Data Labelers Association, the Data Workers Inquiry and Turkopticon, organising that Brookings has documented alongside the retaliation some of it meets.

The research community's own record is telling: a systematic review of machine-learning papers using crowdworkers found zero that reported what the workers were paid. The workforce that produces the ground truth of modern AI is, in most of the literature, not even a line item — the baseline against which any claim to responsible AI data work should be tested.

  • US federal minimum wage (reference line): $7.25/hour.
  • Typical surveyed crowdwork range: $1–$5.50/hour, with a median near $2/hour.
  • Share of surveyed workers earning above the US minimum: roughly 4%.
  • Unpaid work per paid hour, before invisible labour is counted: about 18 minutes.
  • Intermediary companies through which large tech firms source data work, per SOMO's 2026 investigation: 30+.
  • ML-research papers reviewed that reported crowdworker compensation: zero.

How is contributor privacy protected?

Contributors are data subjects, not just data producers — their identities, demographics and voices deserve the same protection discipline as any client dataset.

The operational rule is separation: identity and verification records live apart from the work product, and what ships to a client is the contribution plus the metadata the dataset legitimately claims — age band, region, dialect — never identity records. Demographic information is collected under consent, used for the balance reporting it exists for, and minimised everywhere else; frameworks like CrowdWorkSheets exist precisely because who annotated a dataset shapes it, and documenting that responsibly means aggregates and safeguards, not exposure. Speech is personal data everywhere and treated as biometric-grade data — information sensitive enough to require explicit, use-specific consent — in several jurisdictions, so voice datasets carry the strictest tier: consent naming the uses, no repurposing without re-consent, and secure handling through the pipeline.

Where contributors process other people's data — the PII inside documents, recordings and screenshots that shows up across document and OCR annotation work — protection points both ways: de-identification workflows protect the subjects in the data, while access controls, clean-room environments and confidentiality training protect contributors from carrying risk they never chose. Privacy in data work is one discipline with three beneficiaries: the client, the data subject and the contributor.

Why is fair pay a quality decision, and how should it be structured?

Because the evidence says paying people properly is how accurate data actually gets produced — retention builds the expertise that written guidelines alone cannot.

Oxford Internet Institute research shows clearer guidance and better pay directly improve annotation accuracy; industry analyses document that stable roles let contributors build the task-specific skill complex guidelines demand, and that fair treatment cuts the turnover that quietly injects errors as replacements relearn every edge case; Fairwork's reporting shows the same principles applied in practice. Even research teams publishing datasets now advertise their labour standards — one alignment dataset documents full-time annotators paid roughly $8–$9 an hour against a local minimum near $3.69, on regulated eight-hour days, because reviewers have started to ask. The practitioner QA literature adds a blunt line to the same ledger: paying per task instead of per hour is listed among the practices that fail.

The delivery-centre model — an employment-shaped operating structure, as opposed to piece-rated crowdwork platforms — is one answer to that evidence: contributors work through 40+ delivery centres across 30+ countries, trained and hourly-oriented rather than piece-rated, with local labour law as the floor rather than the ceiling, and progression tied to the quality record. The same core standards should hold regardless of location, because a value that varies by geography is a policy, not a value — the same logic behind treating annotator recruitment and certification as a real training pipeline rather than a sign-up form, including for contributors recruited into African-language data programs. It is also self-interested in the most defensible way: the 56,000+ registered-contributor network a vendor can draw on only exists because people stay, and people stay because the work is worth staying for — the same reasoning that runs through running a global data operation across multiple delivery centres and routing sensitive annotation work to trained, vetted teams rather than open crowds.

A caution on the numbers: wage surveys describe specific platforms and periods, investigation findings are as published by the cited organisations, and figures a vendor reports about its own operations (delivery-centre count, contributor count) should be labelled company-reported, at the level the vendor itself publishes them. Verify figures at the original sources, and treat none of this as legal advice. Buyers comparing human-in-the-loop annotation providers should expect a straight answer to how contributors are engaged, since it is a leading indicator of the quality a vendor can actually hold.

Frequently asked questions

Three reasons: quality (pay and stability measurably improve accuracy), risk (labour and consent failures in the supply chain are now investigated and reported), and provenance (regulators and customers increasingly ask how training data was made). The cheapest label is rarely the cheapest dataset.

How contributors are engaged (employment-shaped or piece-rated), how pay relates to local wage floors, what consent covers and how it is documented, how identity is protected, and whether quality records rather than churn determine who does the work. Specific answers exist or they do not.

Within honest limits, yes: withdrawal stops future use and removes what can still be removed, while the consent document states upfront what already-shipped data means. Claiming unlimited revocability would be as dishonest as offering none at all.

Through disclosure before opt-in, never as a condition of employment, with rotation and support available to those who choose it. This is the area where the content-moderation record shows the industry at its worst, and where explicit policy matters most.

It changes where the cost sits: experienced, retained contributors produce higher first-pass quality with less rework, while the documented alternative's hidden costs — turnover, inconsistency, reputational and legal exposure — land later and larger. The evidence says fair conditions are how quality is produced at all.

Sources and further reading

  1. CrowdWorkSheets (arXiv), compiling platform wage surveys: $1–$5.50/hour ranges, ~$2 medians, ~4% above US minimum, and unpaid invisible labour findings
  2. SOMO, "Big Tech sets unfair terms and conditions for AI data workers globally," the 2026 investigation of 30+ intermediaries and documented labour practices
  3. Brookings, "Reimagining the future of data and AI labor in the Global South," on worker organising, retaliation and ethical alternatives
  4. "Garbage In, Garbage Out?" (arXiv), the review finding zero ML papers reporting crowdworker compensation
  5. "Trustworthy Human Computation: A Survey" (arXiv), on holding crowdwork to research-ethics standards and the Belmont principles
  6. Welo Data, "Beyond Compliance: Building Ethical AI Starts With Fair Work," compiling Oxford Internet Institute, Fairwork and industry evidence linking conditions to accuracy
  7. "SafeSora" (arXiv), a published example of documented labour standards in dataset construction: wages against local minimums and regulated hours
  8. Label Your Data, "Annotation QA: 2026 Strategies," listing per-task pay among documented failure modes
  9. Lifewood, the delivery-centre model, contributor network and human-in-the-loop quality practice

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team