Short answer. Annotation security is not a software question, because a human being has to look at the data. The security boundary therefore includes the platform, storage, network, identity controls, the production facility and the annotator's physical working environment, and most certification portfolios cover only the first few. Evaluate seven things: data residency, access management, encryption, physical controls, named subprocessors, incident response and a documented deletion path. Ask for every certificate together with its scope statement.
Key takeaways
- Annotation exposes human workers to proprietary data by design, so a SOC 2 report on a cloud platform covers only a subset of the actual exposure.
- A certificate is only as useful as its scope statement; a certification covering a corporate head office says nothing about the delivery centre where the data will be viewed.
- Data residency must include viewing, not only storage: a dataset stored in the EU but reviewed by an annotator elsewhere has left the region.
- Annotation vendors subcontract; ask for the current subprocessor list, flow-down obligations and the change-notification process.
- Crowd models and controlled-centre models fail differently rather than one being inherently more secure; match the model to the sensitivity of the data.
Why is annotation security not just a software question?
Annotation security covers every place a human can see or touch the data, so the platform, storage, network, identity controls, production facility and the annotator's physical working environment all sit inside the boundary.
Annotation security is the set of technical, physical and contractual controls that protect data during the period a human worker must view it in order to label it. Annotation exposes human workers to proprietary images, customer conversations, medical information, maps, source code, product plans and internal documents. That is the whole point of the service, and it means a SOC 2 report describing a cloud platform, however good, describes a subset of your actual exposure.
This guide sets out what to evaluate, what evidence to require, and where security reviews of annotation vendors go wrong. Security is one of the nine criteria for choosing AI annotation services, and for many buyers of managed AI data services it is the gating one.
What are the seven control areas an enterprise should evaluate?
The control areas are data residency, access management, encryption, physical security, subprocessors, certifications with scope statements, incident response and data deletion, and each has a specific piece of procurement evidence.
| Security area | Enterprise requirement | Procurement evidence |
|---|---|---|
| Data residency | Restrict processing to approved geography | Contractual and technical geographic controls |
| Access management | Least privilege, rapid revocation | RBAC, identity verification, audit logs |
| Encryption | Protection in transit and at rest | Documented encryption controls and key management |
| Physical security | Controlled worker environment where needed | Secure centres, authentication, device restrictions |
| Subprocessors | Visibility into third parties | Current subprocessor list and flow-down obligations |
| Certifications | Relevant independent controls | SOC 2, ISO 27001 and sector standards, with scope statements |
| Incident response | Defined breach and escalation process | Response plan and contractual notification windows |
| Data deletion | Retention and disposal rules | Documented lifecycle controls and deletion confirmation |
A SOC 2 report is an independent examination of a service organisation's controls against the AICPA Trust Services Criteria of security, availability, processing integrity, confidentiality and privacy; it evidences the platform rows of this table, not the physical one. The table doubles as a starting point for an annotation vendor consolidation and RFP, where each row becomes a scored requirement.
Where do security reviews of annotation vendors go wrong?
Reviews most often fail by reading the certification logo instead of its scope, treating platform controls as complete, accepting regulatory alignment as a control, not asking about subprocessors, and deferring residency to the contract stage.
A scope statement is the part of a certificate or audit report that names the entities, locations, systems and services the assessment actually covered.
Reading the logo instead of the scope. A certification covering a corporate head office says nothing about the delivery centre where your data will be viewed. Request the certificate and the scope statement together, every time, and check that the scope names the location and service you are buying. This single check disqualifies more vendors than any other.
Treating platform controls as complete. Encryption at rest, SSO and audit trails protect data in a system, not what happens on the screen a human is looking at, or the phone in their pocket. Both control families are necessary and they are assessed separately.
Accepting "we are GDPR compliant" as a control. Regulatory alignment is a claim about a legal position, not a technical or physical control. Ask what the vendor does, specifically, that implements it: where data is stored, who can access it, how consent and lawful basis are evidenced, what the deletion path is.
Not asking about subprocessors. Annotation vendors subcontract. That is not a scandal; an undisclosed subcontractor is. Under GDPR Article 28 a processor needs the controller's prior written authorisation to engage another processor and must notify intended changes, so the list is yours to ask for. Ask for the current list, the flow-down obligations, and the notification process when it changes.
Deferring the residency question to the contract stage. Whether work can be confined to a named jurisdiction is a property of the vendor's operating model, not a clause. Find out early; it eliminates shortlists.
What physical and workforce controls should you ask about?
For sensitive data, ask whether annotators are remote or in controlled facilities, what device controls are enforced, how identity is verified at the workstation, how quickly access is revoked, how work is segregated by client, and who supervises the room.
- Are annotators remote, in controlled facilities, or a mix? A mixed model is common and fine, provided you know which of your data goes to which.
- What device controls are enforced? Personal devices, phones, cameras, USB ports, screenshots, copy and paste, printing, and the network the workstation sits on.
- How is identity verified at the point of work? Badge, biometric, two-factor, and whether the same person is verified as being at the workstation for the duration.
- How quickly is access revoked after reassignment or termination, and is revocation verified rather than requested?
- Is work segregated by project and by client? Ask how, not whether.
- Who supervises the room? The most sensitive programmes need a named supervisor and a physical access log, not a policy.
These questions apply differently to a crowd model and a centre model. Neither is inherently more secure; they fail differently. A crowd model has a larger, less verifiable population with less environmental control. A centre model concentrates risk in fewer locations with more control and a clearer audit trail. Match the model to the sensitivity of the data rather than to a preference; the trade-offs are set out in controlled centres versus crowd work for sensitive data.
What should you require for data residency?
Require that processing, including viewing, can be confined to a named country or region, enforced technically as well as contractually, inherited by subprocessors, applied to backups and derived artefacts, and evidenced with a report rather than an assurance.
Data residency is the requirement that data is stored, processed and viewed only within a named geography, enforced by technical controls and contract rather than by policy alone. Residency is where legal requirements and operating models collide.
- Can processing be confined to a named country or region, and is that enforced technically as well as contractually?
- Does "processing" include viewing? A dataset stored in the EU but reviewed by an annotator elsewhere has left the region in every sense that matters.
- Do subprocessors inherit the restriction? Flow-down is where residency guarantees usually leak.
- What happens to backups, logs and derived artefacts? Gold sets, QA records and training extracts are data too.
- How is compliance evidenced? A report you can request, or an assurance you can only believe.
The same five questions apply to collected data, not only labelled data; see data sovereignty for AI training data.
What do other annotation providers publish about compliance?
Several annotation vendors publish detailed compliance portfolios; SuperAnnotate, iMerit and Sama each list certifications and controls on a public security page.
All are company-published statements, checked against each vendor's own page in September 2026, and should be requested with scope statements:
- SuperAnnotate publishes SOC 2 Type II, ISO 27001:2022, GDPR and CCPA controls, along with encryption and key management, restricted production data use and a subprocessor list held in its trust centre.
- iMerit publishes SOC 2 Type 2, ISO 27001, ISO 9001:2015, HIPAA, GDPR and TISAX compliance.
- Sama publishes ISO 9001 and ISO 27001 certified delivery centres, biometric facility access at project level, two-factor user authentication, GDPR and CCPA compliance and TISAX; its page lists ISO 42001 and SOC 2 as in progress.
If a certification is mandatory for your programme, make it a pass/fail RFP requirement and verify the scope covers the platform, location and service you will use. If it is desirable rather than mandatory, score it, but score the scope, not the badge. A wider field of providers, ranked on scale rather than security posture, is in the list of large-scale AI data annotation and labelling companies.
How does Lifewood approach annotation security?
Lifewood's relevant property is structural rather than a certificate list: work is delivered through 40+ delivery centres across 30+ countries, which gives procurement a physical control point that a purely remote pool does not offer.
A named facility can be inspected, access-logged, device-restricted and supervised; a distributed crowd cannot. The same footprint makes geographic scoping possible: distributed capacity supports client-mandated data residency, including confining processing to a specified region, enforced through geographic access controls and contractual flow-down to subprocessors, because there is somewhere specific for the work to happen.
For global AI products, the combination that matters is residency options alongside 50+ languages of native-speaker capability, so localisation and controlled processing are not a trade-off.
Data-protection coverage, certification scope and physical controls should be evidenced per engagement rather than assumed from a footer badge, including for Lifewood. Ask for the scope statement here too.