LIFEWOOD
Ready100
AIGC

What Happens to Your Data at a Generative AI Vendor

Short answer. It depends on terms most buyers never read, and the answer differs between the model provider and the production vendor sitting in front of it. Three questions settle most…

Lifewood Data Technology · August 2026 · 8 min read

Download PDF

Short answer. It depends on terms most buyers never read, and the answer differs between the model provider and the production vendor sitting in front of it. Three questions settle most of the risk: is your input retained, and for how long; is it used to train or improve anything; and where is it processed and stored. Under the GDPR, any vendor processing personal data on your behalf is a processor, and Article 28 requires a written contract with specified terms — including that they act only on documented instructions and engage no sub-processor without authorisation. Beyond the legal minimum, the controls that actually reduce exposure are unglamorous: minimise what you send, segregate client material, pin sub-processors, and record which model version saw what.

Teams assessing AI vendor risk usually focus on the output and underestimate the input. In a content pipeline the material crossing the boundary is routinely more sensitive than the finished asset — and it crosses before anyone has decided whether the output is any good.

This is a practitioner's summary rather than legal advice; data-protection obligations depend on the specific processing, the jurisdictions involved and the roles of the parties, and should be confirmed with counsel and your data protection officer.


What actually leaves your perimeter?

Six categories, most of which never appear on a security questionnaire:

  • Briefs, which frequently contain unannounced products, pricing, launch dates and strategy.
  • Source material — footage, images, documents, recordings — which may contain identifiable people, third-party IP and confidential settings.
  • Customer data, where content is personalised or generated from records.
  • Reference and brand assets, including unreleased identity work.
  • Reviewer commentary, which is usually blunter than anything else in the chain and is retained alongside the asset.
  • The prompts themselves, which encode method and are commonly logged for longer than the content is.

The second structural point is that there are usually at least two vendors in the chain and their obligations differ. The production vendor holds the relationship and the material. The model provider behind them receives whatever is passed to the API. A contract that binds only the first does not reach where the data actually goes.


The three questions, and three that follow from them

Question A good answer Why it matters
Is our input retained, and for how long? A specific retention period, with zero-retention available for the API tier we use Retention is the difference between a transient exposure and a standing one, and it determines what a breach would disclose
Is it used to train or improve models? No, contractually, for the tier we are on — with the tier named Training use is effectively irreversible; content cannot be withdrawn from a trained model
Where is it processed and stored? Named regions, with residency options where we need them Determines whether a transfer mechanism is required and whether sector rules are satisfiable
Who are the sub-processors? A published list with a change-notification commitment Article 28 requires authorisation for sub-processors; an unlisted one is an unassessed one
What are the deletion mechanics? A defined process, a timeframe and confirmation — including backups Deletion that does not reach backups is not deletion
Which model versions saw our data? Recorded per job Needed for incident response, and unreconstructable after the fact

Answers should be contractual rather than descriptions of current practice. Documentation changes without notice; contract terms do not.


What does the GDPR require of a content pipeline?

Where personal data is involved — and footage of identifiable people, voice recordings and customer records all qualify — a vendor processing it on your behalf is a processor and you are the controller. Article 28 sets out what that relationship requires, and its core provisions map directly onto AI vendor arrangements.

  • A written contract is mandatory, setting out the subject matter, duration, nature and purpose of the processing, the type of personal data and the categories of data subjects.
  • Processing only on documented instructions. A vendor using your inputs to improve its own models is processing for its own purposes, which sits outside that instruction unless you agreed to it.
  • No sub-processor without prior authorisation, specific or general, with notice of changes so the controller can object. In an AI pipeline the model provider is a sub-processor.
  • Confidentiality commitments from personnel, plus appropriate technical and organisational security measures.
  • Deletion or return at the end of the engagement, at the controller's choice.
  • Cooperation with audits, with data subject rights and with breach notification.

Separately, moving personal data outside the EEA engages the transfer rules in Chapter V, which require their own lawful mechanism. This is a common gap in AI arrangements, because inference frequently happens in a different jurisdiction from the one where the vendor is established, and buyers assume the vendor's location determines the processing location. It does not.

The single most common gap is a signed data processing agreement with the production vendor and no visibility into which model APIs they call, in which regions, under what retention. The DPA is only as strong as the sub-processor list attached to it — ask for the list, and ask what happens when it changes.


Which controls actually reduce exposure?

Ordered by effectiveness rather than effort. The first two remove risk; the rest manage it.

  1. Minimise what you send. Most briefs contain more than the work requires. Strip customer identifiers, unannounced product detail and internal strategy that does not affect the output. Data never sent cannot be retained, trained on or breached — the only control with no residual risk.
  2. Redact and pseudonymise source material. Remove identifiable faces and personal data that is not the subject of the work, and replace real customer records with representative equivalents wherever the task permits.
  3. Classify content and route accordingly. Three tiers is usually enough — public, internal, restricted — with a rule for which vendors and which model tiers each may reach. Restricted material may need a dedicated-tenancy path, or may simply not be a candidate for generative production.
  4. Contract for zero retention and no training use. Where the vendor and tier support it, make it a term rather than relying on a settings page, and name the tier, because retention behaviour frequently differs between tiers of the same product.
  5. Pin and monitor sub-processors. Require a published list, prior notice of changes and a right to object. Re-assess when the list changes rather than at annual review; the list is what actually determines where your data goes.
  6. Segregate client material. No shared prompt libraries containing client-specific material, no cross-client reference sets, no fine-tuning on one client's material for another's benefit. This is the confidentiality failure most likely to end a relationship.
  7. Log which model version processed what, per job.
  8. Test deletion. Ask for a specific test asset to be deleted, then ask what happened to backups and logs. A vendor who cannot describe the mechanics has probably not implemented them.

How should you read vendor assurance material?

Certifications and framework alignment are useful evidence and are routinely over-read. Knowing what each does and does not establish saves a great deal of misplaced confidence.

Artefact What it establishes What it does not
ISO/IEC 27001 certificate An information security management system exists and is audited against the standard, within a defined scope Anything about retention or training use of your specific inputs — and the scope may exclude the service you are buying
SOC 2 report Controls against defined trust criteria; a Type II covers a period, a Type I is a point-in-time design assessment Obligations owed to you; it describes the vendor, not your engagement
NIST AI RMF alignment Adoption of a recognised risk framework covering governance, measurement and documentation Independent verification — it is voluntary and self-asserted unless separately assessed
Data processing agreement Enforceable obligations to you specifically Nothing about vendors further down the chain unless the sub-processor list is attached

The practical hierarchy: read the DPA and the sub-processor list first, the retention and training-use terms second, and the certifications last, as corroboration that an organisation capable of honouring the first two exists. A badge on a footer is a claim; ask for the certificate, the scope statement and the audit period, and treat an unproduced certificate as absent.


How Lifewood approaches this

Lifewood processes client briefs, source material and, in its annotation work, substantial volumes of client data — which makes it a processor under the arrangements described above rather than a commentator on them. The commitments a buyer should require from any vendor in that position, this one included, are exactly the ones enumerated here: a signed processing agreement with a named sub-processor list, contractual retention and training-use terms with the tier named, per-job records of which model version processed what, client-segregated material with no shared reference sets, and a deletion process that can be demonstrated on request rather than described.

A vendor answering those with documents is doing the work. A vendor answering with certifications alone has answered a different question — one about their organisation in general rather than about your data specifically. See AIGC services, the delivery methodology, and the companion guide on AI content governance, disclosure and provenance for the per-asset record that makes model-version logging usable.


Sources and further reading

  • GDPR Article 28 (Processor — obligations and required contract terms) and Article 44 (transfers to third countries) — Regulation (EU) 2016/679.
  • ISO/IEC 27001, Information security management systems — International Organization for Standardization.
  • AI Risk Management Framework (AI RMF 1.0) — National Institute of Standards and Technology, January 2023.
  • Article 50, Transparency Obligations — EU Artificial Intelligence Act (consolidated text).

Frequently asked questions

It depends entirely on the vendor and the tier. Many enterprise tiers contractually exclude training use while consumer tiers do not, and the same product can behave differently between them. Get the exclusion in the contract with the tier named, rather than relying on a documentation page or an account setting, because both can change without notice.

It is necessary and not sufficient. A DPA binds your direct vendor; the model providers behind them are sub-processors, and Article 28 requires authorisation for those. Ask for the sub-processor list, the change-notification commitment and the right to object — the DPA is only as strong as that list.

Ask where processing happens, not where the vendor is established, because inference frequently runs in a different jurisdiction. Where personal data leaves the EEA, the GDPR's transfer rules require their own lawful mechanism regardless of how secure the vendor is. Some providers offer regional processing options; whether they cover the specific model you need is worth confirming rather than assuming.

Sometimes, with a lawful basis, a compliant processor arrangement and appropriate minimisation. The better question is usually whether you need to: representative or synthetic equivalents produce comparable results for most content tasks, and data never sent carries no residual risk. Reserve real customer data for work that genuinely requires it.

It means an audited information security management system exists within a defined scope. Read the scope statement, because it may not cover the service you are buying, and note that it addresses security management rather than retention or training use. Certifications corroborate; the processing agreement and the retention terms are what bind.

Sending less. Most briefs contain internal strategy, unannounced product detail and personal data the work does not require, and stripping them costs one review pass. Every other control on the list manages risk that minimisation would have removed outright.

Because incident response needs it. If a retention or training-use term turns out to have been breached, or a model provider discloses an exposure window, the only way to know which of your assets were affected is a per-job record of the model and version that processed each one. That record cannot be reconstructed later.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team