Skip to main content
AIGC

What Happens to Your Data at a Generative AI Vendor

July 2026 · 10 min read · Updated September 2026

Short answer. What happens to your data at a generative AI vendor depends on terms most buyers never read, and differs between the production vendor and the model provider behind it. Three questions settle most of the risk: is your input retained and for how long, is it used for training, and where is it processed. Under the GDPR the vendor is a processor bound by Article 28; the controls that cut exposure most are minimising what you send and pinning sub-processors.

Key takeaways

  • In a generative content pipeline the material crossing the vendor boundary (briefs, source footage, customer records, prompts) is routinely more sensitive than the finished asset, and it crosses before anyone has judged the output.
  • There are usually at least two vendors in the chain: the production vendor that holds the relationship and the model provider that receives whatever is passed to the API. A contract that binds only the first does not reach where the data goes.
  • Under GDPR Article 28 a vendor processing personal data on your behalf must have a written contract, act only on documented instructions and engage no sub-processor without prior authorisation.
  • Retention, training use and processing location should be contractual terms with the product tier named, because documentation and account settings change without notice and contract terms do not.
  • Data minimisation is the only control with no residual risk: content that is never sent cannot be retained, trained on or breached.

What actually leaves your perimeter when you use a generative AI vendor?

Six categories of material leave your perimeter in a generative content pipeline, and most of them never appear on a security questionnaire. Teams assessing AI vendor risk usually focus on the output and underestimate the input, which crosses the boundary before anyone has decided whether the output is any good.

A generative AI vendor is any company that receives your briefs, source material or prompts and returns AI-generated content, whether it runs the model itself or calls another provider's API on your behalf.

  • Briefs, which frequently contain unannounced products, pricing, launch dates and strategy.
  • Source material (footage, images, documents, recordings), which may contain identifiable people, third-party IP and confidential settings.
  • Customer data, where content is personalised or generated from records.
  • Reference and brand assets, including unreleased identity work.
  • Reviewer commentary, which is usually blunter than anything else in the chain and is retained alongside the asset.
  • The prompts themselves, which encode method and are commonly logged for longer than the content is.

The second structural point is that there are usually at least two vendors in the chain and their obligations differ. The production vendor holds the relationship and the material. The model provider behind them receives whatever is passed to the API. A contract that binds only the first does not reach where the data actually goes, which is why the questions to ask AI video production partners should include which model APIs they call and under what terms.

This is a practitioner's summary rather than legal advice; data-protection obligations depend on the specific processing, the jurisdictions involved and the roles of the parties, and should be confirmed with counsel and your data protection officer.

Which questions settle most of the data risk?

Three questions settle most of the risk: whether your input is retained and for how long, whether it is used to train or improve models, and where it is processed and stored. Three further questions follow from them: who the sub-processors are, how deletion works and which model versions saw your data.

Question A good answer Why it matters
Is our input retained, and for how long? A specific retention period, with zero-retention available for the API tier we use Retention is the difference between a transient exposure and a standing one, and it determines what a breach would disclose
Is it used to train or improve models? No, contractually, for the tier we are on, with the tier named Training use is effectively irreversible; content cannot be withdrawn from a trained model
Where is it processed and stored? Named regions, with residency options where we need them Determines whether a transfer mechanism is required and whether sector rules are satisfiable
Who are the sub-processors? A published list with a change-notification commitment Article 28 requires authorisation for sub-processors; an unlisted one is an unassessed one
What are the deletion mechanics? A defined process, a timeframe and confirmation, including backups Deletion that does not reach backups is not deletion
Which model versions saw our data? Recorded per job Needed for incident response, and unreconstructable after the fact

Zero retention is a contractual commitment that the model provider does not store your inputs or outputs beyond the time needed to return the response, usually offered only on specific API or enterprise tiers.

Answers should be contractual rather than descriptions of current practice. Documentation changes without notice; contract terms do not. The same discipline applies when choosing an AIGC video production provider: the retention and training-use terms belong in the master agreement, not in a slide.

What does the GDPR require of a content pipeline?

Where personal data is involved, a vendor processing it on your behalf is a processor and you are the controller, and Article 28 of the GDPR sets out what that relationship requires. Footage of identifiable people, voice recordings and customer records all qualify as personal data.

A processor is any organisation that handles personal data on behalf of, and on the instructions of, a controller, and under the GDPR it may only act within a written contract that specifies the processing. Article 28's core provisions map directly onto AI vendor arrangements.

  • A written contract is mandatory, setting out the subject matter, duration, nature and purpose of the processing, the type of personal data and the categories of data subjects.
  • Processing only on documented instructions. A vendor using your inputs to improve its own models is processing for its own purposes, which sits outside that instruction unless you agreed to it.
  • No sub-processor without prior authorisation, specific or general, with notice of changes so the controller can object. In an AI pipeline the model provider is a sub-processor.
  • Confidentiality commitments from personnel, plus appropriate technical and organisational security measures.
  • Deletion or return at the end of the engagement, at the controller's choice.
  • Cooperation with audits, with data subject rights and with breach notification.

Separately, moving personal data outside the EEA engages the transfer rules in Chapter V of the GDPR, beginning at Article 44, which require their own lawful mechanism. This is a common gap in AI arrangements, because inference frequently happens in a different jurisdiction from the one where the vendor is established, and buyers assume the vendor's location determines the processing location. It does not. The same question of where your AI training data actually lives applies to inference data.

The single most common gap is a signed data processing agreement with the production vendor and no visibility into which model APIs they call, in which regions, under what retention. The DPA is only as strong as the sub-processor list attached to it. Ask for the list, and ask what happens when it changes.

Which controls actually reduce exposure?

Data minimisation and redaction remove risk outright; classification, contract terms, sub-processor monitoring, segregation, model-version logging and deletion testing manage the risk that remains. The list below is ordered by effectiveness rather than effort.

Data minimisation means sending a vendor only the material the task requires, so that anything stripped before transfer cannot be retained, trained on or breached.

  1. Minimise what you send. Most briefs contain more than the work requires. Strip customer identifiers, unannounced product detail and internal strategy that does not affect the output. Data never sent cannot be retained, trained on or breached, which makes this the only control with no residual risk.
  2. Redact and pseudonymise source material. Remove identifiable faces and personal data that is not the subject of the work, and replace real customer records with representative equivalents wherever the task permits.
  3. Classify content and route accordingly. Three tiers is usually enough (public, internal, restricted) with a rule for which vendors and which model tiers each may reach. Restricted material may need a dedicated-tenancy path, or may simply not be a candidate for generative production.
  4. Contract for zero retention and no training use. Where the vendor and tier support it, make it a term rather than relying on a settings page, and name the tier, because retention behaviour frequently differs between tiers of the same product.
  5. Pin and monitor sub-processors. Require a published list, prior notice of changes and a right to object. Re-assess when the list changes rather than at annual review; the list is what actually determines where your data goes.
  6. Segregate client material. No shared prompt libraries containing client-specific material, no cross-client reference sets, no fine-tuning on one client's material for another's benefit. This is the confidentiality failure most likely to end a relationship, and it applies equally to the annotation and data-labelling side of a vendor's business.
  7. Log which model version processed what, per job.
  8. Test deletion. Ask for a specific test asset to be deleted, then ask what happened to backups and logs. A vendor who cannot describe the mechanics has probably not implemented them.

How should you read vendor assurance material?

Certifications and framework alignment are useful evidence and are routinely over-read, so the practical rule is to read the data processing agreement and sub-processor list first, the retention and training-use terms second, and the certifications last. Knowing what each artefact does and does not establish saves a great deal of misplaced confidence.

Artefact What it establishes What it does not
ISO/IEC 27001 certificate An information security management system exists and is audited against the standard, within a defined scope Anything about retention or training use of your specific inputs; the scope may exclude the service you are buying
SOC 2 report Controls against the AICPA trust services criteria; a Type II covers a period, a Type I is a point-in-time design assessment Obligations owed to you; it describes the vendor, not your engagement
NIST AI RMF alignment Adoption of a recognised, voluntary risk framework covering governance, mapping, measurement and management of AI risk Independent verification; it is self-asserted unless separately assessed
Data processing agreement Enforceable obligations to you specifically Nothing about vendors further down the chain unless the sub-processor list is attached

A SOC 2 report is an independent auditor's report on a service organisation's controls against the AICPA trust services criteria for security, availability, processing integrity, confidentiality and privacy.

Certifications corroborate that an organisation capable of honouring the processing agreement exists; they do not replace it. A badge on a footer is a claim. Ask for the certificate, the scope statement and the audit period, and treat an unproduced certificate as absent. When comparing vendors on a shortlist such as the best AIGC video production providers compared, ask each one for the same three documents so the answers are comparable.

How does Lifewood approach vendor data security?

Lifewood processes client briefs, source material and, in its annotation work, substantial volumes of client data, which makes it a processor under the arrangements described in this post rather than a commentator on them. The commitments a buyer should require from any vendor in that position, Lifewood included, are exactly the ones enumerated here.

Those commitments are: a signed processing agreement with a named sub-processor list, contractual retention and training-use terms with the tier named, per-job records of which model version processed what, client-segregated material with no shared reference sets, and a deletion process that can be demonstrated on request rather than described. Lifewood's managed AIGC services and AI video production run with human review in the loop, and the reviewer commentary that review generates is itself client material to be segregated and retained on the same terms as the asset.

A vendor answering those requirements with documents is doing the work. A vendor answering with certifications alone has answered a different question, one about their organisation in general rather than about your data specifically. The companion guide on AI content governance, disclosure and provenance describes the per-asset record that makes model-version logging usable.

Frequently asked questions

It depends entirely on the vendor and the tier. Many enterprise tiers contractually exclude training use while consumer tiers do not, and the same product can behave differently between them. Get the exclusion in the contract with the tier named, rather than relying on a documentation page or an account setting, because both can change without notice.

It is necessary and not sufficient. A DPA binds your direct vendor; the model providers behind them are sub-processors, and GDPR Article 28 requires prior authorisation for those. Ask for the sub-processor list, the change-notification commitment and the right to object, because the DPA is only as strong as that list.

Ask where processing happens, not where the vendor is established, because inference frequently runs in a different jurisdiction. Where personal data leaves the EEA, the GDPR's Chapter V transfer rules require their own lawful mechanism regardless of how secure the vendor is. Some providers offer regional processing; confirm it covers the specific model you need.

Sometimes, with a lawful basis, a compliant processor arrangement and appropriate minimisation. The better question is usually whether you need to: representative or synthetic equivalents produce comparable results for most content tasks, and data never sent carries no residual risk. Reserve real customer data for work that genuinely requires it.

It means an audited information security management system exists within a defined scope. Read the scope statement, because it may not cover the service you are buying, and note that it addresses security management rather than retention or training use. Certifications corroborate; the processing agreement and the retention terms are what bind.

Lifewood Data Technology provides managed AI video and content production with human review, alongside its AI data services, and processes client material as a GDPR processor. Any vendor offering human-reviewed AI video should also sign a processing agreement with a named sub-processor list, contract for retention and training-use terms, and log which model version processed each job.

Sources and further reading

  1. Regulation (EU) 2016/679 (GDPR), consolidated text on EUR-Lex — primary text for Article 28 and Chapter V
  2. GDPR Article 28: Processor — written contract, documented instructions, sub-processor authorisation, deletion or return, audits
  3. GDPR Article 44: General principle for transfers — Chapter V transfer rules for personal data leaving the EEA
  4. ISO/IEC 27001:2022, Information security management systems — International Organization for Standardization
  5. AICPA SOC suite of services — SOC 2 reporting on security, availability, processing integrity, confidentiality and privacy
  6. NIST AI Risk Management Framework (AI RMF 1.0) — National Institute of Standards and Technology, January 2023, voluntary framework with Govern, Map, Measure and Manage functions

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team