Skip to main content
AIGC

Content Provenance: C2PA, SynthID and What Survives

August 2026 · 9 min read · Updated September 2026

Short answer. Content provenance uses two complementary mechanisms that fail in opposite directions, which is why serious pipelines run both. C2PA Content Credentials attach a cryptographically signed manifest to a file recording what made it and what was done to it — strong evidence, easily removed, because stripping metadata is trivial and ordinary re-encoding does it by accident. Invisible watermarking such as Google's SynthID embeds a signal in the pixels or audio itself, surviving re-encoding and re-upload far better, but proving only that a generator family made the file. Absent provenance is not evidence of anything.

Key takeaways

  • C2PA Content Credentials are a cryptographically signed manifest bound to a file's metadata; they carry a detailed edit chain but are destroyed by routine steps like resizing or screenshotting.
  • Watermarking such as SynthID embeds a signal directly in pixels or audio, so it survives re-encoding and cropping, but reveals only that a generator family produced the file, not its edit history.
  • A missing Content Credential does not mean content is fake — most genuine content in circulation, including camera originals, carries no credential at all.
  • Neither mechanism verifies truth: a signed, disclosed synthetic image can still be used to deceive, and a genuine photo can carry a false caption.
  • The EU AI Act, China's labelling rules and California's AI Transparency Act converge on the same technical answer — machine-readable marking, a watermark for hostile distribution paths, and an internal ledger that survives every platform's handling.

What can content provenance actually tell you?

Provenance technology answers what made a file and what has been done to it since; it is silent on whether what the file shows is true. Provenance is the checkable record of a file's origin and edit history — it raises the cost of certain deceptions but does not certify accuracy.

A perfectly signed Content Credential can accompany a completely misleading image — an authentic photograph of a real event with a false caption, or a synthetic image honestly labelled as synthetic and still deployed to deceive. That is not a defect in the design; it is the correct scope. Provenance is frequently marketed as an answer to misinformation, and organisations then build policy on the assumption that it settles authenticity, which it does not.

The asymmetry that trips people up: provenance is meaningful when present and meaningless when absent. A file carrying a valid manifest tells you something. A file carrying none tells you nothing at all — it might be a camera original from a device that does not write credentials, or a synthetic file whose manifest was destroyed by a resize. Any policy that treats missing provenance as a negative signal will eventually produce a false accusation.

What does a C2PA Content Credential establish?

A Content Credential is a signed manifest bound to a specific file, and it establishes what the producing software claimed, not the truth of that claim. C2PA Content Credentials are the open specification from the Coalition for Content Provenance and Authenticity for writing and signing that manifest.

When an asset is created or modified, the producing application writes assertions into the manifest — what device or model produced it, what actions were taken, what ingredients were used — and signs it with a certificate. A validator checks the signature, confirms the manifest is unaltered, and reads the chain. Where an edited asset derives from earlier assets, those appear as ingredients, so the record forms a chain rather than a single stamp.

The specification has moved quickly: version 2.3 published in January 2026 and 2.4 in April 2026, adding support for live video streaming and manifests for unstructured text. The practical marker of adoption is hardware and platform integration — provenance written by capture devices at the point of photography, and platforms surfacing credentials rather than discarding them.

Claim Established? Why
The manifest is unaltered since signing Yes Cryptographic signature over the manifest contents
The signer is who they say they are Conditionally Depends on the certificate and the trust list used to validate it
This asset was produced by the named model or device As asserted The manifest records the producing software's claim; the validator checks the signature, not the truth of the claim
These edits, in this order, were applied As asserted Only for steps performed by C2PA-aware tools; a step taken elsewhere leaves a gap
Nothing else was done to this asset No An asset can be exported, edited elsewhere and re-signed. The chain shows what was recorded, not what was omitted
The content is truthful No Out of scope by design

The middle rows are where the misunderstanding lives. C2PA validates the integrity of a claim; it does not independently verify the substance of the claim. Published critiques have made this point in detail and are worth reading before a policy leans hard on credential presence.

What does watermarking like SynthID establish?

Watermarking writes a signal into the content itself rather than into the file's metadata, so it establishes that a generator family made the asset without recording an edit history. SynthID is Google's watermarking system that embeds an imperceptible signal into generated images, audio, video and text.

Designed to remain detectable after transformations that remove metadata entirely — re-encoding, compression, cropping, colour adjustment — SynthID has been used to mark more than 100 billion images and videos and tens of thousands of years of audio since it launched, with a public SynthID Detector portal that accepts an upload and reports whether a watermark is present. Adoption has since extended beyond Google's own products to other vendors.

The trade-offs are the mirror image of C2PA's. A watermark carries very little information — essentially "this came from a system in this family" — where a manifest carries a detailed chain. Detection generally depends on the embedding party providing a detector, which makes it vendor-scoped rather than an open standard, although cross-vendor adoption has widened. Robustness is a research question rather than a settled property: watermarks are designed to resist removal, and adversaries are designed to remove them.

  • Watermarking answers "did a generator make this" across a hostile distribution path where metadata will not survive.
  • C2PA answers "what exactly happened to this asset" inside a pipeline you control, and publishes a checkable record alongside the asset.
  • An organisation's own records answer both once a file has left its control and come back stripped — the common case, as covered in AI content governance: disclosure and provenance.
  • Neither answers "is this true," and no policy document should imply otherwise.

How do you build a provenance record that survives your own pipeline?

Most provenance failures are self-inflicted and happen inside the producer's own workflow, long before anything adversarial occurs, so the fix starts with auditing the pipeline rather than the threat model.

  1. Audit every step for metadata survival. Run one test asset end to end — generation, edit, transcode, upload, download — checking at each stage whether the manifest is still present. The result is usually worse than expected, and it identifies exactly which tool is destroying the record.
  2. Sign as close to generation as possible. A credential written at generation and carried forward records more than one applied at export. Where a tool in the middle is not C2PA-aware, document the gap rather than re-signing at the end as though the chain were continuous.
  3. Keep an internal ledger independent of the file. Asset ID, generating model and version, prompt or brief reference, licence, human review record, labels applied, and publication destinations. When the file comes back stripped, this is the only thing that still substantiates the claim, and it is what human-in-the-loop review is meant to produce for an audit.
  4. Layer a watermark where the destination is hostile to metadata. Social platforms, messaging apps and third-party syndication routinely re-encode. A watermark that survives re-encoding is worth more there than a manifest that will not.
  5. Add a visible disclosure where the law or the audience needs one. A label rendered into the picture is the only mechanism guaranteed to survive every pipeline, because it is the picture. It is also the crudest, so reserve it for asset classes where disclosure is a legal duty or a genuine audience expectation — a question examined further in can people tell when content is AI generated, and do they care?
  6. Validate on ingest, not only on export. If you accept assets from agencies, contributors or licensors, check credentials on arrival. Provenance you did not verify at ingest is provenance you are republishing on trust.

How does provenance connect to labelling laws?

Provenance technology is the implementation layer for obligations now written into law, so the same mark-at-export, watermark-for-hostile-paths approach tends to satisfy several regimes at once.

The EU AI Act's Article 50 requires providers of generative systems to mark synthetic outputs in a machine-readable format detectable as artificially generated, with transparency obligations taking effect from 2 August 2026 — machine-readable marking is precisely what a manifest and an embedded watermark provide. China's labelling measures, in force since 1 September 2025, distinguish explicit labels perceivable by users from implicit labels written into file metadata, which maps closely onto the visible-disclosure and embedded-provenance split above, and the full landscape is set out in AI content labelling law: the EU, China and the US. California's AI Transparency Act requires covered providers to embed latent, machine-readable provenance in generated image, video and audio, with later phases placing duties on large platforms not to knowingly strip it.

The convergence is useful: one technical implementation — mark at export, watermark for hostile paths, visible label by asset class, internal ledger throughout — satisfies the mechanics of all three regimes without maintaining separate per-market pipelines.

How does Lifewood approach content provenance?

Lifewood applies provenance metadata and disclosure at the delivery stage of its AIGC programmes and maintains the asset-level ledger described above, because in a fifty-language delivery the ledger is the only record that survives every platform's handling. Delivery is also where the marking decision is applied once across every variant rather than retrofitted per market, an approach covered in more detail on AIGC services and in how the studio handles finished AIGC video production.

Two sentences are worth having ready for stakeholders, because the gap between how provenance is marketed and what it does causes real internal confusion. First, provenance lets a team prove what it made and how, which is a claim about its own content that can be substantiated to a regulator, a client or a platform. Second, it does not let a team prove that someone else's content is fake, and any tool claiming to detect AI content reliably from the file alone should be assumed unreliable until it publishes its false-positive rate.

The second sentence prevents the more damaging mistake. An organisation that adopts a detection tool and starts acting on its outputs will eventually act on a false positive, and the cost of a wrong accusation is considerably higher than the uncertainty it was meant to remove.

Frequently asked questions

Easily, and often accidentally. Stripping metadata is trivial, and ordinary steps — resizing, transcoding, screenshotting, uploading to platforms that re-encode — discard it with no intent to do so. This is the central practical limitation of manifest-based provenance and the reason watermarking and independent record-keeping are complements, not alternatives.

No, and treating it that way is the most common policy error in this area. Most content in circulation carries no credential at all, including genuine camera originals from devices that do not write them. Missing provenance means unknown provenance; only presence is informative.

Detection depends on access to a detector for that watermark. Google operates a public detector portal that accepts uploads and reports whether its watermark is present, and adoption has extended to other vendors' products. It is not a universal detector — it identifies its own family of watermarks, not synthetic content in general.

Start with C2PA if assets stay inside a controlled pipeline and the goal is substantiating an internal production record; it carries far more information and is an open standard. Start with watermarking if assets go straight to platforms that re-encode everything, because a manifest that does not survive the first upload is not doing any work.

Detectors looking for an embedded signal they know about, such as a specific watermark, work within their scope. Detectors inferring synthetic origin from statistical properties of the content are considerably less reliable, and their false-positive behaviour is what matters once a decision is attached to the output. Ask any vendor for the false-positive rate on content resembling yours, and treat the absence of that figure as an answer.

Asset identifier, generating model and version, date, brief or prompt reference, licence and rights position, human review record with reviewer and rubric, labels applied, and every publication destination. This ledger is independent of the file, survives every transcode, and is what answers an audit. It is far cheaper to maintain from the start than to reconstruct later.

Sources and further reading

  1. EU AI Act, Article 50 — transparency obligations
  2. California AI Transparency Act (SB 942)
  3. SynthID — Google DeepMind
  4. SynthID Detector: identify content made with Google's AI tools
  5. China's Measures for Labeling of AI-Generated Synthetic Content (translation)

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team