Treadstone Associates
Article · 8 min read

Why image metadata gets stripped

A photo can leave a generator carrying a full record of where it came from, and arrive at a viewer’s screen with none of it — not because anyone removed it deliberately, but because most of the systems a file passes through were never built to keep it.

Treadstone Associates · Updated 2026

Key takeaways

  • • A Content Credentials record travels inside the file itself, so any step that re-encodes, screenshots or reprocesses it can leave the record behind.
  • • Social platforms and messaging apps were built to optimize file size and strip location data long before content provenance existed as a design goal.
  • • Google’s SynthID is engineered specifically to survive cropping, filters and compression, an explicit response to how much ordinary handling otherwise destroys.
  • • The absence of a credential says nothing about origin — only that a record, if it ever existed, did not survive this particular path.

What actually travels with a file, and what doesn’t

Two different mechanisms are easy to conflate, and they behave differently under pressure. Content Credentials is data attached to a file: C2PA describes it as “an open technical standard for publishers, creators and consumers to establish the origin and edits of digital content,” functioning “like a nutrition label for digital content.” (c2pa.org) SynthID works the opposite way, embedding a signal directly inside the content: Google’s own description is that the watermark for an image or video “doesn’t change the image or video quality” and is “added the moment content is created,” engineered to “stand up to modifications like cropping, adding filters, changing frame rates, or lossy compression.” (deepmind.google) A record attached alongside a file is structurally more fragile than a signal woven into the file’s own pixels or audio samples, which is exactly why Google states a specific resilience claim for its watermark and no equivalent survival claim exists for a Content Credentials manifest in any material fetched for this piece.

Why resizing, screenshotting and re-uploading are the biggest threats

A screenshot captures pixels only. It has no mechanism to read or forward an attached manifest, because a screen-capture tool was built to record what is displayed, not to inspect the file underneath it — so a screenshot of a provenance-carrying image starts life with zero provenance, through no deliberate act by anyone. A platform’s own re-encoding on upload behaves similarly: services standardize file formats and compress for bandwidth long before any of them supported C2PA, and a manifest is only carried forward if that specific step was rebuilt to preserve it. Content Credentials’ own framing of its progress is telling on this point — its site describes “a collaboration with hundreds of companies” as the measure of adoption, which is itself an admission that the record only survives at each hop where the software handling it was specifically built to keep it.

Why a watermark is a different kind of survivor

SynthID’s design choices target exactly the operations that destroy an attached record. Google states that its audio watermark, embedded through the Lyria model or NotebookLM’s podcast feature, “can’t be altered by common modifications like adding noise, MP3 compression, or changing the speed of the track,” and that its image and video watermark tolerates the same category of change. That resilience is a deliberate engineering answer to the fact that ordinary handling — the same resizing, compression and re-encoding that strips an attached manifest — was the expected threat model from the start, not an afterthought.

What this means for a business relying on provenance

If a workflow depends on a credentials record surviving to final publication, the record has to be checked at each hop, not just at creation — the upload tool, the content management system, the CDN, and any social scheduler in the chain each has to specifically implement C2PA support, or the record is silently dropped somewhere along the way with no error and no notice to anyone. Verifying the published, public-facing copy is the only reliable check; verifying the original file proves only that the record existed once, not that it reached the reader.

Why the gap matters beyond aesthetics

A missing record is not just a lost technical nicety. Canada’s Centre for Cyber Security names the consequence directly, under the heading “Misinformation and disinformation”: “Content not clearly identified as being AI-generated can result in the spread of misinformation, disinformation and confusion. Threat actors use AI in scams and fraudulent campaigns against individuals and organizations.” (cyber.gc.ca) A provenance record is precisely the mechanism that would let a viewer tell the difference — which means every hop that silently strips it is not a cosmetic loss, it is one more image that reaches a viewer with the very identification the Cyber Centre says matters already gone, through nobody’s decision at all.

A worked example

A brand creates a product photo with a generation tool that embeds a Content Credentials record, then hands it to its social media scheduling tool for posting. If that scheduler re-compresses the image on its way out and was never built to preserve a C2PA manifest, the published post carries no record at all — indistinguishable, to a viewer checking for a pin, from a photo that never had one. The gap here is not malicious and nobody made a decision to remove anything; it is simply a tool built before the standard existed, sitting in the middle of an otherwise provenance-aware workflow. The only way to catch this is to check the published post itself, not the file the brand originally exported. If the same brand later fields a complaint that the photo was misleadingly undisclosed as AI-generated, pointing to the credential embedded at creation is not, on its own, a complete answer — the credential has to have survived to the copy the complaint is actually about, and checking that requires looking at what the audience saw, not what the design team produced.

Related: the two different mechanisms this piece compares, why the record matters more than testing the artifact afterward, what to do once a record is gone

Common questions

If a photo has no Content Credentials pin, was it never watermarked?

Not necessarily. It may have carried a record originally that was stripped somewhere in the path from creation to the copy being viewed — a screenshot, a re-upload, or a platform that does not preserve C2PA data will all produce the same visible absence.

Does resizing an image always remove a SynthID watermark?

Google’s own description states the watermark is designed to survive cropping, filters, frame-rate changes and lossy compression, but no accuracy rate for that survival across every possible downstream tool is published in the material reviewed here.

Whose responsibility is it to keep a provenance record intact?

Every tool the file passes through, in practice — the generator that created the record, the CMS or scheduler that touches the file next, and the platform that ultimately serves it to a viewer. A record’s survival is only as strong as the weakest link that handles the file.

Is a stripped record the same problem as a forged one?

No, and the difference matters. A stripped record is an honest gap — the file simply carries no information either way. A forged one actively misrepresents origin, which is the harder problem C2PA’s own framing acknowledges it cannot fully solve: the standard helps a good actor prove authenticity, it does not stop a determined bad actor from fabricating a false one.

Provenance is only as useful as the workflow that preserves it end to end.

Checking whether a credential survives to the published, public-facing copy is a workflow question, not a one-time setup step.