Do AI images carry hidden watermarks?
Yes, and usually two of them at once, doing opposite jobs. One is a pattern woven into the pixels that you cannot see and cannot easily remove. The other is a signed record attached to the file, which any screenshot destroys. Understanding which is which explains why an AI image can be both traceable and untraceable at the same time, and why neither fact is the thing that decides whether a figure belongs in your paper.
If you only came for the practical rule: a figure that reports evidence has to come from the evidence, and no watermark question changes that. Section seven is the checklist.
Two marks, doing different jobs
Every current provenance scheme for images is one of two kinds, and they fail in opposite directions.
| Pixel watermark | Content Credentials | |
|---|---|---|
| Example | SynthID | C2PA |
| Where it lives | In the pixel values | In signed metadata beside the pixels |
| Says what | This came from our model | This file was made by X and edited by Y |
| Survives a screenshot | Largely, yes | No. It is gone |
| Readable by | The provider's detector | Anyone, with a public verifier |
So the durable one is unreadable to you, and the readable one is fragile. That is not a temporary state of the technology, it is a straight consequence of where each one is stored, and every claim you read about image provenance should be sorted into one column or the other before you evaluate it.
The pattern in the pixels
Google's SynthID adjusts pixel values across the image in a pattern keyed to a secret, spread widely enough that no local region carries the whole signal and subtly enough that the picture looks unchanged. Because the mark is in the image data rather than beside it, ordinary handling does not remove it: cropping, resizing, saving as a different format and moderate compression all leave enough of the pattern for the detector to find. Google applies it across Gemini's image, audio and video output.
OpenAI applies an invisible SynthID watermark to its image output as well as C2PA metadata, and marks generated audio with SynthID. Anthropic's August 2026 marking policy covers generated .svg, .png and .jpg files with signed provenance metadata rather than a pixel watermark, which is a meaningfully weaker commitment for exactly the reasons in the next section.
Reading a pixel watermark needs the provider's detector. Google's SynthID Detector portal exists, is waitlisted, and covers content from Google's own models only. There is no universal image checker, and the same key-secrecy problem that blocks one for text blocks one here.
The record attached to the file
C2PA takes the opposite approach. Instead of hiding a signal in the image, it attaches a cryptographically signed manifest recording what created the file and what has edited it since. Anyone can read it: drag a file onto the Content Credentials verifier and the manifest appears if it survived.
The adoption list is genuinely broad now. Adobe Firefly, Microsoft Designer and OpenAI's image tools write manifests. Some cameras do, including Leica's M11-P. Google Pixel phones mark images touched by AI editing features, and Google has been rolling C2PA verification into Gemini, Search and Chrome. In principle this is the better system: it is open, inspectable, and records edits rather than just origin.
In practice it evaporates. Because the manifest is metadata, anything that re-encodes the file discards it. Independent testing of six major social platforms found C2PA manifests fully stripped on five of them. Work on image metadata generally has found the large majority of images uploaded to websites lose their metadata in processing, a pattern that predates AI provenance entirely. Tim Bray, testing the ecosystem end to end in 2025, put it plainly: nearly every online photo is delivered either by social media or by publishing software, and in both cases the metadata is routinely stripped. He also found the editing tools themselves mishandling credentials across version mismatches, and Adobe's own inspector failing mid-test.
The result is a system whose coverage is inversely correlated with need. The images that most want a provenance trail are the ones that travel, and travelling is what removes it.
What survives a real workflow
Trace one image through the path a student actually takes. Generate a diagram in a chatbot. Screenshot it because that is faster than downloading. Paste it into your document. Export to PDF. Upload to the submission portal.
The manifest died at step two. The screenshot produced a new file with no C2PA data at all, so every downstream verifier will report nothing found, which is indistinguishable from an image that never had credentials. The pixel watermark, if the generator wrote one, is probably still there through all five steps, because screen capture, format conversion and PDF embedding all preserve enough of the pixel data. Nobody in that chain can read it.
Now the same image with a download instead of a screenshot. The manifest survives into your document folder and then usually dies at the PDF export or the portal upload, depending on what re-encodes the image. Whether your figure carries a verifiable credential comes down to which of two keyboard shortcuts you used and how one server was configured. That is not a foundation to build an academic-integrity process on, and to their credit, nobody serious is trying to.
Can they be removed?
Metadata, trivially, by anyone, on purpose or by accident. That is the whole of the C2PA story.
The pixel watermark is a real research problem, and the answer so far is "harder, but not secure." An open-source project published in 2026 reverse-engineered SynthID on Gemini images by analysing the frequency domain to locate its carrier frequencies, then built a removal pipeline out of regeneration through an autoencoder, elastic deformation, spectral subtraction and JPEG re-encoding. Its README reports around 90% accuracy at detecting the watermark and claims the detector was bypassed on recent Gemini image models with visually lossless output, validated on twenty images.
Take the specifics lightly. Those are the author's own figures on a validation set of twenty, not an independent evaluation, and the target moves with every model release. Take the direction seriously: a mark that a hobbyist project can characterise from the outside is a signal, not a proof. It is the same conclusion Anthropic reached in writing about its own text marks, that a detected mark is evidence content was processed by a model and an absent mark proves nothing whatsoever.
Why this is not your figure's real problem
Here is the part that matters more than any of the above, and it has nothing to do with detection.
A figure in a research paper is an evidence claim. A chart says these are the numbers, a micrograph says this is what the sample looked like, a map says this is where the readings came from. When a generative model produces an image, it produces something that resembles those things without any of them having happened. There were no readings. The problem is not that the picture is undeclared, it is that it is not evidence and is shaped exactly like evidence.
Publishers wrote their policies around that fact rather than around detection. Springer Nature does not permit generative AI images in its journals and books, with a narrow exception where AI is integral to the research and the process is reproducible. Elsevier prohibits using generative AI to create or alter images in submitted manuscripts, allowing help with schematics and flow charts while ruling it out for research images, with a similar exception for AI-based imaging methods. Notice that neither policy promises to catch you. They are prohibitions, enforced by declaration, by forensic examination of the image itself when something looks wrong, and by asking for the underlying data.
Which means the watermark question and the integrity question barely touch. A generated figure with an intact SynthID mark that nobody reads is still a fabricated figure. A hand-drawn diagram with no provenance data at all is still perfectly honest.
What to do about the figures in your paper
- Sort your figures into two piles. Does it present data or observation, or does it explain a concept? A chart, a table image, a photograph of a sample, a map of your sites: evidence. A process diagram, a conceptual schematic, a timeline: explanation. The rules differ sharply and the two piles are easy to tell apart.
- Never generate anything in the evidence pile. Plot your own numbers in a spreadsheet, a notebook or any plotting tool. If the figure is someone else's, reproduce it with a proper credit line and the permission your venue requires, and cite the source it came from.
- Check before you generate anything in the explanation pile. Your instructor or target journal decides, and the safe default is no. Ordinary drawing software takes less time than an exception request, and produces something you can edit later.
- Declare it if you used it. If a generated illustration is permitted, say so in the caption or the methods, naming the tool and the date. That is the same discipline as citing an AI tool in APA, MLA or Chicago, and it costs one sentence.
- Keep the originals. The spreadsheet behind the chart, the raw photograph, the export settings. This is the evidence that answers a question about a figure, and it beats any credential the file may or may not still be carrying.
- Do not screenshot a figure you own. Export it. You lose resolution and you lose whatever provenance the file had, for no benefit at all.
If you want to inspect what your own files are carrying before you submit, the document inspection walkthrough covers reading a C2PA manifest along with the file properties that say considerably more about you than any watermark does.
What CiteOwl puts in a figure
Nothing that was not already yours, and the reason is structural rather than a policy we could quietly change.
CiteOwl does not generate images. Every figure in a CiteOwl document arrives from one of three places: a page or chart extracted from a PDF you imported, an image file you uploaded, or an equation. Equations are rendered server-side from your LaTeX into a path-only SVG, the same bytes in the editor, the PDF preview and the export, so a formula is a deterministic drawing of what you typed rather than model output. The agent can read a figure, which is how it writes an accurate caption, and reading leaves nothing behind.
So there is no pixel watermark, no C2PA manifest and no content credential in anything CiteOwl adds to your document, because there is no generated image for one to describe. What sits in your figures is the material you brought, which is also the only kind of figure that belongs in a paper.