Anthropic will start embedding invisible watermarks in text and images generated by Claude

Anthropic will start embedding invisible watermarks in text and images generated by Claude

Anthropic has revealed that Claude will quietly embed invisible watermarks into AI generated text, with the markers potentially remaining even when the text is copied and pasted elsewhere. The company will also add digitally signed provenance metadata to supported image files using the C2PA standard, as part of its compliance with new transparency requirements under the EU AI Act that took effect on August 2.

The markers will be applied at the model level across Claude products and interfaces, including Claude, Claude Code, Claude Cowork, Claude Tag, the Claude Platform API, and third party services such as AWS, Google Cloud, and Microsoft Foundry. New Claude models will include the system from launch, while Anthropic is still working to add it to existing models. The rollout will be global rather than limited to the European Union.

Anthropic is also developing tools and technical documentation so users and third parties can detect the markers. It's also worth noting that a detected marker only indicates that Claude may have processed the content, not that it originally wrote it. This could apply even if the user only used Claude for minor tasks like correcting grammar or translating text. However, the system is not definitive, since C2PA metadata can be removed, while heavy editing, paraphrasing, translation, or short passages may make text watermarks harder to detect.

by Mauricio B. Holguin

Ro
J6
shazmatazMyrano
RobertasTa found this interesting
Add as a preferred source on Google
Claude iconClaude
  226
  • ...

AI assistant for chat and API, performing summarization, search, writing, Q&A, and coding with high reliability, customizability, privacy, and reduced risks.

Other apps mentioned in this article

Comments

RobertasTa
0

The article already hints at it, but the asymmetry deserves more attention: the marker means "Claude touched this", yet everyone will read a detector hit as "AI wrote this". Fix one comma with Claude and your whole letter carries the mark. I suspect the practical effect won't be catching AI slop — it will be honest people having to prove they wrote their own words.

2 replies
BorisF

Claude is just one of many AIs. Unless all of them implement the same protocol, I do not see it as absolute proof of AI none-involvement.

RobertasTa

Agreed — and that cuts both ways. The mark can only ever prove one thing: "Claude was somewhere in the pipeline." Its absence proves nothing — any other model leaves no mark, and a decent paraphrase likely strips it anyway. And its presence says nothing about how much of the text is actually human. A detector that can neither clear you nor convict you is a strange foundation for any rule about authorship.

Darlene Sonalder
0

Wait Claude can generate images? I didn't knew that x)

Creative_joe
0

https://www.mixfont.com/experiments/decoy-font The most efficient Claude model cannot even quickly pass the test above, whereas the free version of Mistral, Gemini and GPT perform better with some tinkering. Utter joke.

2 replies
Creative_joe

I mean the test from the hyperlink, ugh... I can't edit comments under tech news.

Mauricio B. Holguin

Yeah sorry, it’s on the list 😅

BorisF
2

I am puzzled. How do you keep watermarks if somebody performs "copy as plain text"?

5 replies
fabriziob

The "watermark" is based on the composition of the phrases themselves. To keep the signal, I predict we will see longer and longer responses, which coincidentally is what I hate the most of all the things AI does.

That also means that copy-pasting only short portions of a response and possibly recombining them with fuck it up.

Dino C

perhaps specia lsymbols that show as plain text in reality idk

BorisF

Do you mean each letter in a slightly different font that makes it almost indistinguishable? I do not know any special symbols that can pass as letters or numbers. That's why they are called "special."

BorisF

In my understanding, and it may be outdated, "watermark" means some invisible to the naked eye picture or QR code embedded in a video or another picture. If somebody is using the composition of the phrases themselves, it should be called "fingerprinting." "Fingerprinting" and "watermarking" serve similar purposes, but they are different processes.

mimumu

the text the models outputs has a statistical unique pattern associated with claude, the words and/or letter have certain order or certain repetitions that those who know the pattern (ai detectors) can identify as calude's. its just like certain people uses certain words or phrases, or talk in a certain way and just with that you can identify them, except that is not plainly visible and instead is on a deeper level that only those who know a certain formula or algorithm can parse.

Gu