// what it actually is
When a model writes, it picks one word at a time — and very often several candidates fit equally well. Anthropic's own example: "The weather today was cold and…". Overcast or grey; the sentence means the same either way. Normally that choice is settled by a random number.
The words are still effectively random. But afterwards, someone holding the key can check whether the sequence matches the choices Claude would have made using it. Anthropic's analogy is a good one: a game of Monopoly where the dice rolls come from the digits of pi. The play is unchanged; later, someone who knows pi can tell.
The method is a version of Google DeepMind's SynthID-Text, published in Nature in 2024, in a line going back to a 2022 proposal by Scott Aaronson.
Correction one: it cannot point at you
This is the assumption almost everyone brings to the word "watermark", and Anthropic denies it twice, in different sections of the same page:
The key answers exactly one question:
So it can say Claude was probably involved. It has no mechanism for saying this person used Claude.
Correction two: it cannot clear you either
This is the half that rarely gets mentioned, and for anyone who has been wrongly accused by a style-guessing "AI detector", it is the one that matters:
A clean check is not evidence you wrote something. It is only evidence that this particular model probably didn't. Both halves are true at once, and a version of this story that gives you only one of them is misleading you.
Note what this is not. Existing AI detectors guess from style and do falsely accuse people. This is key-based and answers a much narrower question. Blurring the two makes the situation harder to reason about, not easier.
There is nothing hidden to delete
That kills a popular folk remedy — the belief that AI text carries invisible Unicode you can strip by pasting into a plain-text editor. There is no payload to remove. The mark is the pattern of word choices.
On editing, Anthropic is straightforward: "Light editing probably won't remove the watermark completely; a complete rewrite where every word is replaced will" — at which point, of course, the words are yours.
It is weakest exactly where you'd worry most
The mark lives in free choices. Where there is only one right answer, there is nowhere to put a signal — and four consequences follow, all stated on the same page:
It also costs nothing: no extra tokens, no slowdown, and in DeepMind's testing on live Gemini traffic, no statistically significant difference in quality ratings.
Why now, and who else
This is not an Anthropic decision so much as an industry one:
As of 2 August 2026 the EU requires providers serving its market to mark AI-generated content, and other major developers signed the same code — so expect their own marks, with their own keys. Anthropic applies it globally "because we don't yet have a durable way to scope it by region".
One practical caveat: nobody outside Anthropic can check anything yet. A detection API is described as coming, with implementation details still being worked out. Files are a separate mechanism entirely — images get C2PA content credentials in metadata, which is signed provenance rather than a watermark, and nothing in the file changes.
Source: Anthropic — "How Claude's text watermark works", 14 August 2026. Every quote above is from that page; fetched 2026-08-17. Related: why models sound sure when they're wrong.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
