toolcall() ← all concepts

// concept · news · august 2026

Claude marks everything it writes. It can't say it was you.

Anthropic has started watermarking Claude's text. Most coverage gets two things wrong in opposite directions: it cannot point at you, and it cannot clear you either. Both halves are on Anthropic's own page.

// what it actually is

When a model writes, it picks one word at a time — and very often several candidates fit equally well. Anthropic's own example: "The weather today was cold and…". Overcast or grey; the sentence means the same either way. Normally that choice is settled by a random number.

Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick.

The words are still effectively random. But afterwards, someone holding the key can check whether the sequence matches the choices Claude would have made using it. Anthropic's analogy is a good one: a game of Monopoly where the dice rolls come from the digits of pi. The play is unchanged; later, someone who knows pi can tell.

The method is a version of Google DeepMind's SynthID-Text, published in Nature in 2024, in a line going back to a 2022 proposal by Scott Aaronson.

Correction one: it cannot point at you

This is the assumption almost everyone brings to the word "watermark", and Anthropic denies it twice, in different sections of the same page:

Watermarking carries no identifying information and can't be traced to a specific person, organization, or chat.
There's nothing in the watermark, or its key, that would allow anyone to recover any information about the user, their organization, or their chats with Claude.

The key answers exactly one question:

Using our key, one can only answer the question "What is the likelihood this was partly written by Claude?"

So it can say Claude was probably involved. It has no mechanism for saying this person used Claude.

Correction two: it cannot clear you either

This is the half that rarely gets mentioned, and for anyone who has been wrongly accused by a style-guessing "AI detector", it is the one that matters:

It doesn't confirm whether the text was human-written, and it can't tell whether the text was written by a different AI (even if that other AI uses watermarking, it would have a different key; it might also use a different watermarking method altogether).

A clean check is not evidence you wrote something. It is only evidence that this particular model probably didn't. Both halves are true at once, and a version of this story that gives you only one of them is misleading you.

Note what this is not. Existing AI detectors guess from style and do falsely accuse people. This is key-based and answers a much narrower question. Blurring the two makes the situation harder to reason about, not easier.

There is nothing hidden to delete

Nothing is added to the text and there are no hidden characters.

That kills a popular folk remedy — the belief that AI text carries invisible Unicode you can strip by pasting into a plain-text editor. There is no payload to remove. The mark is the pattern of word choices.

On editing, Anthropic is straightforward: "Light editing probably won't remove the watermark completely; a complete rewrite where every word is replaced will" — at which point, of course, the words are yours.

It is weakest exactly where you'd worry most

The mark lives in free choices. Where there is only one right answer, there is nowhere to put a signal — and four consequences follow, all stated on the same page:

# short text "doesn't work well on small samples, where there are fewer word choices and thus less information to go on" # plain facts "Isaac Newton's most famous work was called Principia …" → one right next word. Nothing for the mark to act on. # code "has generally less watermarking than some other forms of text" # text Claude only proofread "nearly all the words are the person's, there's very little (if anything) for the watermark to attach to"

It also costs nothing: no extra tokens, no slowdown, and in DeepMind's testing on live Gemini traffic, no statistically significant difference in quality ratings.

Why now, and who else

This is not an Anthropic decision so much as an industry one:

Anthropic, along with several other major AI model providers and around 190 total signatories, signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026.

As of 2 August 2026 the EU requires providers serving its market to mark AI-generated content, and other major developers signed the same code — so expect their own marks, with their own keys. Anthropic applies it globally "because we don't yet have a durable way to scope it by region".

One practical caveat: nobody outside Anthropic can check anything yet. A detection API is described as coming, with implementation details still being worked out. Files are a separate mechanism entirely — images get C2PA content credentials in metadata, which is signed provenance rather than a watermark, and nothing in the file changes.

Source: Anthropic — "How Claude's text watermark works", 14 August 2026. Every quote above is from that page; fetched 2026-08-17. Related: why models sound sure when they're wrong.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click