Connect with us

AI

Claude Stamps Text Worldwide and Locks the Detector

New Claude models watermark text worldwide for the EU AI Act, while Anthropic’s detector stays a private preview for regulators, schools, and media.

Published

on

Anthropic now weaves an invisible watermark into text from new Claude models, then withholds the detector from ordinary users. The stamp is meant to satisfy the EU AI Act. It still rides along in chats far outside Europe.

Schools, police, and publishers can ask whether Claude touched a passage. The person who pasted that passage into an email cannot.

Invisible Marks on Every New Claude Reply

On 2 August 2026, new Claude models began marking output at launch. Anthropic’s help article says the company marks AI-generated content worldwide, not only in the EU. The help page lists Fable 5.1 and Mythos 5.1 as models that already support the scheme. Older models are still being updated.

The mark sits at the model level, so the product surface does not get a choice. It applies across Claude Platform (the API), Claude, Claude Code, Claude Cowork, and Claude Tag. Anthropic says it also applies when those models run on AWS, Google Cloud, or Microsoft Foundry.

Text carries the hidden pattern. Supported files such as.png,.jpg, and.svg get a different treatment: a digitally signed C2PA origin note in the file’s metadata. That note says Claude made or processed the file. It does not name the user. Strip the file, screenshot it, or convert the format, and that label can vanish. The text mark is harder to flick off, because it is not a footer or a hidden character. It is in the words.

Anthropic says the scheme adds no extra tokens, does not raise the price, and does not slow the model in any way users would notice. It also says the watermark carries no identifying information and cannot be traced to a person, a company, or a chat. The argument against it was never really about a secret ID. It is about who can read the stamp, and what they will do with a hit.

Who Gets the Detector, and Who Does Not?

A watermark you cannot test is a rumour with a logo on it. Anthropic’s 1 September 2026 update says the detection API is a private preview. Access goes to the groups EU law points toward: regulators, law enforcement, media, fact-checkers, independent researchers, educational organisations, and EU civil society groups. Enterprises that must verify marks for their own compliance can also ask. Everyone else is told to wait. The company says it plans to widen access over time.

We are releasing a detection API in private preview. It is currently available to eligible organizations as required under EU law (such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups).

Anthropic, How Claude’s Text Watermark Works, 14 August 2026

That split is the product. Transparency, in this design, is a window that opens toward institutions. It does not open toward the writer. A student cannot run their own draft through Claude’s key and see the score before a university does. A freelancer cannot check a client memo. A reporter outside the preview list cannot test a quote.

On 9 September 2026, Alexander Nemecek, Vipin Chaudhary, and Erman Ayday put that gap in a paper titled “Watermarks Without Verification.” Their point is blunt: the public cannot test Anthropic’s quality claims, or the mark’s toughness, because no public tool can query the deployed system. They ran the open-source SynthID-Text code on two open-weight models instead. On prose, they wrote, the measured effect did not exceed a change of sampling seed. On code, detection stayed near chance. That is not a verdict on Claude’s private key. It is a verdict on a rollout that asks the world to trust a lockbox.

Anthropic already warns against treating a hit as a gavel. A detected mark, it says, only answers how likely it is that Claude was involved at some point. It does not prove a human never wrote the piece. It cannot spot text from another model, even one that watermarks with a different key. Short passages give the test too little to chew on. Factual lines with one right next word give it even less. None of those caveats will be in the subject line of a disciplinary email.

The Mark Lives in Word Choice

Anthropic set out how Claude’s text watermark works on 14 August 2026. Language models pick one word at a time from a list of decent next options. Plenty of those picks are toss-ups. The company walks through a weather line to show where the stamp hides.

Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses.

Anthropic, How Claude’s Text Watermark Works, 14 August 2026

Watermarking changes the source of the randomness that breaks those ties. Instead of an ordinary random number, the model uses a key plus a few earlier words. The sentence still reads like English. Anyone who holds the key can later ask whether the sequence of choices matches what that key would have produced, and attach a probability that Claude wrote it.

Nothing is inserted into the file. There are no hidden Unicode tricks. Claude’s method is a version of SynthID-Text, published by Google DeepMind in Nature in 2024, in a family of ideas that goes back to a 2022 proposal by Scott Aaronson. Anthropic points to DeepMind’s live test on Gemini traffic, which found no statistically significant gap in thumbs-up and thumbs-down ratings between watermarked and plain output. Human raters in a side-by-side study, Anthropic adds, saw no quality difference either.

The stamp is sparse where the model has no real choice. After “Isaac Newton’s most famous work was called Principia,” the next word has to be “Mathematica.” Code is the same kind of trap. If the next token has to be exact or the program breaks, the nudge is skipped. Comments and other toss-up phrases can still carry a little signal. Anthropic says the effect on the code that actually runs is negligible.

Google Marks Gemini. OpenAI Omits ChatGPT Text.

Anthropic is not the only lab under the same law. In July 2026 it signed the EU Code of Practice on Transparency of AI-Generated Content, alongside other major providers and about 190 organisations in total, according to the Commission’s 31 July 2026 tally. That list is signatures, not unique names: 82 on the provider section and 152 on the deployer section, with some firms on both. Named provider-section examples include Anthropic, Google, OpenAI, Meta, Microsoft, Mistral, Cohere, and Aleph Alpha. Signing is voluntary. Article 50’s marking duty is not.

What each lab actually ships still splits on text. Google has applied SynthID watermarking for generated text in the Gemini app and web experience since 2024. People can also ask Gemini whether an image, video, or audio clip carries a Google mark, and a SynthID Detector portal is being tested with journalists. OpenAI’s provenance page, updated 31 July 2026, describes C2PA Content Credentials and SynthID watermarks on ChatGPT images, then extends SynthID to supported audio from ChatGPT and the API. It does not describe a production watermark on ChatGPT text. A public OpenAI verification tool in preview checks images and audio, not essays.

HOW THREE LABS MARK OUTPUT

Lab Text Files and media Who can check
Anthropic (Claude) SynthID-Text version on models launched on or after 2 August 2026 C2PA on.png,.jpg,.svg Private preview for named institutions and obligated firms
Google (Gemini) SynthID-Text in the Gemini app and web experience since 2024 SynthID on image, audio, and video Ask Gemini on media; Detector portal in journalist tests
OpenAI (ChatGPT) No production text watermark described C2PA plus SynthID on images; SynthID on supported audio from 31 July 2026 Public preview tool for images and audio

So the first big consumer assistant to put a statistical stamp on everyday prose, then hide the scanner, is Claude. Gemini already marked text. ChatGPT still has a cleaner page for anyone who wants unmarked words and does not care about image credentials. Open weights sitting on a laptop were never in this deal. The honest lab becomes the easy lab to catch. That is a strange prize for going first.

Keep the Words, Keep the Stamp

Because the signal is the wording, copy and paste does not wash it out. Paste a Claude answer into Word, Notion, or a CMS and the words are still the words. Anthropic’s own rule of thumb is simple, and it is also the loophole.

WHAT HAPPENS TO THE SIGNAL

  • Copy and paste: The mark travels with the words, because nothing extra was attached to the file.
  • Light edits: Anthropic says they probably will not remove the watermark completely.
  • Grammar-only proofreading: Almost all the words are still the person’s, so there may be too little for a detector to fire.
  • Translation by Claude: Every word is Claude’s choice, so the passage carries a watermark.
  • Full rewrite: Replace every word and the mark goes; Anthropic adds that the text may no longer be AI-generated in any useful sense.
  • Code: Less room for the stamp, and formatters can thin it further; comments can still hold a little pattern.

That map is why hybrid writing is the messy case. A heavy Claude edit can look, to a detector, like Claude wrote the piece. The company says the tool cannot tell those two jobs apart. Light polish can go either way, which is already the fight around watermarks that survive light AI edits. A native speaker who asks for a comma check may walk away clean. Someone who needs Claude to put a finished argument into English walks away fully marked, because a translation is not a nudge. It is a new sentence, word by word.

The people with the most to lose did not design this. A determined user can kill the stamp by running the draft through another model and keeping the rewrite. Ordinary users who never heard of SynthID will not. Unmarked open-weight models, and labs that have not shipped a text mark, become the quiet off-ramp. Screening tools will then start treating the absence of a Claude hit as a human tell. All it proves is that the writer used a system that does not report itself.

A Brussels Rule, Applied Worldwide

Article 50 of the EU AI Act requires providers of generative systems to mark synthetic audio, image, video, and text in a machine-readable way, and to make that output detectable as artificial. New systems placed on the market on or after 2 August 2026 had to comply from that day. Anthropic could have tried to watermark only EU traffic. It did not.

“We’re applying watermarking globally at launch because we don’t yet have a durable way to scope it by region,” the company wrote. One inference path is cheaper than two. A rule written for one market becomes the default in every Claude chat, including API jobs and cloud resale. There is no regional off switch in the docs. Paying customers get the same stamp as free users.

THE MARKING CALENDAR

  1. 2 August 2026: Article 50 takes effect, and new Claude models mark text from launch.
  2. 14 August 2026: Anthropic publishes the SynthID-Text method note and the detector plan.
  3. 1 September 2026: The company updates the detection API as a private preview for named groups.
  4. 2 December 2026: Generative systems already on the market before 2 August 2026 must meet the marking duty.

That last date is the Article 50 marking grace period for older systems. Anthropic says it is rolling watermarking onto pre-August models over the coming months. Content produced before 2 August 2026 does not need a retroactive label. Everything those older models emit after the grace window closes is supposed to carry a mark.

Until 2 December 2026, a slice of Claude output can still leave the building without the new stamp, depending on which model wrote it. After that date, the unmarked path inside Anthropic’s own lineup is supposed to shrink. The detector list may grow. For now, the feature billed as public transparency remains a mark on the words and a key held by someone else.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending