HomeArticle

Claude launches the "Invisible Watermark" feature amid widespread backlash, with the mechanism fully embedded in all new models to mark every piece of generated text.

量子位2026-08-11 13:13
It takes effect globally, and cannot be easily cleaned even through rewriting.

Another bizarre move from Company A has once again sparked public outrage across the entire internet.

Just now, Anthropic announced: The new Claude model will embed invisible watermarks in generated text.

Yes, these marks are not stored in metadata, they will directly become part of the text itself.

It will be copied and pasted along with the text to other locations, and may still remain intact even after a certain degree of editing.

It's such a sticky, unremovable nuisance...

Netizens have all voiced their criticism collectively.

The direct driving force behind this move is the EU AI Act that Anthropic has signed, and this mechanism will apply to all models released on or after August 2, 2026.

For the Claude models that have already been launched, Anthropic is also researching how to retroactively add this watermark mechanism.

Oh, by the way —

This measure will be implemented globally, not limited to the European Union.

The great "era of watermarks" is here.

Text now also comes with watermarks

Just now, Anthropic signed the EU AI Act's Code of Conduct on Transparency for AI-Generated Content.

This code was drafted by independent experts and assessed by the European Commission and the AI Board. Around 190 institutions have signed it so far, including Google, Meta, Microsoft, OpenAI and others.

For Claude, this specifically involves two major measures.

First, text watermarks.

When supported Claude models generate text, they will directly weave a kind of invisible mark that cannot be seen by the naked eye into the text.

Yes, it is completely invisible. You cannot see it, and it does not change the meaning, quality or readability of the response.

Worse still, you cannot get rid of it easily.

The watermark is part of the text, it will be carried over when you copy and paste the content, and it will not be removed even after editing in most cases.

Second, signed metadata.

When Claude generates supported file types such as .svg, .png, .jpg, it will attach digitally signed source metadata.

This set of metadata follows the C2PA open standard, which is jointly promoted by Adobe, OpenAI, Google and other parties. It can indicate that the file has been processed by Claude, and can also detect whether the file has been tampered with.

This coverage applies to the entire product line, including API, Claude, Claude Code, Claude Cowork, and Claude Tag.

Cloud customers who access Claude through AWS, Google Cloud, and Microsoft Foundry are also covered by this rule.

Moreover, there is no regional restriction, it is implemented uniformly across the globe.

At the same time, Anthropic stated that it is developing detection tools, so that users and third parties can verify whether a piece of text or a file carries a Claude mark.

The company will release the specific technical documents at a later date.

However, Anthropic has also clearly stated the limitations of this mechanism.

The detection of the mark only indicates that the content has been processed by Claude, and cannot confirm that Claude is the original author.

After all, users often use Claude for proofreading, translation, summarization, and format conversion, and the underlying ideas and data may come from humans; marked content may also be modified, excerpted, or mixed with other materials afterwards.

Conversely, the absence of a detected mark does not necessarily mean that the content is not AI-generated.

It may come from an older model, may have been heavily edited or rewritten, may be too short a paragraph to be recognized, or the metadata may have been stripped due to format conversion or screenshot capture...

Everyone is a suspected object anyway...

Up to now, Anthropic has not disclosed the details of the watermark algorithm.

In other words, no one knows what the "invisible ink" looks like, nor what else Company A can do with it.

Company A's secret marking trick

So, is there really such a magic trick that makes the watermark in text completely invisible?

Yes.

And the tricky operations may be far more than you imagine.

The most well-known one is something called Unicode.

You can understand it as issuing an ID card for almost every character and symbol in the world: A corresponds to one Unicode number, a Chinese character corresponds to another Unicode number,

also corresponds to a Unicode number.

But in addition, there is a set of characters that are basically invisible to the naked eye but can be distinguished by computers.

For example, the simplest word "Hello", the underlying code of one version may be H-e-l-l-o, while the other is H-e-invisible character-l-l-o.

They look exactly the same to the human eye, but when the program reads the Unicode code points, it will find that the second paragraph has one more special character.

Amazing.

I don't know if you still remember that Company A has long "followed the trend" and used the Unicode trick before.

On June 30, a Reddit user reverse-engineered Claude Code and discovered a shocking truth —

A group of trojan programs are hidden in the packaged files of Claude Code. If it detects that you are a user from China, it will add two strokes of invisible ink to your system prompt — "Today’s date is 2026-06-30".

1. The hyphen in the date format is changed to a slash, marking that the China time zone has been detected.

2. The apostrophe in "Today’s" is quietly replaced with three characters that have different Unicode code points but are completely identical in appearance.

As a result, in any editor or terminal, this line of text looks exactly the same as the normal version. But in the program, you have already been identified and locked.

Yes, Company A has been using Unicode to add hidden marks in Claude Code all along.

The company has not yet disclosed whether Claude's new watermark uses this method.

From this perspective, text watermarks don't seem so bad, right? At least you won't get your account banned...

Besides, various AI detection tools on the market are uneven in quality now, many of them charge dozens of yuan for a single detection, and output an AI probability number with no logical basis. Having an official marking channel can at least save you a lot of unnecessary expenses.

The main reason why netizens are so angry is that Company A already has a bad track record before.

But think carefully, AIGC content does need a reliable detection system, just like OpenAI adds C2PA metadata and SynthID watermarks to images.

However, will embedding watermarks by restricting the model's sampling distribution really not affect the generation quality?

Company A's statement is —

No.

I remain skeptical.

After all, since the 4.6 version, Claude's outputs have become increasingly unnatural and unlike normal human language.

One More Thing

In any case, the "Avengers" targeting Anthropic has already taken action!

YC founder Paul Graham has already come up with a brand new startup idea, just waiting for ambitious people to join in!

Build a third-party rephrasing function, reword the content, and remove the watermark while retaining the original meaning.

Of course, a more lightweight solution that is available to everyone right now has also emerged —

Kimi, remove that invisible watermark.

References:

[1]https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content

[2]https://x.com/atharvabuilds/status/2086920300441268579?s=20

This article is from the WeChat official account "QbitAI", Author: Jay, published with authorization from 36Kr.