Anthropic’s new AI watermark sparks backlash from Claude subscribers


Anthropic said editing can affect whether the watermark can be detected, but acknowledged that minor revisions are unlikely to delete it. — Unsplash

Some Claude users say they are cancelling paid subscriptions after Anthropic announced on Aug 14 that future versions of its AI models will automatically add invisible watermarks to AI-generated text, a move the company says is necessary to comply with the European Union’s AI transparency rules that went into effect Aug 2.

Dozens of Claude subscribers on X criticised the decision, expressing concerns that content generated or revised with Claude could still be identified as AI-assisted even after the text is reworked, reformatted, or incorporated into another piece of writing.

Some critics described the policy as “a conspiracy against innocent Claude users.” Others characterised the watermark as a modern-day “scarlet letter,” a mark that could remain identifiable long after any piece of AI-assisted text leaves the platform.

“Who will get caught? You,” a user named visionode wrote in a lengthy post on Reddit. “The student who used Claude to reorganise a paragraph. The journalist who asked the AI to summarise a two-hundred-page transcript. The writer who had creative block and asked for synonyms. Those guys come out of the process with a digital tattoo on their forehead.”

Anthropic disputes that characterisation. The company said the watermark “carries no identifying information and can’t be traced to a specific person, organisation, or chat.”

The new watermark, based on Google DeepMind’s SynthID-Text technology, allows AI-generated text to be identified without a visible label. Instead, it embeds a statistical pattern into ordinary word choices made by the model. The pattern is invisible to readers, Anthropic said, but can be picked up by someone using a verification key to determine “the likelihood that Claude was involved in writing the text.”

The company said editing can affect whether the watermark can be detected, but acknowledged that minor revisions are unlikely to delete it. “Light editing probably won’t remove the watermark completely,” Anthropic said. “A complete rewrite where every word is replaced will.”

The company said it also plans to release a watermark-detection API, a cloud-hosted service that would allow third parties to identify Claude-generated text. The company did not say when it would become available.

The controversy coincides with Anthropic’s decision not to release its experimental Model 2 system, which the company said has shown “a noticeable improvement” over Mythos, its current flagship system, on “many tasks relevant to internal use.”

“As part of our standard R&D process, we internally train and evaluate many different exploratory versions of models that we don’t intend to release,” Anthropic said in its August risk report. “Model 2 is one of these.” – Inc./Tribune News Service

Follow us on our official WhatsApp channel for breaking news alerts and key updates!

Others Also Read