Monday, August 17, 2026
English edition

Development

Claude's new Scarlet Letter watermark is invisible — for now

August 13, 2026 Development Source: Ars Technica

Claude's new Scarlet Letter watermark is invisible — for now

Share this article

The approach described by the EU and implemented by Anthropic is unfortunately trivially easy for bad actors to bypass, while potentially punishing users who trust the system to accurately label their outputs. Text watermarks work by biasing the model’s word choices in a pattern spread across the entire document, only detectable in aggregate by the right tool. The catch is that “invisible” can also mean the model occasionally trades the best word for a slightly worse one, just to keep the signal intact. Anthropic noted that those marks “will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.” But if watermarked text is pasted into another chatbot system that edits the text, the watermark could be destroyed. With image and video content, screenshotting/recording or using any decent metadata editing tool will suffice to remove this information, too. And once Anthropic tells the world how to identify these watermarks, building a system to remove them would be trivial. Further, the potential for misinterpretation seems high; the watermark is not particularly informative. Anthropic explained that a “detected mark provides a signal that content was processed by Claude, but is not fully conclusive.” The only real message the mark sends is that “the content may have been processed by Claude,” Anthropic said, and the mark may even appear on content that was not generated by Claude. On top of this, you have the general public, who may not grasp the difference between processed text and wholly generated text. If the system watermarks human-authored text simply because it was edited in a workflow that touches Claude, suddenly it carries the same denotation as wholly generated text does. And all of this, in a system where the “Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed,” Anthropic said. Ars reached out to Anthropic to see if there’s a timeline for details on detection to be released or results from any testing the company can share assessing the likelihood for false negatives or positives. We also asked if Anthropic could address how its watermarks may conflict with standard editing and other exemptions from the AI Act, but Anthropic did not immediately respond. In the EU, transparency requirements are meant to ensure AI tools like Claude don’t upset “the integrity and trust in the information ecosystem, raising new risks of misinformation and manipulation at scale, fraud, impersonation, and consumer deception.” One EU support article forecasted that the obligations would be the “primary compliance challenge” for many AI firms. “People should know when they are interacting with AI or exposed to AI-generated content,” the European Commission’s guidelines said. “This will help them make informed decisions, calibrate their trust and reliance on AI, and avoid misinformation or deception.” Still, it’s hard to square this with the fact that a wholly generated article on a matter of public interest gets a watermark, but not a reader-facing label, if an editor properly reviews it. AI firms like Anthropic are best positioned to develop watermarking solutions, the EU expects, since AI moves fast and there will be an ongoing “need for new methods and techniques to trace origin of information.” But that largely leaves the societal value of such marks up to tech firms to decide, with the EU only stipulating that “techniques and methods should be sufficiently reliable, interoperable, effective and robust as far as this is technically feasible.” In its post, Anthropic said it plans to continue working on its watermarks and detection methods that meet the EU’s demands. If Claude’s labels fail, the AI Act carries steep penalties for violations, including fines up to 15 million euros or 3 percent of a company’s worldwide annual revenue. Ars Editor-in-Chief Ken Fisher contributed to this report.