Anthropic has introduced a new system that places invisible watermarks on text generated by its Claude models, while supported files receive signed provenance metadata. The policy took effect for Claude models launched from Aug. 2 and applies globally, with no opt-out available. Anthropic says the change is intended to provide greater transparency about the origin of AI-generated content.
The company announced the policy through a support page, explaining that Claude-generated text will contain a hidden statistical signal. Anthropic said that as AI-created material becomes more common, information about where content comes from can provide useful context for people consuming it.
Files receive a different type of identification. When Claude creates supported formats including SVG, PNG and JPG files, Anthropic attaches signed provenance information based on the Coalition for Content Provenance and Authenticity, or C2PA, open standard. The same provenance framework is used by companies including Google and Adobe. A valid signature indicates that Claude processed the file and can also help determine whether the file was altered after generation.
The watermarking system extends across Claude’s various access points, including its API, Claude Code and Claude Cowork, as well as deployments through AWS, Google Cloud and Microsoft Foundry. Anthropic also plans to retrofit older models, although it has not provided a timetable for doing so.
The move comes as the European Union’s AI Act Article 50 became enforceable on Aug. 2. The provision requires providers of generative AI systems to mark AI-generated outputs in machine-readable formats, allowing downstream users, platforms and regulators to identify AI-generated material. Violations can carry fines of up to €15 million or 3% of a company’s total worldwide annual turnover, whichever amount is higher.
Anthropic has signed the EU Act’s Code of Practice on Transparency of AI-generated Content, a voluntary framework that provides a presumption of compliance with the Article 50 requirements. By late July, nearly 200 companies had signed the code, including Microsoft, Google, Meta and OpenAI.
Google has already used watermarking for AI-generated images since 2023 and has expanded the approach to other forms of content, including text, audio and video. Elon Musk’s xAI has not signed the EU code. OpenAI has reportedly possessed technology capable of watermarking ChatGPT text for years but has not deployed it, with concerns including false positives and competitive risks cited as reasons. Anthropic’s decision could add pressure for that approach to change.
Anthropic’s text watermark does not appear as a visible label. Instead, the system statistically influences Claude’s selection of words according to a key controlled by Anthropic. Individual word choices are intended to appear normal, but patterns across a sufficiently large amount of text can become detectable. Because the watermark is embedded within the text itself, Anthropic says it can travel when content is copied and pasted and may survive some forms of editing. The system operates at the model level, meaning the watermark is present regardless of which Claude product or interface generates the material.
The mechanism for files works differently. C2PA creates a digitally signed record that establishes provenance and can be checked to determine whether a file has been modified since Claude generated it. That metadata can disappear if a file is resaved or converted into another format, making the text watermark and file-based provenance system complementary approaches.
Anthropic says it plans to release tools that will allow users and third parties to determine whether text or files contain Claude’s watermark. However, the company has emphasized that detection of a watermark does not necessarily mean Claude originally wrote the underlying material. Someone could use Claude to proofread, translate or summarize work they created themselves, resulting in content carrying a Claude signal.
Likewise, the absence of a watermark cannot conclusively establish that material was written entirely by a human. Signals may not be detectable in older-model output, very short passages or heavily paraphrased text. File provenance can also disappear when metadata is removed through processes such as screenshots or format conversion.
The policy has generated criticism from lawyers, academics, researchers and writers who use Claude to revise or edit material they substantially created themselves. Critics are concerned that work receiving limited AI assistance could subsequently be labeled as AI-generated, potentially creating misleading accusations about authorship.
Questions have also been raised about how Anthropic’s planned detection system will handle disputed results. The company has not yet released the promised detection tools, published accuracy thresholds or established a dispute process for people who believe a detection result is incorrect. Some users have reportedly canceled subscriptions in response to the policy.
There are also concerns about the lack of technical detail surrounding the watermark itself. Anthropic maintains that the system does not reduce the quality or readability of Claude’s output, but it has not yet released enough implementation information for outside parties to independently verify that claim.
Users on X and Reddit have described the policy as problematic, while some software developers have raised separate concerns about cryptographic provenance signatures being attached to AI-generated code. They have questioned whether such markings could affect software workflows or output and have also debated broader questions surrounding credit and attribution.
Anthropic says more technical documentation, detection capabilities and information about extending watermarking to older models are on the way. Until those resources become available, the new system will largely operate out of sight, with its presence becoming apparent primarily when reliable detection tools are available.
Leave a comment