Wednesday, September 2, 2026

Anthropic Adds Invisible Watermarks to Claude AI-Generated Text and Signed Metadata to Files

Anthropic has launched a machine-readable content-marking initiative for materials generated by its Claude models. This initiative combines invisible text watermarks with digitally signed provenance metadata for supported files.

This move follows Anthropic’s commitment to the transparency guidelines outlined in Article 50(2) of the European Union AI Act.

Anthropic Adds Invisible Watermarks to Claude

According to the plan, Claude models that are released in the EU on or after August 2, 2026, will support content marking from the outset.

Anthropic has stated that these measures will apply globally, not just to users in Europe, across various Claude services, including the Claude Platform API, Claude, Claude Code, Claude Cowork, and Claude Tag.

Additionally, the company is working on enabling marking capabilities for models released before the August 2026 deadline, which will be covered under the EU law’s transition period.

Anthropic’s approach uses two distinct mechanisms: embedded watermarks for AI-generated text and signed provenance records for supported files.

For text, Claude embeds an invisible watermark directly into the generated output at the model level. Anthropic asserts that this marking does not alter the visible wording, meaning, readability, or quality of the response.

Because the watermark is embedded in the text itself, it can persist when content is copied and pasted to other platforms and may withstand some forms of editing.

This model-level design is crucial for enterprise security and content moderation teams. It ensures that the watermark is applied regardless of how Claude is accessed, whether through an end-user application, a developer API, or partner-hosted infrastructure.

Furthermore, the same text watermarking will be in effect when supported Claude models are accessed through services like AWS, Google Cloud, or Microsoft Foundry. However, support for file-based metadata may vary depending on the capabilities of individual cloud platforms.

For supported file types, including SVG, PNG, and JPG, Claude will attach signed provenance metadata on accordance with the Coalition for Content Provenance and Authenticity (C2PA) open standard.

This C2PA metadata can confirm that Claude processed a file and provide signals that indicate whether the file has been tampered with. If the signed metadata remains intact, downstream tools can determine whether the file has been altered after the provenance record was added.

This mechanism can help security teams, publishers, and investigators distinguish between content with a verifiable Claude-processing signal and files lacking reliable provenance.

Anthropic has warned that a detected mark does not definitively prove that Claude created all underlying content.

For instance, a document may still carry a Claude watermark after the model has proofread, translated, summarized, reformatted, or otherwise processed human-authored material. Additionally, the absence of a watermark does not confirm that a human created the content.

Detection may also fail when the text is heavily edited, paraphrased, translated, mixed with other writing, or is too short to provide a reliable signal. File metadata can also be removed through processes such as conversion, re-saving, taking screenshots, or unsupported workflows.

For developers creating products powered by Claude, Anthropic emphasized the need for them to independently assess their own transparency obligations under the EU AI Act. The company plans to provide further guidance detailing its marking and detection methods in more technical depth.

Stop new phishing & malware before they compromise your business. Integrate live intel from 15K SOCs around the world

Divya
Divya
Divya is a Senior Journalist at GBhackers covering Cyber Attacks, Threats, Breaches, Vulnerabilities and other happenings in the cyber world.

Hot this week

How To Access Dark Web Anonymously and know its Secretive and Mysterious Activities

What is Deep Web The deep web, invisible web, or...

How to Build and Run a Security Operations Center (SOC Guide) – 2023

Today’s Cyber security operations center (CSOC) should have everything...

Russian Hackers Bypass EDR to Deliver a Weaponized TeamViewer Component

TeamViewer's popularity and remote access capabilities make it an...

Web Server Penetration Testing Checklist – 2026

Web server pentesting is performed under three significant categories: identity,...

ATM Penetration Testing – Advanced Testing Methods to Find The Vulnerabilities

ATM Penetration testing, Hackers have found different approaches to...

Threat Intelligence: Definition, Benefits, and Use Cases

Security teams rarely struggle because they lack data. More...

255 Fake Accounts Used to Send Malicious Excel Files to 80,000 Freelancers

A Russian national has been extradited to the United...

Singularity Rootkit Bypasses Elastic Defend eBPF Module Load Detection

Security researcher has disclosed a technique used by the...

Claude AI Develops Working RCE Exploit Against WAGO PLC With Researcher Assistance

Researchers have demonstrated that Anthropic’s Claude AI can assist...

Hackers Exploit LiteLLM Admin API Flaw to Turn Read-Only Access Into Full Server Takeover

Attackers are actively exploiting a critical authorization flaw in...

Palo Alto Networks Acquires Console to Add Agentic AI Workflows to Cortex

Palo Alto Networks has acquired Console, an AI-native platform...

Related Articles

Recent News