METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Anthropic to release watermark API letting third parties verify Claude-written text

Detection tech adapts Google's SynthID, but struggles with code and short text — EU regulation driving global rollout

Anthropic to release watermark API letting third parties verify Claude-written text

Image: METAL

Summary

  • Anthropic will soon launch a watermark detection API that lets third-party developers verify whether text was written by Claude
  • The technology adapts Google DeepMind's SynthID Text method, published in Nature in 2024, by adjusting randomness in word selection
  • Detection accuracy drops for short passages, fact-heavy sentences, code, and text that has been fully rewritten
Video from the source

Text written by Claude can now be verified by others

Anthropic is preparing to release a watermark detection API that third-party developers can embed in their own apps. Until now, the notable step was Claude leaving an invisible trace in the text it generates; the significance of this announcement is that companies other than Anthropic will now be able to read that trace. Services that need to verify whether text came from Claude — such as school plagiarism checkers or AI-detection features on content platforms — will be able to plug in this API.

An approach built on SynthID

The technology is reportedly a variation of SynthID Text, the method Google DeepMind published in Nature in 2024. The underlying principle involves finely adjusting the randomness factor that comes into play when selecting each word, embedding a traceable pattern within the sentence. Anthropic explained that this process "does not affect the content, creativity, or readability of text written by Claude." To a human reader, the sentences look no different from ordinary writing, but running them through the detection API can reveal the likelihood that Claude was involved.

Where it doesn't work

However, this method doesn't function equally well across all types of text. Short passages or sentences that simply list facts offer few alternative word choices, making it hard to embed a pattern. The same goes for code. Pure edits where a human manually adjusts word by word don't retain the watermark, since the words are human-selected. Conversely, Anthropic says watermarks tend to persist well in translations, since in translation Claude ultimately selects every word. According to Anthropic's published FAQ, a substantial rewrite of the original text can also erase the watermark.

The watermark also only indicates the likelihood that Claude was involved in writing the text — it cannot distinguish whether Claude wrote the entire piece or only touched up part of it. Nor can it determine whether a piece was written by a human or by another company's AI.

A different approach from pattern-scanning tools

Anthropic's method is fundamentally different from existing AI detection tools like Pangram. Services like Pangram don't have access to Anthropic's watermark key, so instead they statistically scan for AI-characteristic writing style or commonly used expression patterns. Watermark detection, by contrast, checks for traces that Claude actually left behind, which in principle can produce more reliable results.

EU regulation is driving global adoption

The EU AI Act is the backdrop for Anthropic's adoption of watermarking. In July, Anthropic signed the EU's Code of Practice on transparency for AI-generated content, alongside roughly 190 other signatories. Since there's no practical technical way to restrict the feature by region, watermarking applies simultaneously worldwide, not just in the EU. All Claude models released after August 2, 2025 support watermarking by default, and earlier models will gain support progressively over the coming months. For non-text files, Anthropic uses the open standard C2PA, which attaches provenance metadata without altering the file itself.

Anthropic previously announced on August 11 that it would make watermarking mandatory across Claude's outputs; this detection API represents the next step — opening up the ability to read that trace outside of Anthropic itself.

How it will work in practice

This API is aimed at developers, not general users. A company wanting to add Claude-verification capability to its service would call Anthropic's detection API and receive a response indicating whether the text carries Claude's watermark pattern. The exact launch date and application process haven't been specified yet, but the likely use cases can be inferred. For example, schools or academic journals could use it as a first-pass filter to check for AI involvement in assignments or papers, and news or community platforms could use it to label AI-generated posts. Hiring platforms could also use it as a reference signal for detecting AI-written cover letters.

Editor's view

Embedding a watermark and letting others verify it are two entirely different decisions. When Anthropic announced on August 11 that it would apply watermarking globally, that was a declaration of "we will leave a trace." This API is a declaration of "we will hand others the key to check that trace." That's a high-risk decision for Anthropic. It effectively opens the door for outside parties to verify how and where its models were used, and a buildup of false positives or evasion cases could actually erode trust rather than build it. Still, it's reasonable to assume that the external pressure of the EU AI Act played a major role in this choice. Without the regulation, there would have been little reason to voluntarily shoulder a transparency burden that competitors don't carry.

OpenAI and Google have each experimented with their own approaches to marking AI-generated text, but opening up detection authority itself to third parties has been rare. Google's SynthID has mostly been used within its own ecosystem, whereas Anthropic has taken that idea and opened it externally in the form of an API. Anyone who has worked hands-on with AI-detection tools knows that statistical style-analysis-based detectors have always struggled with false-positive rates. Watermark-based detection is inherently more accurate in principle, but the catch is that the exceptions — code, short text, and full rewrites — are exactly the cases most commonly encountered in practice. Ultimately, even once this API launches, a "not Claude" result should not be understood to mean "definitely written by a human."

Domestic companies aren't likely to adopt this API immediately, but content platforms and educational institutions should start considering it now. That said, it would be safer to use watermark detection results as a supplementary signal alongside existing style-analysis tools, rather than as standalone evidence. In the coming weeks, it's likely that news will emerge of OpenAI or Google considering similar externally accessible detection options. This pattern — where regulation moves one player and the rest follow — has already played out repeatedly.

Comments