
Summary
- Anthropic published an explainer on August 14 confirming that Claude's text watermark belongs to the same family of technology as Google DeepMind's SynthID. Because it only alters how words are chosen rather than adding characters, pricing, speed, and writing quality remain unchanged.
- The marker only appears where Claude "could have written it differently." It is rarely present in factual sentences or code, where there's only one correct answer, and the same holds true when Claude only lightly edits human-written text. Translations, however, are entirely subject to the marker, and short pieces of text are hard to detect.
- With Article 50 of the EU AI Act taking effect on August 2, roughly 190 companies have been pulled in the same direction. South Korea's AI Basic Act leaves machine-readable labeling as optional rather than mandatory, while Google automatically applies visible watermarks in South Korea, India, and Vietnam.
First, the bottom line — what happens to my writing
The marker Anthropic turned on isn't a matter of secretly inserting characters. It embeds a hidden rule only in the places where Claude could have used any word when choosing its wording. We'll walk through the mechanics later — first, let's look at what happens to your writing.
The watermark only appears where Claude could have written it differently. That single sentence explains almost everything. Where there was no room for choice, there's nowhere to plant a marker either.
- If you hand off an entire draft, it sticks. The more words Claude decided on, the more places there are for the marker to attach.
- If you have it polish your own writing, it barely sticks. Most of the words are still yours.
- Translation is entirely covered. In a translation, Claude chooses every single word anew.
- It's rare in factual sentences. In "Isaac Newton's major work is the Principia ___," the blank has exactly one correct answer: "Mathematica." Any other word simply makes the sentence wrong.
- It barely appears in code. Anthropic doesn't apply it at all in places requiring exact output. However, it can show up in places where wording is flexible, such as the code comments Anthropic cites as an example. The claim that "watermarks are embedded in code" is an exaggeration.
- It can't be detected in short text. Since the number of word choices is essentially the amount of signal, confidence rises the longer the text is.
Anthropic's own example captures this principle neatly. "2 + 2 = ___" has only one correct answer: 4. But in the middle of a retelling of George Orwell's 1984, 5 might be the "correct" answer. The moment context narrows the answer to one, the watermark has no room to exist.
An important limitation is worth noting here too. The watermark cannot distinguish between "Claude wrote this" and "Claude heavily edited this." Conversely, the absence of a marker doesn't mean AI wasn't used — other companies' AI systems use different keys. Ownership is a separate matter as well. Anthropic has explicitly stated that the watermark says nothing about copyright or authorship, and it does not change any rights under its terms of service.
Images and files are a different story — a point many English-language outlets have blurred. PNG, JPG, and SVG files don't get a pixel-level watermark; instead they receive a C2PA Content Credential, essentially a tamper-resistant electronic tag attached to the file's metadata. It's an open standard also used by camera makers and photo-editing software, so the image itself is left untouched. But because it's a tag, it carries the weakness of being removable simply by stripping it out.
Not adding characters — changing the dice
When an AI chooses its next word, it picks probabilistically among several candidates. In "The weather today is cold and ___," "sugary" wouldn't fit, but "gloomy" and "overcast" both work naturally. The meaning doesn't change no matter which one is chosen. Such choices where either option is fine occur hundreds of times in a long piece of text.
Watermarking doesn't force a change in the outcome of these choices. What it changes is the "dice" used to make the choice. Instead of using ordinary randomness, it derives a number from a secret key combined with the preceding few words, and uses that number to determine the next word. On the surface it still looks random, but anyone holding the key can trace back and find a regularity that can't be explained by chance.
Anthropic's own analogy captures this precisely. Imagine that instead of rolling dice in a board game, you move your piece by reading off the digits of pi in sequence. During the game, this is indistinguishable from dice rolls. But looking back at the full record afterward, someone who knows pi can calculate whether the game used it.
One misconception should be cleared up here. This method doesn't bias Claude toward always choosing "gloomy" as a matter of taste. Anthropic has explicitly stated that it doesn't force in obscure synonyms Claude wouldn't normally use. It's not about reshuffling the ranking of candidates — it's about changing the seed of the lottery.
The underlying technology is Google DeepMind's SynthID-Text. It was published in Nature in October 2024 under the title "Scalable watermarking for identifying large language model outputs." It uses tournament sampling, pitting candidate words against each other head-to-head, tournament-style (30 rounds by default) to determine a final winner — with the secret key hidden in the rule that decides who wins each matchup. Under settings that preserve each word's original selection probability, writing quality is not degraded in principle. DeepMind compared user feedback across about 20 million real Gemini responses and reported a difference of just 0.01 percentage points in likes and 0.02 points in dislikes — effectively no difference. Anthropic has also stated that its internal testing showed no impact.
Completely different from commercial "AI detectors"
This is where confusion is most common. Services like GPTZero or Pangram have no key. So they look at writing habits instead. Anthropic itself gives an amusing example: AI models are unusually fond of constructions like "this isn't X, it's Y," and use the adverb "quietly" far more often than humans do.
Watermark detection works the opposite way. Since it doesn't look at style but computes using a key, well-written text isn't suspected any more than poorly written text. But the catch is that only Anthropic, which holds the key, can do this. Neither teachers, companies, nor writers themselves currently have any way to check — the detection API remains a "coming soon" promise.
Can it be erased? — academics disagree
The common belief that "paraphrasing erases it" is only half true.
In 2023, a research team led by Krishna used a dedicated 11-billion-parameter AI to fully rewrite text, driving the accuracy of a leading detector down from 70.3% to 4.6%. But a team at the University of Maryland led by Kirchenbauer reached the opposite conclusion in a June 2023 paper (ICLR 2024). Even after heavy human rewriting, they found that when averaging across roughly 800 tokens (the "word fragments" AI models process text in), the watermark could still be detected — even under a strict threshold of one false positive in 100,000. This is because fragments of the original word groupings survive scattered throughout the rewritten text.
The axis on which these two conclusions diverge is text length. Short text erases the watermark; long text preserves it. Anthropic's own explanation stays within these bounds: "Light editing won't fully remove it, but a complete rewrite that changes every word will. Of course, if you've gone that far, it's debatable whether the result can even be called AI-written."
There's a more fundamental problem too. In 2024, researchers at ETH Zurich reverse-engineered the watermark rules simply by repeatedly querying the API, successfully performing both removal and forgery. Forgery is the more troubling of the two — maliciously planting a fake signal in human-written text to frame it as AI-generated is far harder to deal with than simply failing to catch AI text. GPTZero's Alex Cui also stated that "it won't survive strong paraphrasing that attacks both word choice and sentence structure simultaneously" — meaning that overhauling both wording and structure at once defeats it.
Why now — a technology nearly four years old meets regulation
Watermarking isn't a sudden idea. A technology that had been circulating in academia met regulation and surged forward all at once.
November 2022. Theoretical computer scientist Scott Aaronson, then on leave and working at OpenAI, revealed his main project in a talk at UT Austin. The idea: pick words using randomness generated from a secret rule known only to OpenAI — the prototype of the very mechanism Anthropic has now turned on. He also acknowledged that the signal becomes detectable after a few hundred tokens, but that rewriting with another AI would break it.
January 2023. A team at the University of Maryland led by Kirchenbauer published "A Watermark for Large Language Models," which won the Best Paper Award at ICML that year (out of roughly 1,800 accepted papers). The method pre-splits vocabulary into two groups and detects usage based on which group was favored more heavily — and this is also where the limitation that it "can't be applied to sentences with no choice" was first pointed out. This is the same team behind the "long text preserves it" paper discussed in the previous section.
2023–2024. Google DeepMind applied SynthID to images in August 2023, and to Gemini's text output in May 2024. In October, alongside the Nature paper, it open-sourced the code on Hugging Face for anyone to use. That was the moment the technology went from being one company's proprietary asset to a shared industry-wide component.
And OpenAI's choice not to turn it on. In August 2024, The Wall Street Journal reported that OpenAI had built text watermarking with 99.9% confidence but held off launching it for nearly a year. A major factor was an internal survey from April 2023 in which about 30% of users said they would use ChatGPT less if OpenAI alone applied watermarks while competitors didn't. OpenAI also publicly cited another, weightier reason on the same day — that watermarking could stigmatize people who aren't native English speakers.
August 2, 2026. The transparency obligations under Article 50 of the EU AI Act took effect. Providers of systems that generate AI content must ensure that outputs are marked in a machine-readable format. Anthropic signed the "Code of Practice on Transparency for AI-Generated Content" in July 2026, and about 190 organizations have signed on. Violations carry fines of up to €15 million or 3% of global revenue.
The most contentious point follows from this. Anthropic turned on the watermark not just in Europe, but worldwide. Its stated reason is blunt: "there is not yet a robust way to scope this by region." It's a textbook example of the so-called "Brussels Effect," in which a regulation created in Europe becomes a de facto global standard.
South Korea is different — watermarking isn't mandatory
The AI Basic Act took effect on January 22, 2026. Article 31, Paragraph 2 requires labeling generative AI outputs as such, and Article 23, Paragraph 2 of the enforcement decree specifies the method as either a human-perceptible label or a machine-readable marker. Meeting either one suffices. Even when choosing the machine-readable option, providers must still notify humans at least once.
In other words, South Korea does not mandate watermarking. This is where it diverges from Europe. The administrative fine (up to 30 million won) applies to a violation of the prior disclosure obligation under Paragraph 1 of Article 31, not to a failure to mark the output itself — a point that is often misread.
| European Union | South Korea | China | California (US) | |
|---|---|---|---|---|
| Effective date | 2026-08-02 | 2026-01-22 | 2025-09-01 | 2026-08-02 |
| Machine-readable marking | Mandatory | Optional | Mandatory | Mandatory |
| Human-visible label | Distributor's responsibility | Mandatory | Mandatory | User's choice |
| Obligation on users | None | None | Yes | None |
China stands out as uniquely comprehensive, imposing obligations not just on developers but also on users and distribution platforms. California will also require detection capability from platforms starting January 2027.
The relevant statutes are Article 50 of the EU AI Act, Article 31 of South Korea's AI Basic Act, China's "Measures for Labeling Management," and California's SB 942. Marking methods differ slightly — China designates an "implicit label" embedded in file metadata, while California designates a "latent label"; California's human-visible label requirement also comes with the qualifier "to the extent technically feasible."
The same day, Google went the other way
On the very day Anthropic's explanation came out, Google made the opposite announcement.
Josh Woodward, VP of Google Labs, said on August 14 (10:39 p.m. Korean time) that a toggle to turn visible watermarks on and off had been opened up across Gemini and Flow. It applies to Nano Banana (images), Omni (video), and Lyria (music) alike. The criterion wasn't subscription tier but whether the country legally requires watermarks to remain in place.
This is where South Korea comes in. Google Flow's help documentation states plainly that visible watermarks are automatically applied for residents of India, South Korea, and Vietnam. In practice, however, country, subscription plan, and account type all overlap — Digital Trends reported that even in these three countries, paid AI Ultra subscribers can turn it off, while work and school accounts have no such setting at all.
More significant is the caveat Woodward attached: only the visible watermark can be turned off — the invisible SynthID and the electronic tag attached to files remain in place regardless. The volume of content Google has tagged grew from 10 billion pieces in May 2025 to over 100 billion images and videos alone by May 2026.
The two companies' moves look like opposites, but they point in the same direction: stripping away human-visible labels while shifting toward machine-readable markers. OpenAI, too, introduced SynthID for images in May 2026 and extended it to audio in late July — though it still does not apply it to text. Text watermarking itself was pioneered by Google, so what's new about Anthropic isn't the technology but the justification. This is the first case of a company turning it on worldwide explicitly citing regulatory compliance.
The real reason people are angry
The way it was announced was itself part of the problem. Anthropic first disclosed this not in a news release but in a quiet help document. It spread after gaining 449 points on the developer community Hacker News on August 10, and was picked up the next day by TechCrunch, Forbes, and Euronews. But that document never explained how the marking actually works. Tech columnist John Gruber (Daring Fireball) hit this gap on August 11 in a piece titled "Claude Explains How It Marks AI Content Without Explaining How It Marks It." His critique: if it worked by inserting invisible characters, text length would inflate on platforms with character limits; and if it worked by biasing word choice, that would contradict Anthropic's own claim that it "doesn't change meaning, quality, or readability."
It was the community that filled in the blanks first. Developer John J. Wang scanned 7.2 million characters of Claude-written text on August 12 and confirmed no invisible special characters were mixed in, concluding that "it must be a method that alters word choice using a secret key." He explicitly noted, however, that he could not prove it was SynthID specifically — a gap Anthropic filled on August 14. The nature of this episode is that Anthropic's statement wasn't a fresh announcement but a after-the-fact explanation prompted by four days of backlash.
The backlash split into two camps. The first was framed around catching cheaters. TechCrunch summed up the reaction on August 12 as coming from "users angered by the risk of being caught using AI at work or in class." One side argued, "I directed and edited it, so Claude is just a tool"; the other countered, "I only gave instructions — Claude actually made it." The pro-watermark argument was simple: "The only reason you wouldn't want this is to deceive someone."
The second camp is far weightier: people who used Claude as a proofreading tool. Radio host Erick Erickson wrote that "I switched from Grammarly to Claude because it proofread better — and now my own writing gets tagged as Claude's work." A widely shared response on Hacker News was even more direct: "I'm autistic, have ADHD, and dyslexia. It's hard enough just getting my meaning across clearly — now I have to worry about watermarks too. The stigma I already face from society is enough." This echoes precisely the "non-native speaker stigma" concern OpenAI cited two years ago when it shelved its own launch.
Anthropic's response is that "light editing leaves almost no trace." The principle is sound. But no threshold has been disclosed for what counts as "light" editing, nor any procedure for disputing a mistaken flag. What happens when a probabilistic signal like this flows unchecked into schools and workplaces is exactly what years of false positives from AI detectors flagging human writing as AI-generated have already shown.
Editor's take
The essence of this decision isn't detection — it's the redistribution of responsibility. What Anthropic gains is regulatory compliance; what it loses is the trust of users who relied on Claude as a proofreading tool. The gap in between — thresholds, appeals, remedies for false positives — remains unfilled by anyone. What actually matters going forward is when and under what criteria the detection API is released.
Three practical takeaways for now. ① If you hand off an entire draft, the marker sticks; if you use it to polish your own writing, it barely does. ② Code is largely outside its reach — comments being the exception. ③ Translation is entirely subject to marking, so if you're publishing translated output under your own name, now is worth checking.
What's changed isn't the technology. Between Aaronson's 2022 proposal, OpenAI's 2023 hesitation, and Anthropic's 2026 rollout, what shifted was regulation and competitive dynamics. OpenAI concluded that going it alone would cost it users; now, roughly 190 organizations are bound together in the same direction under the EU's code of practice. Watermarking was always a coordination problem, not an algorithmic one — exactly as the Nature paper noted in its section on limitations, stating that "cooperation among watermarking parties is a prerequisite."
Still, optimism is premature. Once a technology that academics have already demonstrated can be forged and stripped begins to be used as grounds for judging an individual's honesty, the harm falls first on those following the rules. The question the watermark needs to answer isn't "did AI write this," but "is it acceptable to punish a person based on this marker." That answer appears nowhere in Anthropic's explainer.





Comments