
Image: METAL
Summary
- On October 5, OpenAI announced that it will roll out an invisible watermark called textGrain to ChatGPT and Codex output in the European Union over the coming weeks, in response to the EU AI Act.
- API customers worldwide can switch the watermark on for select models, with the default left off, while the detector is open only to approved researchers and expert organizations.
- At a 1% false positive target, detection reached about 95% for 400-token passages but fell to 17% when 25% of the words were replaced with synonyms, and the company said benchmark quality on Astra showed no meaningful difference.
OpenAI announced on October 5 (local time) its plan to introduce text watermarking in response to the text provenance rules of the European Union's AI Act. The watermarking technology, built in house, is called textGrain. Over the coming weeks, eligible text written by ChatGPT and Codex in the EU will carry an invisible watermark, and API customers around the world can switch the watermark on for select models starting from the day of the announcement. The API default is off. The detector that looks for the watermark will not be released to the public; instead, OpenAI is taking applications from approved researchers and expert organizations.
The EU AI Act requires generative AI providers to make AI-generated text identifiable in a machine-readable way. According to reports, Article 50, which carries this obligation, took effect on August 2, 2026, with existing systems given a transition period until December 2. In its announcement, OpenAI said "text watermarking and detection remain early technologies with significant limitations," explaining that its phased approach reflects both the law's requirements and the state of the technology. OpenAI had earlier written on its support site that its "goal is to expand provenance signals to all modalities including text."
textGrain mixes a statistical signal, generated with a secret key, into the probabilities the model uses to choose its next word. Readers see no difference, but a detector holding the same key looks for the bias left in word choices to determine whether an OpenAI watermark is present. According to the 20-page technical report METAL reviewed, the method divides the vocabulary into keyed blocks and solves an optimal transport problem to tie the choice of block to the key. Within each block, the original relative probabilities of the words are kept. The design ensures that averaging over randomly drawn keys recovers the model's original distribution, so the watermark does not push the model's answers in one direction.
From an engineer's point of view, the central mechanism is the entropy budget. The earlier Gumbel-max approach always picks the same word for the same context and the same key, so asking the same question repeatedly can produce identical answers. The report addresses this by measuring the sampling randomness removed by the watermark with a single KL divergence value and capping it at a fixed fraction of the entropy. The detector needs only the generated text and the secret key, with no need to know the model or the budget used during generation. The report said the method can also be combined with speculative sampling, in which a draft model and a target model share the key. The report has nine authors, five OpenAI researchers and four from the University of Pennsylvania and Yale University, and the corresponding author is OpenAI's Weijie Su.
The performance figures OpenAI released varied widely by condition. In tests with a target false positive rate of 1%, the detector found the watermark in about 80% of 200-token psychology answers and about 95% of 400-token ones. Detection dropped sharply for mathematics answers, where word choice is less flexible; in the chart the company published, it stayed below 40% at 200 tokens and around 60% at 400 tokens. Edited text was harder still. In 400-token passages, replacing 10% of the words with synonyms lowered detection from about 92% to 66%, and replacing 25% brought it down to 17%. OpenAI said textGrain matched or exceeded the other approaches it tested, including Google's SynthID for text. The announcement also noted that strong results under ideal conditions do not guarantee reliable detection in everyday use.
The company said quality loss was negligible. OpenAI ran its latest frontier model, Astra, at its maximum setting and compared watermarked and unwatermarked output across eight benchmarks. The Artificial Analysis Intelligence Index went from 49.57 points to 49.76 points, GPQA Diamond from 94.44% to 93.94%, DeepSWE v1.1 from 72.80% to 71.68%, and Terminal-Bench 4.0 from 53.90% to 56.06%. With results moving in both directions, the company concluded that it saw no meaningful performance difference.


The decision not to release the detector publicly follows from these figures. OpenAI cited the risk of false positives, reporting a watermark where there is none, and false negatives, missing one that is present, as the reason it is not making the tool public at launch. Access will be granted case by case in accordance with the Code of Practice, and the tool will only report whether an OpenAI watermark is detected, without revealing who the user is or their prompts and conversations. The company also set out five things a watermark result cannot tell you. A watermark does not say how much a person contributed, who owns or is responsible for the text, who wrote it, or whether the content is accurate. Nor is the absence of a detected watermark proof of human authorship, because the text may be too short, edited or translated, or may have been produced with another company's tools.

Among competitors, Anthropic moved first. METAL previously reported Anthropic's decision to put an invisible watermark in text written by Claude. At the time, Anthropic said it would "mark AI-generated content from day one" and applied the measure across all supported regions. OpenAI, by contrast, said it will turn the ChatGPT watermark on only in the EU and will not make it a global default, so that it can start regionally and learn from real-world use and feedback. According to reports, Microsoft, Meta and Mistral also signed the same code, while xAI did not, leaving Grok without an equivalent commitment.
Measures for images and audio remain in place. OpenAI attaches Content Credentials to supported images, is C2PA conformant, and embeds SynthID watermarks in images and audio. The openai.com/verify web tool and the Content Provenance API will also stay publicly available. The company said that within the coming weeks it plans to offer watermarking for OpenAI model output accessed through cloud partners, expand the technical report, and release textGrain as open source.
The number this announcement leaves behind is 25%. The fact that changing one word in four drops detection to 17% means a text watermark is less a detector that catches AI writing and more a stamp that confirms the origin of untouched output. OpenAI chose to publish that boundary alongside its figures and to open detector access narrowly. The company said it will work to narrow the gap between the machine-readable marking that regulation demands and the reliability current technology can offer as evidence and standards accumulate.





Comments