On 5 October 2026 OpenAI said it will add an invisible statistical watermark, called textGrain, to ChatGPT and Codex text for users in the European Union. In the API the same feature is available worldwide but switched off by default, and the detector is not public. The announcement is easy to read as a simple ‘AI text will now be labelled’ story. OpenAI’s own published test results say something more careful: the signal is strong in long, unedited prose and fades quickly when text is short, technical, translated or lightly rewritten. A missing watermark, OpenAI says, does not prove a person wrote the text.
What OpenAI is rolling out, and to whom
According to TechCrunch and OpenAI’s own post, watermarking for ChatGPT and Codex text will reach eligible EU users on all plans over the coming weeks. It is not a global default at launch. For developers, API watermarking opened on 5 October for select models, worldwide, and OpenAI states it will remain off by default. Availability through cloud partners is expected in the coming weeks. TechCrunch ties the timing to the EU AI Act’s transparency rules, which took effect on 2 August 2026, and notes that Anthropic said on 11 August it would watermark future Claude models, drawing user backlash. We covered the legal regime itself in an earlier piece and do not repeat it here; this article is about the vendor implementation and its measured reliability.
The detector is the part ordinary businesses should read twice. Applications opened on 5 October, but access is granted case by case and, in OpenAI’s words, initially only to approved researchers and expert organisations. OpenAI gives the risk of missed watermarks and false positives as the reason it is not public at launch. The tool reports whether an OpenAI watermark is detected; it does not identify users or reveal prompts. So a publisher, school or HR team cannot simply paste a suspect document into a public checker.
How a statistical text watermark works
A language model produces text one token at a time. At each step it assigns a probability to every possible next token and samples from them, so the likeliest word usually wins but not always. A statistical watermark exploits that slack. TechCrunch reports that textGrain uses a secret key to sort the model’s next-word predictions, subtly shaping word choice, and that a detector holding the key looks for the resulting pattern. OpenAI has not described the exact algorithm in what we read, so the following is illustrative of the published family of methods and not a description of textGrain’s internals. In the 2023 red/green-list scheme by John Kirchenbauer and colleagues, which IEEE Spectrum describes, a key and the preceding word split the vocabulary into a green list and a red list. Green-list words get a small probability boost. No single word proves anything, because green words appear in unwatermarked text by chance too. But over hundreds of tokens, a text that picks green words far more often than chance is statistically unlikely to be innocent.
That is why the watermark is probabilistic rather than a stamp. The signal lives in the statistics of many choices. It survives copy and paste, which OpenAI confirms, because the words themselves carry it. It weakens when words are changed, when the passage is too short to accumulate evidence, and when the text leaves little room for choice, as with a maths answer where the next token is largely forced.
Where it fails, by OpenAI’s own numbers
OpenAI reports its detection rates at a target false-positive rate of 1%. On 400-token psychology passages it reports about 95% detection, and about 80% at 200 tokens. For editing, it reports that replacing 10% of the words in 400-token passages with synonyms cuts detection from about 92% to 66%, and replacing 25% cuts it to 17%. These are OpenAI’s tests on its own system, not an independent evaluation, and the post does not say why the 400-token baseline is 92% in the editing test and about 95% in the clean test; they may use different passage sets, so we do not merge them. Mathematics, OpenAI says, is substantially harder to detect, and TechCrunch adds that short passages and translated text are harder too.
Here is the arithmetic the announcement does not do. A drop from 92% to 66% is 26 percentage points, a relative loss of about 28%: after a light synonym pass, roughly one watermarked passage in three goes undetected. At 25% replacement the fall is 75 points, a relative loss of about 82%, so only about one passage in six is still caught. Halving length from 400 to 200 tokens costs about 15 points on the clean tests. And a 1% target false-positive rate, applied to 10,000 human-written passages, would flag about 100 of them wrongly. The 1% is a target OpenAI states for its experiments, not a guarantee in deployment, but it shows why the company is cautious about opening the detector to everyone: at scale, even a small error rate produces many wrongly accused authors.
OpenAI is explicit about the limits. A detected watermark, it says, can indicate that an OpenAI system generated or processed part of a passage, but not how much human judgement or editing went into it. A watermark does not establish ownership or responsibility, identify the user, or verify accuracy. No watermark does not prove human authorship, because the text may be too short, edited, translated, from an unsupported model, from before watermarking began, or from another company’s tools. The practical consequence is asymmetric. A positive result is informative; a negative result is close to uninformative.
The quality dispute and the counter-argument
Every watermark trades some output freedom for detectability, so the quality question is the first one sceptics ask. OpenAI reports no meaningful difference with and without watermarking on its benchmarks for its Astra model, citing Artificial Analysis Intelligence Index scores of 49.57 unwatermarked and 49.76 watermarked. That is a vendor figure on one composite index, and a gap of 0.19 points in the other direction says little about stylistic effects on specific kinds of writing. Google tested its SynthID-Text watermark differently: IEEE Spectrum reports Gemini queries were randomly routed to watermarked and unwatermarked variants across 20 million responses, with no significant difference in user feedback, though detection reached up to 95% in best cases and fell below 50% for short replies.
The named counter-argument comes from Vinu Sankar Sadasivan, an AI research scientist at Meta, quoted by IEEE Spectrum. He disputes Anthropic’s claim that its watermark leaves quality unchanged, noting that watermark strength can be tuned to protect quality but that lowering it also weakens detection. He also points out that when few word choices are available the method struggles: for a 20-word tweet, he says, 50 or 60 percent of the words would need to come from the green list for reliable detection. The tension he describes is built into the design, and textGrain’s published numbers sit somewhere on that trade-off curve that OpenAI has chosen.
There is also a history worth knowing. TechCrunch reports that the Wall Street Journal wrote in 2024 that OpenAI had built a text watermark earlier but held it back, partly over concerns users would switch to rivals. Making it mandatory in the EU, where the law applies regardless of competitive pressure, and optional elsewhere is consistent with that old worry. OpenAI’s decision to leave the API default off is the same instinct applied to developers: it avoids imposing detectability on customers who have not asked for it.
What developers and compliance teams should do
- Treat a watermark detection as supporting evidence, never as proof of authorship, and never treat a missing watermark as evidence of human writing. OpenAI says the same.
- Decide the API setting deliberately. Enabling it may help you demonstrate good-faith labelling if you distribute generated text in the EU; leaving it off keeps outputs unmarked. The default is off, so doing nothing is a choice.
- Do not plan workflows around running the detector yourself. Access is limited to approved researchers and expert organisations, so most firms will rely on a third party or on provenance records they keep.
- Expect weaker results on short, technical and translated text, which are common in business use, and log which model produced which output so you do not depend on detection alone.
What the evidence does not establish
We have not seen independent tests of textGrain, so every robustness figure above is OpenAI’s. OpenAI says a technical report, co-written with researchers from the University of Pennsylvania and Yale, will be updated in coming weeks and that it plans to release the technology as open source; both could change the picture once outsiders can attack it. We also cannot say how well the system holds up against deliberate removal tools beyond synonym swaps, how false-positive rates behave on real-world text rather than test sets, or how regulators will judge whether a restricted detector satisfies the transparency requirement. This analysis is current as of 6 October 2026.

