metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
OpenAI
ChatGPT
Watermarking
EU AI Act
AI Transparency

OpenAI Text Watermark: How textGrain Marks ChatGPT Output

OpenAI is adding an invisible text watermark called textGrain to ChatGPT and Codex in the EU. How it works, where it breaks, and how it compares with SynthID.

Metir AI TeamOctober 6, 20268 min read
OpenAI Text Watermark: How textGrain Marks ChatGPT Output

On October 5, 2026, OpenAI announced an invisible text watermark, called textGrain, for ChatGPT and Codex output in the European Union. The OpenAI text watermark is being added to eligible text over the coming weeks, according to PYMNTS. API customers anywhere can opt in to watermarked text for select models starting immediately, but the setting is off by default, and access to the detector is limited to approved researchers and expert organizations. The change is a response to Article 50 of the EU AI Act, and it follows Anthropic's similar move in August.

OpenAI logoOpenAI
Anthropic logoAnthropic
Google logoGoogle
Three providers with text watermarking tied to EU transparency rules or, in Google's case, an earlier research effort.
Oct 5, 2026textGrain announcedChatGPT and Codex, EU first
~80%Detection at 200 tokensreported
~95%Detection at 400 tokens400-token psychology passage
17%Detection after 25% word replacement400-token passages

This piece explains how statistical text watermarking works, what OpenAI has said about textGrain's design and limits, how it compares with the Claude watermark and Google's SynthID Text, and what Article 50 actually requires and when.

What OpenAI announced

Based on the coverage reviewed for this post, the key terms are as follows. OpenAI's own blog post could not be retrieved directly, so the details below come from its technical report plus press coverage of the announcement.

  • Where it applies by default: ChatGPT and Codex text in the EU, rolling out over the coming weeks. ActuIA reports it covers all ChatGPT and Codex plans there.
  • API: opt-in for customers worldwide on select models, off by default.
  • Detector: not public at launch. OpenAI cites the risk of missed watermarks and false positives, and says it is working with researchers, with partners reported to include Cornell, ETH Zurich and the Slovak institute KInIT. It has also said it plans to open-source textGrain.
  • Stated limits: per PYMNTS, the watermark can indicate that an OpenAI system generated or processed part of a text, but it cannot show how much a human edited it, establish ownership, link a user to content or verify accuracy. In OpenAI's wording as quoted there, "the absence of a watermark does not prove the passage was authored by a human."

The technical report, titled textGrain: Entropy-Calibrated Watermarking for Language Model Text, lists authors from OpenAI, the University of Pennsylvania and Yale University.

How statistical text watermarking works

A language model writes one token (a word or part of a word) at a time by sampling from a probability distribution over the next token. A statistical watermark leaves the words readable but changes how that sampling is done, so the choices carry a signal only a key holder can test for.

The general recipe, which the textGrain report itself describes, has three parts:

  1. A secret key plus recent context. At each position, a pseudorandom value is derived from a secret key and a window of preceding tokens.
  2. Keyed token-choice biasing. Instead of using fresh random numbers, the sampler uses that keyed value to decide which plausible token to emit, nudging choices toward tokens the key favors without changing which tokens are plausible.
  3. Detection by correlation. A detector that has the key recomputes those values from the text and tests whether the observed tokens depend on them more than chance would predict.

There is a built-in tradeoff. If tokens are independent of the keyed randomness, there is nothing to detect. If the randomness fully determines every token, the model loses its sampling variety. The report frames this as an entropy budget: how much sampling randomness may be given up in exchange for signal.

“

If tokens are independent of the keyed randomness, there is no watermark signal to detect. If that randomness determines every token, no further sampling randomness remains.

Paraphrase of the framing in OpenAI's textGrain technical report

Text is also a harder medium than images. The signal lives in word choice, so it needs enough tokens to accumulate evidence, and passages where the model has little freedom (a factual answer, a fixed code snippet) carry little of it.

What is distinctive about textGrain

According to the technical report, textGrain couples token generation to keyed randomness through an optimal transport problem, using costs based on Gumbel random variables and KL regularization. The KL divergence equals the reduction in conditional entropy, which gives the watermark strength setting a direct information-theoretic meaning. To make this fast, it applies the transport to blocks of vocabulary tokens, preserving the relative probabilities of tokens within each block. The detector needs only the text and the secret key, not the budget used at generation time.

The report also points at a known weakness of earlier schemes. Gumbel-max watermarking selects the token with the largest sum of its log probability and a keyed Gumbel value, so at a fixed context and key it always picks the same token, and repeated generation from the same prompt can produce identical answers. Keeping some sampling randomness is the stated motivation for the entropy budget. ActuIA adds that the signal is carried by the words themselves, with no hidden characters or invisible spaces to strip out.

Robustness: what the reported numbers show

ActuIA's summary of OpenAI's material gives the clearest picture of the tradeoff between length, editing and detection.

textGrain detection rate by passage length and editing

Share of passages in which the watermark was detected. Green bars are unedited text; grey bars are text with words replaced. Figures as reported from OpenAI's technical material.

Source: ActuIA's reporting on OpenAI's October 5, 2026 textGrain technical material. Different passages and test setups, so compare the shape, not exact deltas. Hover a bar for details.

  • A 200-token passage was detected about 80% of the time, and a 400-token psychology passage about 95%.
  • In 400-token passages where 10% of words were swapped for synonyms, detection fell from about 92% to 66%. At 25% replacement it fell to 17%.
  • Mathematical content was detected around 60% at 400 tokens, which fits the low-freedom problem above.
  • Across 24 EU languages, detection ranged from 42.2% in Romanian to 69.0% in Spanish.
  • Light edits and copying are tolerated, but substantial paraphrasing or translation can make the watermark undetectable.

The chart mixes passages and test setups, so the exact deltas should not be over-read. The pattern is the useful part: more text helps, domain matters, and heavy rewriting erodes the signal. In practice that means a negative result is weak evidence, which matches OpenAI's own statement about absence of a watermark.

The Berlaymont building in Brussels, headquarters of the European Commission, with a large European Commission banner on its central tower
The Berlaymont building in Brussels, headquarters of the European Commission, the EU institution that publishes the AI Act guidance and code of practice behind these watermarking changes. Photo by Romaine via Wikimedia Commons, CC0.

How it compares with Claude and SynthID Text

All three approaches are statistical and share the same weaknesses, so the differences are mostly about deployment and access.

OpenAI textGrainAnthropic ClaudeGoogle SynthID Text
MechanismKeyed optimal-transport coupling over vocabulary blocksDescribed by Anthropic as an adaptation of SynthID-TextModulates the likelihood of tokens as they are generated
Where appliedChatGPT and Codex in the EU; API opt-in, off by defaultNew Claude models worldwide, per our earlier post on Claude's watermarkGemini app and web experience
DetectorResearchers and expert organizations only at launchAnthropic has said it plans a detection APINot covered in the sources reviewed here
Known limitsParaphrase, translation, short or low-entropy textRewriting, translation, short passagesThorough rewriting, translation, factual prompts

Google's description of SynthID Text says it is less reliable on factual prompts with little variation and that its effectiveness drops with thorough rewriting or translation, which is the same profile OpenAI reports. The notable policy difference is scope: Anthropic applied marking globally, while OpenAI limited the default to the EU and left everything else opt-in. Neither choice is wrong in itself; they reflect different readings of how much of the market Article 50 reaches.

What Article 50 requires, and when

Article 50(2) of the EU AI Act requires providers of AI systems that generate synthetic audio, image, video or text to mark outputs in a machine-readable format so they are detectable as artificially generated or manipulated. The European Commission's page on the code of practice says the transparency obligations apply from August 2, 2026, that the code was finalized on June 10, 2026, and that roughly 190 companies and organizations had signed by late July. Signing is voluntary, but the legal obligation is not.

Two timing details matter here. First, per ActuIA, systems placed on the market before August 2, 2026 have until December 2, 2026 to comply with Article 50(2), which is the reported grace period under the Digital Omnibus changes. Second, the code of practice exempts very short text, defined as under 200 tokens, from watermarking, a threshold OpenAI says it matched. That aligns with the weak detection signal on short passages.

Penalties can reach 15 million euros or 3% of global annual turnover, whichever is higher, as we noted in the earlier Claude post. The code of practice also leaves detector access mostly to regulators, researchers, media and fact-checkers rather than private uses such as HR screening, per ActuIA's reading.

What to watch next

  • The December 2 date: whether other providers with legacy systems ship marking before the grace period ends.
  • Open-sourcing: if OpenAI releases textGrain, independent teams can test its robustness rather than relying on vendor-reported figures.
  • Detector access: how the researcher program widens, and whether false-positive rates are published.
  • Interoperability: each vendor holds its own key, so a platform wanting to check text across providers needs separate detectors. Teams that use several models, for example through a model-agnostic workspace such as Metir, will see these differences first-hand.
  • Evasion research: paraphrase and translation attacks are the open question, as the numbers above suggest.

FAQ

What is textGrain? OpenAI's text watermark, which adjusts the statistical pattern of word choices during generation so a detector with the secret key can estimate whether a passage came from an OpenAI system.

Does it apply to everyone? By default only to ChatGPT and Codex text in the EU. API customers worldwide can opt in on select models, and it is off by default.

Can I use the detector? Not publicly at launch. Access is limited to approved researchers and expert organizations.

Does a missing watermark mean a human wrote it? No. OpenAI states that the absence of a watermark does not prove human authorship.

Sources:

  • OpenAI rolls out text watermarks to meet European Union AI Act mandates, PYMNTS (Oct 2026)
  • OpenAI will watermark ChatGPT in the EU but leaves the API opt-in, ActuIA (Oct 5, 2026)
  • ChatGPT text is getting an invisible watermark in Europe, Tech Startups (Oct 5, 2026)
  • textGrain: Entropy-Calibrated Watermarking for Language Model Text, OpenAI technical report (Oct 5, 2026)
  • Weijie Su on X announcing textGrain
  • EU Code of Practice on Transparency of AI-Generated Content, European Commission
  • Watermarking AI-generated text and video with SynthID, Google DeepMind
  • EU AI Act Watermarking Grace Period Ends December 2026, Cloud Security Alliance Labs
  • How Claude's text watermarking works, Anthropic

Image credits

  • Hero: flags in front of the Berlaymont building, headquarters of the European Commission in Brussels, by almathias via Wikimedia Commons, CC0.
  • In-body: the Berlaymont building facade, by Romaine via Wikimedia Commons, CC0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational