OpenAI plans to add an invisible watermark to text generated by ChatGPT and Codex for users in the European Union over the coming weeks. The move complies with the EU AI Act’s transparency rules, which took effect on August 2 and require artificial intelligence developers to mark machine-generated output so external systems can detect it. While the regional consumer rollout is slated to reach eligible accounts across all tiers in the near future, developers using OpenAI’s API globally can manually enable the feature starting today, though it remains off by default.
How textGrain Encodes Machine Text
Unlike image watermarking, which embeds visual indicators or metadata that can be stripped during a copy-paste, text watermarking must reside directly inside the words themselves. To accomplish this, OpenAI worked with researchers from the University of Pennsylvania and Yale to design a statistical method called textGrain.
The system operates during the model’s token-prediction process. As the language model calculates potential next words to complete a sentence, a secret pseudo-random key nudges the selection toward specific subsets of words. These adjustments remain subtle enough that sentence flow, coherence, and readability do not change, and OpenAI reported no measurable drop in model output quality. However, when hundreds of these nudges accumulate throughout a passage, a detector equipped with the secret key can evaluate the word pattern and verify that the text originated from OpenAI. The company stated that the watermark does not identify individual users or link back to specific accounts.
Detection Limits and Evasion
OpenAI’s technical report documents practical limits to statistical text watermarking:
- Susceptibility to light editing: In testing, replacing 10% of generated words with synonyms caused detection accuracy to drop from approximately 92% to 66%.
- Short and formulaic formats: Short answers, translations, and deterministic outputs such as mathematical solutions do not contain enough words to sustain a readable mathematical pattern.
- Inconclusive negative results: OpenAI cautioned that the absence of a watermark does not prove human authorship, as text may have been edited, translated, created by a competitor’s system, or kept too brief for analysis.
- Lack of collaborative context: The system can detect whether an OpenAI model generated or processed parts of a passage, but it cannot measure human input, prompt refinement, or editorial revision.
Because of these false-negative risks and detection boundaries, OpenAI is restricting initial detector access to vetted researchers and specialized oversight organizations.
Regulatory Differences and Content Provenance
OpenAI’s regional rollout highlights how compliance requirements are leading to different user experiences across borders. Rival AI developer Anthropic previously announced plans to apply text watermarking to Claude globally. That decision drew pushback from users who argued they supply the core context, instructions, and editorial choices while using the model as an assistive tool. OpenAI had built text watermarking previously but held back on deploying it, partly over concerns that users might move to rival models that do not watermark text.
Major AI companies, including OpenAI, Anthropic, Google, Meta, and Microsoft, have all committed to the EU’s code of practice for identifying generated content. However, the technical limits of textGrain illustrate the challenge of relying purely on word distribution for verification. Because basic synonym swaps degrade detection, industry efforts toward content provenance continue to point toward cryptographic signing, verified compute records, and open provenance standards rather than statistical text modifications alone.
