Artificial intelligence firm Anthropic has pledged to include hidden watermarks in text created by its popular chatbot Claude to make computer-generated content easier to detect, with major competitor OpenAI also agreeing to new transparency rules which could see it follow suit.
The moves come amid new obligations under the European Union's AI Act, but are expected to be applied worldwide by both companies.
Anthropic will initially apply a text watermarking technique to future Claude models but also plans to add it to earlier ones, the company confirmed last week after it committed to the European Commission’s Code of Practice on Transparency of AI-generated Content.
The likes of Anthropic, OpenAI, Google, Meta, and Microsoft have so far agreed to the EU's new rules, which took effect on 2 August and legally require AI providers to make computer-generated text detectable, "with the exception of very short text".
Claude's AI text watermark will be a version of Google DeepMind's SynthID system, which Google already uses to embed watermarks into most of its AI-generated text, images, audio, and video.
OpenAI quietly told ChatGPT users earlier this month that it plans to "expand provenance signals to all modalities including text", but has not yet detailed how or when it will do so.
The company previously developed an AI text classifier which it shelved in 2023 "due to its low rate of accuracy", but has also reportedly developed a text watermarking tool that it has not released due to circumvention risks.
While AI text detection software already exists, it has faced accuracy issues because it works by looking for words or text structures favoured by AI, instead of searching for hidden patterns created by text watermarks.
Dr T J Thomson, an associate professor of visual communication and digital media at RMIT University, said text watermarking techniques "are much more precise and reliable for a particular AI model compared to general-purpose detectors".
"The main benefit of AI text watermarking is a higher level of reliability for a particular model, with the caveat that it only works for that particular model and not for any text generated through other AI applications," he told Information Age.
How do AI text watermarks work?
While watermarks and metadata used to label AI-generated images and videos can usually be read by humans, text watermarking alters the way an AI model chooses its words in a way which is not perceptible to humans but creates a machine-readable pattern.
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself," Anthropic told users last week.
"You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response."
Claude's text watermarking will not be used in software code it processes "where an exact output is required", but it can be used in more flexible content "such as comments within code", Anthropic said.
AI companies are developing computer programs which can identify AI text watermarks from their own models – which, if found in a piece of text, shows that content was in some way processed by that specific model.

Anthropic says its text watermarks won't change 'the meaning, quality, or readability' of Claude's output. Image: Shutterstock
Anthropic said it is working to allow users and third parties to detect Claude's text watermarks.
If a Claude watermark is found during detection checks "it indicates that the content may have been processed by Claude", the company said – but it "is not fully conclusive" about where the text came from.
This is because watermarked Claude output could still be altered after it is created, and because the model's output "can carry a Claude mark even if the underlying ideas, text, or data originated from another source", the company said.
The lack of a text watermark on a piece of text also does not guarantee that it was not wholly or partly generated by AI.
This is largely because a watermark can be weakened if the text is heavily edited, cut, or mixed into other writing, or run through other platforms and software.
Will AI users try to 'cover their tracks'?
Following Anthropic's announcement, technology analysts at Gartner wrote that the rise of new AI transparency requirements "will make it harder" for people to "pass off AI-generated content as their own" and may increase workers' use of so-called 'shadow AI' systems not approved by their employer.
"Any means to further investigate content provenance, or further restriction of the usage of certain tools, will have unintended side effects as staff look to circumvent controls by using nonapproved providers, models, and tools," analysts Robert Stoneman and Bart Willemsen said in a research note.
Dr Thomson said he does not believe the increasing use of AI text watermarks will change ordinary people's use of large language models (LLMs) for basic tasks "at this early stage".
However, those with greater technical understanding may try to "cover their tracks and game the system".
"I would suspect these users would either try to edit the text or use AI models which don’t have watermarking features," he said.
'Cancelled the subscription'
There have been some reports of people cancelling their Claude subscriptions due to Anthropic's plan to introduce AI text watermarking.
This is largely over concerns about their AI work being detected, the influence of the EU, and whether watermarking may change Claude's output when handling sensitive text such as legal contracts and medical records – despite the company's reassurances.
One software engineer asked Anthropic on social media platform X, "Why should I, being a non EU citizen watermark my work generated by a paid subscription of Claude? Cancelled the subscription."
While Dr Thomson suggested there should be "more safeguards and greater transparency" for AI use in higher-risk situations such as legal research and medical advice, he said "a major issue at present is the lack of interoperability among watermark detectors".
"Each AI service has its own proprietary system and can only detect patterns created by its own software," he said.
"Until we have tools that can detect AI-generated patterns across platforms, verifying AI or human input will remain a piecemeal and fraught activity."