Who are AI watermarks for?
Anthropic's new watermark feature unlikely to change much on the ground

Anthropic announced this week that from August 2, it has been embedding invisible watermarks into the text outputs of all new Claude models.
The company is understandably not giving all the details about how it works, as this would make the marks easier to remove. But from what it has disclosed, we know that it has to do with the selection of the words themselves, not with any hidden or lookalike characters. It is likely an application or derivative of Google’s SynthID technology that subtly perturbs the patterns of a model’s token selection in ways that can be detected in reverse by an algorithm that knows exactly what to look for.
One would naturally assume this must impact output quality — that a model would be forcing itself to phrase things in ways that run at least somewhat against its strongest instincts — but Anthropic insists that it doesn’t. My take on that question is that with models changing so frequently, I don’t know how anyone would be able to attribute any minor style change to watermarking.
I think the more important question is, “Who is AI text watermarking for?” The short answer is the EU. Its AI Act includes transparency requirements about this that went into effect on August 2. A slightly longer answer is companies with compliance requirements that require them to do due diligence on their inputs or outputs, even if this diligence is known to be inadequate.
I’m not complaining — it’s always nice when you can say with 100% confidence that something is AI generated, even if those occasions are rare — but I don’t expect watermarking to help much in education or with information hygiene more generally.
People who want to pass AI writing off as their own can just play the usual game of laundering the outputs through AI detectors and “humanizers” that paraphrase and introduce deliberate small errors until the text comes up clean. These will definitely defeat watermarks if the ability to check for the mark is broadly disseminated, as the cheater tool would just need to keep making changes until the mark is no longer detectable. If a watermark’s creators avoid this problem by reserving detection for themselves and government investigators, then casual AI-plagiarists will continue to fly under the radar.
We know which side of this divide Anthropic will fall on: It says it plans to roll out a free detection tool to allow third parties to check text themselves.
The cheater’s more foolproof workaround, of course, is to just use models from companies that don’t do watermarking, or use existing open-weights models, where any watermarking machinery (unlikely) could be easily removed.
For teachers, I don’t think watermarking will catch any but those who are both very lazy and very inexperienced at cheating — two traits seldom found together. To get caught by a watermark, a student would have to be using raw outputs straight from a corporate model, rather than from any of the many wrapper applications that cater to students. Because if the watermark is readable by the teacher, it will also be readable by CheatGPT or whatever, which will reword the output until the mark is undetectable. A student, remember, doesn’t have to convince a teacher or administrator that their work isn’t AI generated, only that there’s enough reasonable doubt to make an investigation and accusation messy.
So real-world AI detection is likely to continue to be a cat-and-mouse game between AI-based detectors like Pangram and the AI-based laundering tools that try to defeat them. In this environment, just a little extra effort allows most cheaters to squeak by on plausible deniability — at least at time of deadline.
But I continue to predict that more powerful AI will excel at detecting AI plagiarism that earlier detectors missed. If a cheater’s work is the kind where anyone with an axe to grind might run it through a detector a few years later — like, say, a doctoral thesis — then past deception could become plain as day.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


