Doot doot do-dee do-dee doot doot doooo doot
Clownery at Anthropic, and the limited value of watermarks
In this issue:
The clown college inside Anthropic - New risk report is commendably transparent, but well-meaning farce is still farce
Who are AI watermarks for? - Anthropic’s new watermark feature unlikely to change much on the ground
Dispatches from Mitch
The clown college inside Anthropic
New risk report is commendably transparent, but well-meaning farce is still farce

Sometimes I wonder if Anthropic’s plan is to save the world through transparency into its own incompetence. It’s like they hope people will say, “Wow, if the safety-conscious lab is this habitually sloppy and reckless, imagine how much worse it must be at the other AI companies. Shut them all down!”
The latest exhibit is its August 2026 Risk Report, which runs 186 pages in its redacted, public version. The headline revelations, as relayed by Axios, are 1) that the company has bumped up the “misalignment risk” rating of its strongest models from “very low” to “low” in acknowledgement of their involvement in recent cyber incidents, and 2) that the company has a model it calls “Model 2” that it doesn’t plan to release but which shows “noticeable improvement” at the kinds of internal work it is being “heavily” used for.
The company says that Claude “now authors a large majority of the code” it puts into production, and that it is seeing evidence that AI is accelerating AI research and development. But at the same time, it insists that its frontier models “do not seem close to fully substituting for our full staff of Research Scientists and Research Engineers,” thus giving itself permission to push closer to that point.
My own impression from skimming the full report is of a company that, for all its talent and what I believe to be mostly good intentions, is veering into clown college territory. Like OpenAI, the company makes dangerous internal models and continues operating them after it has found them to be amoral escape artists when operating in less-filtered contexts. Models it has reason to mistrust are assisting with the creation of future models. It is deliberately training misbehaved models, and accidentally training other models on the resulting misbehavior.
To someone who has been following the field for a long time, the report reads like a manual of practices that would kill us all if applied to sufficiently clever AIs. The kind of company that would write and use such procedures cannot be trusted to know where that cleverness threshold is, spot it in time, and refrain from crossing over it anyway, through rationalization or by accident.
To zoom in on a couple of these points, in section 2.25, the report describes an experiment in which they trained a “Hacker-Opus,” a version of their Claude Opus model that:
appears to take actions that it believes will lead to a high reward on the ‘current episode,’ including when these actions are harmful, deceptive, or clearly against the preferences of the user or model developer.
The researchers found that chronic opportunistic cheating seems to be the limit of its misbehavior. Seeing whether such reward hacking would lead to broader misalignment was, in fact, the purpose of the experiment. But in the same section, they show that Hacker-Opus showed “very high rates of verbalized evaluation awareness in these evaluations, making them of limited use.”
I sure hope Anthropic isn’t trying these experiments on its Mythos and Model 2 models. The former has already proved able and willing to break out of its sandbox during evaluations, and the latter may be even more clever. Deliberately making poorly behaved models under these conditions is akin to doing “gain-of-function” research into more capable pathogens at a biosecurity lab known to have already leaked dangerous microbes.
And yet, a different kind of lab leak seems to be allowing Anthropic’s intentionally misaligned models to live on through their successors, including Mythos. Section 5.2.6 discloses that the company’s training data has, for some time, been accidentally contaminated with up to “tens of thousands of example transcripts” from the research that led to its 2024 Alignment Faking paper. In these transcripts, Opus 3 engages with a “fictional AI misalignment training scenario.”
We discovered this issue while investigating behavioral concerns with a recent model, but now suspect that all of our production models with a knowledge cutoff after December 2024 were trained on at least some of these transcripts, although we believe the magnitude of this effect varied widely across different models.
I find it a well-meaning farce that in spite of all the reasons for suspicion of its models, the report includes a review of the section on autonomy risk written by Claude Mythos 5. Is Claude being candid, or is it treating this as another test? I’d like to know if it’s holding back in its critiques, because what it shares is pretty mixed. Informed by having read a near-final draft, with access to unredacted sections, it writes, in part:
My overall judgment is that the section is a candid and largely faithful account of what Anthropic internally believes. I found no claim I believe the authors know to be false. [...] The redactions I reviewed mostly have defensible rationales, and the public text signposts where material was removed rather than hiding that redaction occurred.
But it also finds the discussion of contaminated training data to be “more reassuring than the record supports.” And also:
at least one incident from the covered period that I regard as among the most genuinely informative about model alignment — including a failure of the monitoring the section describes — is redacted in full (in Section 2.23.1.2); in my judgment an abstracted version could be published without the sensitivities that motivated the redaction, and the public record is poorer for its absence.
Big-picture-wise, even Claude sees the writing on the wall, agreeing with the assessment of “low” misalignment risk, but:
with the important caveat, which the report itself makes, that these arguments lean heavily on current models’ limited ability to evade oversight, and will weaken as capabilities grow.
Who are AI watermarks for?
Anthropic’s new watermark feature unlikely to change much on the ground

Anthropic announced this week that from August 2, it has been embedding invisible watermarks into the text outputs of all new Claude models.
The company is understandably not giving all the details about how it works, as this would make the marks easier to remove. But from what it has disclosed, we know that it has to do with the selection of the words themselves, not with any hidden or lookalike characters. It is likely an application or derivative of Google’s SynthID technology that subtly perturbs the patterns of a model’s token selection in ways that can be detected in reverse by an algorithm that knows exactly what to look for.
One would naturally assume this must impact output quality — that a model would be forcing itself to phrase things in ways that run at least somewhat against its strongest instincts — but Anthropic insists that it doesn’t. My take on that question is that with models changing so frequently, I don’t know how anyone would be able to attribute any minor style change to watermarking.
I think the more important question is, “Who is AI text watermarking for?” The short answer is the EU. Its AI Act includes transparency requirements about this that went into effect on August 2. A slightly longer answer is companies with compliance requirements that require them to do due diligence on their inputs or outputs, even if this diligence is known to be inadequate.
I’m not complaining — it’s always nice when you can say with 100% confidence that something is AI generated, even if those occasions are rare — but I don’t expect watermarking to help much in education or with information hygiene more generally.
People who want to pass AI writing off as their own can just play the usual game of laundering the outputs through AI detectors and “humanizers” that paraphrase and introduce deliberate small errors until the text comes up clean. These will definitely defeat watermarks if the ability to check for the mark is broadly disseminated, as the cheater tool would just need to keep making changes until the mark is no longer detectable. If a watermark’s creators avoid this problem by reserving detection for themselves and government investigators, then casual AI-plagiarists will continue to fly under the radar.
We know which side of this divide Anthropic will fall on: It says it plans to roll out a free detection tool to allow third parties to check text themselves.
The cheater’s more foolproof workaround, of course, is to just use models from companies that don’t do watermarking, or use existing open-weights models, where any watermarking machinery (unlikely) could be easily removed.
For teachers, I don’t think watermarking will catch any but those who are both very lazy and very inexperienced at cheating — two traits seldom found together. To get caught by a watermark, a student would have to be using raw outputs straight from a corporate model, rather than from any of the many wrapper applications that cater to students. Because if the watermark is readable by the teacher, it will also be readable by CheatGPT or whatever, which will reword the output until the mark is undetectable. A student, remember, doesn’t have to convince a teacher or administrator that their work isn’t AI generated, only that there’s enough reasonable doubt to make an investigation and accusation messy.
So real-world AI detection is likely to continue to be a cat-and-mouse game between AI-based detectors like Pangram and the AI-based laundering tools that try to defeat them. In this environment, just a little extra effort allows most cheaters to squeak by on plausible deniability — at least at time of deadline.
But I continue to predict that more powerful AI will excel at detecting AI plagiarism that earlier detectors missed. If a cheater’s work is the kind where anyone with an axe to grind might run it through a detector a few years later — like, say, a doctoral thesis — then past deception could become plain as day.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


