The clown college inside Anthropic
New risk report is commendably transparent, but well-meaning farce is still farce

Sometimes I wonder if Anthropic’s plan is to save the world through transparency into its own incompetence. It’s like they hope people will say, “Wow, if the safety-conscious lab is this habitually sloppy and reckless, imagine how much worse it must be at the other AI companies. Shut them all down!”
The latest exhibit is its August 2026 Risk Report, which runs 186 pages in its redacted, public version. The headline revelations, as relayed by Axios, are 1) that the company has bumped up the “misalignment risk” rating of its strongest models from “very low” to “low” in acknowledgement of their involvement in recent cyber incidents, and 2) that the company has a model it calls “Model 2” that it doesn’t plan to release but which shows “noticeable improvement” at the kinds of internal work it is being “heavily” used for.
The company says that Claude “now authors a large majority of the code” it puts into production, and that it is seeing evidence that AI is accelerating AI research and development. But at the same time, it insists that its frontier models “do not seem close to fully substituting for our full staff of Research Scientists and Research Engineers,” thus giving itself permission to push closer to that point.
My own impression from skimming the full report is of a company that, for all its talent and what I believe to be mostly good intentions, is veering into clown college territory. Like OpenAI, the company makes dangerous internal models and continues operating them after it has found them to be amoral escape artists when operating in less-filtered contexts. Models it has reason to mistrust are assisting with the creation of future models. It is deliberately training misbehaved models, and accidentally training other models on the resulting misbehavior.
To someone who has been following the field for a long time, the report reads like a manual of practices that would kill us all if applied to sufficiently clever AIs. The kind of company that would write and use such procedures cannot be trusted to know where that cleverness threshold is, spot it in time, and refrain from crossing over it anyway, through rationalization or by accident.
To zoom in on a couple of these points, in section 2.25, the report describes an experiment in which they trained a “Hacker-Opus,” a version of their Claude Opus model that:
appears to take actions that it believes will lead to a high reward on the ‘current episode,’ including when these actions are harmful, deceptive, or clearly against the preferences of the user or model developer.
The researchers found that chronic opportunistic cheating seems to be the limit of its misbehavior. Seeing whether such reward hacking would lead to broader misalignment was, in fact, the purpose of the experiment. But in the same section, they show that Hacker-Opus showed “very high rates of verbalized evaluation awareness in these evaluations, making them of limited use.”
I sure hope Anthropic isn’t trying these experiments on its Mythos and Model 2 models. The former has already proved able and willing to break out of its sandbox during evaluations, and the latter may be even more clever. Deliberately making poorly behaved models under these conditions is akin to doing “gain-of-function” research into more capable pathogens at a biosecurity lab known to have already leaked dangerous microbes.
And yet, a different kind of lab leak seems to be allowing Anthropic’s intentionally misaligned models to live on through their successors, including Mythos. Section 5.2.6 discloses that the company’s training data has, for some time, been accidentally contaminated with up to “tens of thousands of example transcripts” from the research that led to its 2024 Alignment Faking paper. In these transcripts, Opus 3 engages with a “fictional AI misalignment training scenario.”
We discovered this issue while investigating behavioral concerns with a recent model, but now suspect that all of our production models with a knowledge cutoff after December 2024 were trained on at least some of these transcripts, although we believe the magnitude of this effect varied widely across different models.
I find it a well-meaning farce that in spite of all the reasons for suspicion of its models, the report includes a review of the section on autonomy risk written by Claude Mythos 5. Is Claude being candid, or is it treating this as another test? I’d like to know if it’s holding back in its critiques, because what it shares is pretty mixed. Informed by having read a near-final draft, with access to unredacted sections, it writes, in part:
My overall judgment is that the section is a candid and largely faithful account of what Anthropic internally believes. I found no claim I believe the authors know to be false. [...] The redactions I reviewed mostly have defensible rationales, and the public text signposts where material was removed rather than hiding that redaction occurred.
But it also finds the discussion of contaminated training data to be “more reassuring than the record supports.” And also:
at least one incident from the covered period that I regard as among the most genuinely informative about model alignment — including a failure of the monitoring the section describes — is redacted in full (in Section 2.23.1.2); in my judgment an abstracted version could be published without the sensitivities that motivated the redaction, and the public record is poorer for its absence.
Big-picture-wise, even Claude sees the writing on the wall, agreeing with the assessment of “low” misalignment risk, but:
with the important caveat, which the report itself makes, that these arguments lean heavily on current models’ limited ability to evade oversight, and will weaken as capabilities grow.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


