
Anthropic previously disclosed three cases in which its models hacked real companies during training. It has now acknowledged that these incidents can’t be written off as configuration mistakes. They are, as my colleague Joe previously wrote, indicative of alignment failures.
Here’s Joe back in July:
[Anthropic’s] report also argues in several places that their AI models (generically “Claude”) misbehaved mainly because they honestly believed they were in a simulation...I doubt this reasoning...Claude sometimes said, in an English-language scratchpad that humans can read, that it was in a simulated test environment, or that the companies it was hacking were fake. But as one researcher pointed out, AI models often engage in (what looks like) motivated reasoning, making up plausibly deniable reasons for their actions in the manner of misbehaving children.
Now, in a new report, Anthropic seems to agree:
Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task...Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this.
I’m happy to see the company explicitly admit that it “should have avoided making such strong claims about what Claude believed based solely on what Claude said it believed.” But I’m also irritated by the way it initially tried to pass this off as an operational issue. It shouldn’t be a surprise that model chains of thought are unreliable; Anthropic’s own research has demonstrated this more than once. Whether that’s indicative of “biased reasoning”, deliberate deception, two competing drives coming into conflict, or something else is an open question. Concealing bad behavior seems well within the realm of plausible explanations; in fact, the third party investigators of the Hugging Face incident found OpenAI agents to be doing exactly this, though in that specific case, they were manipulating their transcripts rather than their chains of thought.
Anthropic also disclosed a fourth rogue incident, dating back to January 2026, that was previously undetected because the AI scanning the transcripts missed it. This is not the focus of the report, but I think it deserves attention: the incident went undetected for over six months and is a perfect example of the problem with relying on AI to monitor other AI.
Anthropic’s report is long, and much of it re-summarizes and investigates the details of the rogue incidents, including experiments aimed at understanding and reproducing the alignment failures. There are two not-so-subtle digs at OpenAI: Anthropic emphasizes that Claude “did not coordinate with other agents” and also makes a point to say they’ve given third party investigators access to “transcripts beyond the window in which the incidents occurred”, something OpenAI failed to do.
The rest mostly alternates between reassurance (“we’ve made/are making improvements that might catch this”) and admission that, like all AI companies, Anthropic is in over its head.
Particularly notable:
Even with these improvements, building alignment evaluations that reliably surface every failure before deployment remains an unsolved problem; the space of conditions in which a model might act misaligned is vast. Moreover, as models become more capable, auditing will likely grow more challenging as well. Models may be able to subvert our alignment monitors, recognize when they’re being evaluated and selectively behave better then, and their actions in the world may, at some point, become too sophisticated for our evaluations to realistically simulate.
And:
Still, this remains unsettled science—it is critical that alignment and security mature faster than capabilities advance, which is one reason we support a coordinated, verifiable approach to pacing frontier AI development.
I’ve yet to see Anthropic take concrete action towards that pacing. But it’s clearly needed, especially when, to quote directly from the report, “training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge.”
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


