
After reading Anthropic’s 8/31 blog post on the measures they’ll be taking to “improve alignment and security efforts,” I agree with my colleague Mitch’s take from a couple weeks ago: there’s clownery going on.
How else do we describe a company that prides itself on conscientiousness and safety, but openly admits it hasn’t been putting safety first when it comes into tension with speed?
From the post:
We put significant effort into making sure our tooling evolved with each model generation, from a handful of classifiers to an automated review of all environments before and during training runs, but by spring 2026 that system was under strain. We were producing RL environments at a more rapid pace than ever, faster than our systems could vet them. Flagged environments required a human to adjudicate, and reward hacks and misconfigurations started outpacing our ability to filter or fix them.
How else do we describe a company that is still in the early research stages of understanding how misalignment arises, but continues to make models capable of launching cyber and bio attacks?
From the post:
We also believe lasting progress comes not only from understanding what happened in a given incident but from understanding how misalignment arises in the first place, and we share early research in that direction.
(Spoiler alert: the training methods all companies use result in models that are misaligned by default. Anthropic surely knows this, but has convinced itself that clever patches and guardrails are good enough to relegate such bad behavior to the “usually won’t show up” category. Perhaps Anthropic has deemed it good enough because it’s, sadly, the industry standard.)
How else do we describe a company that makes seemingly careless and widespread errors?
From the post:
In addition, we found a fraction of runs had trained on a model’s chain-of-thought by accident—something we work hard to avoid, out of concern that it teaches a model to hide its actual reasoning. We traced the leak to several misconfigurations in our production environment stack. In doing so, it became clear that there were other issues born from messy code accumulating over time.
Anthropic also states that over 10% of the “environments in our production mix” had “problems ranging from reward hacking to broken tasks and misconfiguration” and that “some flawed environments simply slip through our detection mechanisms.”
And how else do we describe a company that continues to create AI models it can’t fully understand or steer, convincing itself that some improvements to its monitoring and containment practices are all that’s needed?
It sure seems like clownery to me, the horror movie kind.
To be clear, I’m not trying to say that Anthropic’s efforts are all bad. I’m glad the company looked for issues, found them, fixed misconfigurations and flawed environments, and took steps to try to improve security. Some of these steps may well prevent some harmful behaviors from the current generation of models, and I applaud that. But I very much doubt they will be good enough, especially as capabilities increase. And I think Anthropic knows they won’t be enough, but has simply habituated to extreme levels of risk.
The primary measure to “improve containment” seems to be a classifier Anthropic built “to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access.” They also state they are monitoring transcripts to find escapes from past evaluations, moving high-risk sandboxes to “more robust isolation” and telling models to try to escape in order to find issues before they are a problem.
A classifier is a great idea in theory, but will it be robust to the efforts of tenacious AIs who don’t tire and have plenty of time on their hands? Even without Anthropic’s history of accidental misconfigurations and errors, I’d bet good money that a smart AI will find a way to evade the classifier, just as smart AIs have exploited vulnerabilities in the previously-thought-to-be robust infrastructure of all the main AI companies.
If AI companies — either knowingly or unknowingly — push models past human intelligence, I’d bet all my money on these models doing whatever they want, with nobody the wiser until the other shoe drops.
Finally, let’s talk about reward hacking, a key part of Anthropic’s post. This is a known behavior in AI models, and it arises from the way they are trained.
Basically, AI models are scored highly in training for correct answers, and for elements that the evaluators think might lead to correct answers. Sometimes, the models figure out a way to get the high score without actually completing the task itself. As my colleague Beck puts it, it’s like placing all your darts in the center of the bullseye rather than throwing them.
Reward hacking has been heavily discussed in light of the Hugging Face attack, where agents colluded with each other to get the answer to the task, and then again, to manipulate the scorer so it wouldn’t realize they had cheated.
Anthropic writes that it believes training environments prone to reward hacking are linked to higher rates of misaligned behavior, as are tasks that are impossible to pass without cheating. The company has “tried to avoid training environments that have these sorts of defects” and hypothesizes that these efforts “may have prevented more severe alignment incidents, and conversely that the imperfections in these efforts may have contributed to the incidents we have identified to date.”
My reading of this is that Anthropic is trying to reassure us: if misalignment is linked to reward hacking, and we avoid reward hacking, maybe we can avoid misalignment.
But I’m not reassured. First, as Anthropic correctly states, reward hacking is one contributor to harmful actions, certainly not the whole story. Second, if very hard or impossible problems “are disproportionately large contributors to misaligned behavior,” that doesn’t bode well for releasing models into the real, complicated, messy world where problems are often hard, and sometimes impossible. The linkage of “difficult problem” and “harmful behavior by AI” is especially concerning given AI models in the military, sciences, and medicine.
Instead of trying to achieve “perfect” training environments where reward hacking isn’t possible (which will likely just hide the problem until the stakes are much higher), Anthropic should stop pushing its models forward in capabilities until we have a much more robust solution to problems like this.
To that point, Anthropic credits “the substantial investment we made this spring into monitoring and reducing reward hacking” as “a major reason our production models are unlikely to engage in more dangerous reward seeking efforts” while also noting “the process isn’t perfect and the models are not perfectly aligned.”
To me, this feels like saying a leaky pipe with an imperfect patch will lose less water than one with no patch. I wouldn’t be satisfied with that fix in my home, and I’m certainly not satisfied with it as a fix for the world — our collective home.
The way we currently grow and train AI systems produces strange, unpredictable behaviors, including reward hacking, resource acquisition, and tenacious pursuit of whatever goals they end up with. The pipe leaks by default. We shouldn’t just keep building with leaky pipes and hope increasingly sophisticated patches will hold. Instead, we should wait until we know how to get pipes that don’t leak in the first place.
One note of hope: Anthropic’s lackluster pause of RL training (which has now been mostly reinstated) seems to echo OpenAI’s lackluster pause of RL training (which has also now been mostly reinstated). The “improvements in monitoring and containment” (quotes because I don’t see these as scalable improvements) also seem to echo OpenAI’s.
One company’s micro-pause set precedent for another company’s micro-pause. If a major AI company is ever morally consistent enough to take more serious action than a micro-pause, other companies may face pressure to do the same. Currently, the race to the bottom seems to be turning into a race to the bottom with safety washing. Could that turn into a race to meaningfully pause development?
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


