
On the anniversary of the September 11th attacks, a slew of articles covered Anthropic’s Threat Intelligence Report, chronicling cases where its AI models were used in dangerous activities like missile development and bioweapons research.
Most of the reporting says things like “Anthropic blocked attempts to do x.” This is accurate, but not the full picture. It’s important to recognize that Anthropic’s real-time safeguards only caught some of the incidents. In other cases, actors successfully circumvented those safeguards, and dangerous activities sometimes continued for weeks before the Threat Intelligence Team detected and blocked them.
That said, many of the bio cases weren’t clear cut. For some, the models didn’t provide much beyond clerical assistance. For others, it wasn’t clear whether the research was indeed malicious, or simply intended to better understand bio threats. Of course, our inability to distinguish between those things, given they are two sides of the same coin, is a huge problem in and of itself. As former CDC director Susan Monarez explains it to the New York Times: Medical breakthroughs are possible via the same capabilities that “let bad actors hide in plain sight, using seemingly legitimate research to create pathogens we may not see coming and may not be able to stop once released.”
One particularly concerning misuse example in the Anthropic report, which was also the focus of a Financial Times piece: A group based in Northern Yemen, which the FT reports was most likely the Houthis, used Claude Code to help build a series of missiles and guided weapons. They got as far as a live rocket test before Anthropic detected and disrupted the activity, and the actors were able to easily circumvent safeguards by “hiding their goals and the products the software was meant for, and [splitting] their work across multiple sessions so no single session revealed their full intent.” This excerpt from Anthropic’s report is worth a read:
The actors used Claude Code in place of human software engineers to develop the guidance, navigation, and control (GNC) software that steers and stabilizes a flying vehicle. For example, they used Claude to integrate an open-source autopilot onto a phone-class flight computer, writing the control and position estimation software, tuning the control settings, running a firmware build pipeline, and performing a flight simulation. The actors managed several Claude instances at once, assigning each one a role, much as a lead would delegate work on a small engineering team: the actors tasked one instance with writing the code, another with research, and a third with reviewing the code the first instance produced. Our safeguards blocked many of their requests, but not all of them.
The incident is detailed on page 117 of the 153-page report, which is part of why I’m becoming increasingly skeptical of the long, extremely detailed posts frequently released by Anthropic and OpenAI. To be fair, I’d rather see a long report than no report, as is Meta’s approach. But the length does bring to mind something called the dilution principle, whereby important information is buried by a wealth of unnecessary, less important details. The reader feels overwhelmed, and can’t separate the wheat from the chaff. I can’t help but wonder if these lengthy reports are a deliberate strategy to appear robustly transparent while actually drowning us in a deluge of difficult-to-sift-through information.
As best I can tell, though, here are two of the important, and broader, takeaways:
-As models get more powerful, ordinary people can do more harm, since much of the work can be outsourced to the models. As Anthropic puts it: “Sophisticated attacks no longer require sophisticated attackers.”
-Anthropic’s real-time safeguards aren’t catching all misuse.
Finally, Anthropic periodically releases threat reports which chronicle the same general cycle: misuse occurs, safeguards catch some of it but not all of it, it’s eventually detected by the threat intelligence team (though we can’t be sure they’re catching all cases), Anthropic tries to use what it has learned to improve safeguards.
But the reports keep coming, which indicates there are always new ways to circumvent safeguards. This, alone, is reason enough to be worried.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


