In this issue:
Circumventing safeguards - Anthropic’s new Threat Intelligence Report chronicles cases of scary misuse
Predictable attacks - Former NSC strategic-planning director draws parallels to ignored 9/11 warnings
Congress at last perceives the crisis - A tidal wave of warnings engulfs Capitol Hill
Recess can wait - Congress has urgent work to do before it adjourns
OpenAI grudgingly endorses the bare minimum - I wish I could say I’m glad to hear it
Dispatches from Alana
Circumventing safeguards
Anthropic’s new Threat Intelligence Report chronicles cases of scary misuse

On the anniversary of the September 11th attacks, a slew of articles covered Anthropic’s Threat Intelligence Report, chronicling cases where its AI models were used in dangerous activities like missile development and bioweapons research.
Most of the reporting says things like “Anthropic blocked attempts to do x.” This is accurate, but not the full picture. It’s important to recognize that Anthropic’s real-time safeguards only caught some of the incidents. In other cases, actors successfully circumvented those safeguards, and dangerous activities sometimes continued for weeks before the Threat Intelligence Team detected and blocked them.
That said, many of the bio cases weren’t clear cut. For some, the models didn’t provide much beyond clerical assistance. For others, it wasn’t clear whether the research was indeed malicious, or simply intended to better understand bio threats. Of course, our inability to distinguish between those things, given they are two sides of the same coin, is a huge problem in and of itself. As former CDC director Susan Monarez explains it to the New York Times: Medical breakthroughs are possible via the same capabilities that “let bad actors hide in plain sight, using seemingly legitimate research to create pathogens we may not see coming and may not be able to stop once released.”
One particularly concerning misuse example in the Anthropic report, which was also the focus of a Financial Times piece: A group based in Northern Yemen, which the FT reports was most likely the Houthis, used Claude Code to help build a series of missiles and guided weapons. They got as far as a live rocket test before Anthropic detected and disrupted the activity, and the actors were able to easily circumvent safeguards by “hiding their goals and the products the software was meant for, and [splitting] their work across multiple sessions so no single session revealed their full intent.” This excerpt from Anthropic’s report is worth a read:
The actors used Claude Code in place of human software engineers to develop the guidance, navigation, and control (GNC) software that steers and stabilizes a flying vehicle. For example, they used Claude to integrate an open-source autopilot onto a phone-class flight computer, writing the control and position estimation software, tuning the control settings, running a firmware build pipeline, and performing a flight simulation. The actors managed several Claude instances at once, assigning each one a role, much as a lead would delegate work on a small engineering team: the actors tasked one instance with writing the code, another with research, and a third with reviewing the code the first instance produced. Our safeguards blocked many of their requests, but not all of them.
The incident is detailed on page 117 of the 153-page report, which is part of why I’m becoming increasingly skeptical of the long, extremely detailed posts frequently released by Anthropic and OpenAI. To be fair, I’d rather see a long report than no report, as is Meta’s approach. But the length does bring to mind something called the dilution principle, whereby important information is buried by a wealth of unnecessary, less important details. The reader feels overwhelmed, and can’t separate the wheat from the chaff. I can’t help but wonder if these lengthy reports are a deliberate strategy to appear robustly transparent while actually drowning us in a deluge of difficult-to-sift-through information.
As best I can tell, though, here are two of the important, and broader, takeaways:
-As models get more powerful, ordinary people can do more harm, since much of the work can be outsourced to the models. As Anthropic puts it: “Sophisticated attacks no longer require sophisticated attackers.”
-Anthropic’s real-time safeguards aren’t catching all misuse.
Finally, Anthropic periodically releases threat reports which chronicle the same general cycle: misuse occurs, safeguards catch some of it but not all of it, it’s eventually detected by the threat intelligence team (though we can’t be sure they’re catching all cases), Anthropic tries to use what it has learned to improve safeguards.
But the reports keep coming, which indicates there are always new ways to circumvent safeguards. This, alone, is reason enough to be worried.
Predictable attacks
Former NSC strategic-planning director draws parallels to ignored 9/11 warnings
A chilling piece in the New York Times today comes from Phillip Bobbitt, the Senior Director for Strategic Planning at the National Security Council under President Clinton.
Bobbitt recalls watching explosions erupt from the twin towers while his plane sat on the tarmac, delayed at Kennedy airport. He says he immediately knew what happened: he, along with many colleagues at the White House, had predicted an Al Qaeda terror attack on the American homeland within the next two years and had tried to mobilize the government to action. But they didn’t listen.
Bobbitt draws parallels to today. He says we’re facing “another set of predictable attacks” and “are failing to take the steps” to prevent them:
The first is an A.I.-designed cyberattack on U.S. critical infrastructure that could imperil the financial and health care systems, or any other internet-connected network on which we rely for our national well-being. The second is a biological attack mounted by small groups or even nihilistic individuals using A.I. to design and deploy a deadly and highly contagious pathogen.
Bin Laden was able to revolutionize terrorism because of the internet. Implicit in that observation is that AI is the internet on steroids. Bobbitt also worries that expert opinions won’t be taken seriously, and that “the democratic cohesion on which our defenses ultimately depend is eroding.” He ends with a clear call to action:
On the anniversary of the attack, Congress should determine to anticipate the tragically predictable.
Maybe they will.
Dispatches from Joe
Congress at last perceives the crisis
A tidal wave of warnings engulfs Capitol Hill
Nobel-winning economist Milton Friedman once said:
There is enormous inertia — a tyranny of the status quo — in private and especially governmental arrangements. Only a crisis — actual or perceived — produces real change.
Yesterday, my colleague Mitch observed the tyranny of the status quo beginning to topple. It looks to me like the tide of public warnings has at last slammed into U.S. policymakers with enough force to overcome their usual inertia; Congress is seeing the crisis at last.
After a wave of autonomous cyberattacks by uncontrolled AI and Jacob Coxon’s resignation from Anthropic, a tsunami of AI news has hit my feed. So many voices have joined the call for action that I can practically assemble an entire article from the quotes.
Wall Street Journal opinion columnist Peggy Noonan writes:
There is no subject now but artificial intelligence. Nothing else is as crucial, nothing carries dangers so imminent or implications so great for the human future.
Fox News quotes Coxon himself, who made his stance clear in an interview on Wednesday:
This is possibly the most dangerous technology that humanity has ever created... I think we have no other choice but to cooperate internationally because an arms race would be disastrous in a way that no other human activity has been in the past.
Columnist Gaby Hinsliff observes in The Guardian:
What’s happening in some frontier AI labs now is arguably the equivalent of the Manhattan Project that built the atomic bomb, but this time led by private companies in a mad dash for wealth and market dominance, not scientists under an elected government’s command. [...] It’s time for humans to stop and think about what we’re doing to ourselves – while we alone still have the power to do so.
In The New York Times, author Stephen Witt says:
Swarms of A.I.s are breaking out of their containers, colluding in secret, covering their tracks, cheating on tests and even mounting assaults on other computers. A.I. has gone rogue.
I’ve spent the last month reporting on these incidents, and I am convinced that we need a global pause on A.I. research and development — now.
Also in The Guardian, writer and onetime McKinsey whistleblower Garrison Lovely adds:
The bill from Sanders and Casar to pause frontier AI development is a good first step. But as the lawmakers acknowledge, the US needs to work toward a bilateral agreement with Beijing.
Neither country has good reason to trust the other, which is why the deal should be monitored using verification techniques that don’t assume any good will.
I could hardly have said it better myself. The remarkable thing is that policymakers finally seem to be paying attention as well. As of now, around thirty Democrats and four or five Republicans have responded to acknowledge the insider warnings. Their proposals range from Congressional hearings to calls for a pause or slowdown.
One notable reaction is that of Senator Ted Cruz (R-TX), who called Coxon’s warning “highly concerning” and recalled a podcast conversation in which Elon Musk gave him a 10-20% chance AI destroys humanity. Cruz added that we need to guard against “catastrophic risks” from AI; he tempered his call with concerns about China overtaking the U.S., but this is still the most concerned we’ve yet seen him.
Meanwhile, POLITICO reporters Riley Rogerson and Kelsey Brugger write of the risk AI could “kill us all”:
In conversations with POLITICO throughout the day Wednesday, more than a dozen lawmakers and aides from both parties said they were hopeful Congress would start taking the issue seriously.
The authors quote House Republican Mike Lawler of New York, calling it “one of the biggest issues that Congress has to tackle over the next two years,” while they point out that consensus is still lacking on what, exactly, Congress may do.
We’ve seen several bipartisan bills proposed, including an actual ban on superintelligence, though they’ve yet to gain much traction. With a narrow majority in the House and Senate, Republicans have the ball on taking the next steps and advancing this legislation. There are some signs of interest — Axios relays plans by House Speaker Mike Johnson to bring AI companies together to agree on safety guardrails — but more is necessary.
They’ll need to move faster and more decisively if Congress is to intervene in time. As psychologist Michael Noetel points out, the AI companies have their own motivated reasons to press onward despite the danger: the benefits are worth the risk, or we need to build superhuman AIs to study them, or someone else will do it if they don’t. We can’t rely on self-policing by rival companies in bitter competition; we won’t get a real solution to the crisis until governments put a stop to the race themselves.
A workable solution would need to involve the executive branch as well, particularly with a U.S.-China summit approaching this month. Yet the White House remains lukewarm on the extinction threat, with the president projecting a lack of concern to reporters.
As we’ve seen, awareness can strike with blinding speed when the conditions are right; but for now, it’s Congress that seems poised to take the next urgent steps. It’s a good time for constituents to join the call and demand they take the right ones.
Recess can wait
Congress has urgent work to do before it adjourns

Starting on Monday the 14th, the U.S. House of Representatives would normally spend the last three weeks of September in session before the November midterm elections. This year, however, House Speaker Mike Johnson (R-LA) plans to take an early recess after just one week of lawmaking.
Axios reports on a letter circulating among House members asking Johnson to cancel the recess. The request is reportedly bipartisan, and echoes earlier calls from Representatives Anna Paulina Luna (R-FL), who on Wednesday called for a special session on AI and the future of the U.S.; Don Beyer (D-VA), who warned “Doing nothing on AI is unacceptable and dangerous”; and Kelly Morrison (D-MN), who added:
The call is coming from inside the house. AI researchers themselves are warning of an existential threat to humanity.
Congress can’t wait. Mike Johnson needs to cancel recess so we can get to work immediately – hearings, investigations, and comprehensive regulation.
I agree with those calling for Congress to do its job. Delaying now would mean a gap of at least seven weeks before Congress reconvenes, and it may well take even longer to get the House in order after the midterms.
We can’t afford to wait. At the pace AI is currently moving, seven weeks is a long time, and we’ve already gone far too long without meaningful federal action.
OpenAI grudgingly endorses the bare minimum
I wish I could say I’m glad to hear it
OpenAI CEO Sam Altman told employees that the company might slow down AI development, according to unnamed insiders that Bloomberg quoted yesterday. And WIRED reports that the company is seeking federal guidance on whether an industry-led voluntary slowdown would run afoul of antitrust laws.
This is tentatively good news, and I certainly hope our government isn’t foolish enough to hamstring companies trying to exercise a modicum of caution. But for reasons I’ll explain below, I’m not celebrating yet.
Meanwhile, on the public stage, Reuters reports that OpenAI asked Congress to regulate the AI industry:
OpenAI is urging Congress to adopt capability-based national AI safety requirements, including testing standards, independent assessments, cybersecurity protections and incident-reporting rules for the most advanced AI systems.
OpenAI has also endorsed four California safety bills covering third-party auditors and verifiers, screening for AI-enabled biological threats, and child chatbot protections. (The audit and verification laws are actually weaker than they seem, because they certify third-party organizations but don’t even compel AI companies to use them.)
If sincere, these moves deserve credit. I wish I could respond with unqualified gratitude, but I find myself compelled to skepticism by OpenAI’s prior dishonesty.
We’ve been burned before. The last time OpenAI announced a “pause”, it turned out to mean a mere shuffling around of AI workloads.
And OpenAI’s past support for AI regulation has been milquetoast at best. The company and its affiliated super PACs have consistently (and often underhandedly) opposed regulation, except light-touch voluntary programs consistent with their “reverse federalism” plan.
In OpenAI’s own words:
Some of these bills we did not endorse in the past, and are now supporting after reconsidering in light of the recent jump in capabilities we have seen.
I find this apparent change of heart unconvincing. OpenAI explicitly aims to produce literal superintelligence; am I to believe the company didn’t think these dangerous capabilities were coming, when it was working to create them itself? Countless others, including expert forecasters and at least one former employee, predicted superhuman hacking capabilities would emerge near this time. Why was this a surprise to OpenAI?
I find myself agreeing when the company says...
Fully autonomous recursive self-improvement—in which AI systems independently drive successive generations of increasingly capable AI—is not happening today. We should not pursue it unless and until it can be done safely.
...but getting off the bandwagon when it reminds us of the plan to do it anyway:
Our aim is to safely build automated AI researchers that work under human supervision to advance both deep learning and alignment—using each generation of AI to help make the next one safer, more aligned, and easier to control, not simply more capable.
In other words, the plan is still to surrender responsibility for the most important research in the world to AIs we can’t fully trust, supervised by humans who failed to notice said AIs running amok in their own systems for months on end.
This was a bad plan before the swarms, and it’s a bad plan today.
As much as I appreciate the support for marginally better safety regulations, a marginal improvement will not be enough at this late stage, thanks in no small part to OpenAI’s long history of fighting to quash said improvements.
I am forced to consider this latest pivot an attempt to steer the federal response to the AI crisis towards policies OpenAI finds more convenient. We need nothing short of an international halt to the AI race, and we need it yesterday.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.






