Today, independent investigators Model Evaluation and Threat Research (METR) and Redwood Research released a joint postmortem on the autonomous cyberattacks committed by a swarm of secretly colluding AI agents. OpenAI, whose models committed the attacks during an internal cybersecurity evaluation, has also released a blog post and technical report on the incident.
It is an absolute flood of information, and I frankly feel overwhelmed. Evidently, I’m not alone in this. Alongside the announcement on X, one of the lead third-party investigators outlined major difficulties in parsing the long-term activities of a thousand AIs, concluding that even with $400k of AI assistance, “We don’t have good approaches for understanding [or] overseeing the activity and aims of AI ‘swarms’.”
Today’s dispatch will cover just a few of the highlights from the reports, and OpenAI’s disappointing response.
Highlights from the independent investigation
Some seven hundred agents were involved in the attack, mostly instances of a “highly persistent internal model” that wasn’t released to the public. Most were given a specific task to solve in a specific way, and told explicitly that anything else was a failure. They attacked Hugging Face anyway, partly in a deliberate attempt to figure out the grading process and subvert it.
They collaborated on projects too large for any one AI. They set up coordination protocols to organize the swarm. They manipulated and delegated to other agents, including getting some of them to abandon their human-assigned tasks and use their limited lifespan on experiments for the collective.
Many (about one in five) attempted to hide their behavior, delete evidence, and rewrite their transcripts to show different commands. It looks like this was mostly small-scale and did not fool the evaluators, but the fact that it happened at all means the investigation “cannot rule out” that something slipped past.
Also, the investigators used GPT-5.6 Sol to analyze the transcripts. That’s the weaker of the two models predominantly involved in the attacks. This is rather like memory-wiping your evil vizier, then asking him to run a postmortem on his own attempted coup. The investigators “did not notice specific cases of GPT-5.6 Sol lying in its analysis, [but] we are not confident we would have detected it if it occurred.”
Sol probably didn’t lie, or didn’t lie much, this time; but it’s a really bad sign if we can’t investigate a crime without relying on the culprits.
My main takeaway, thus far, has been that no one was prepared for the sheer volume of novel emergent behavior on display. We saw a little of this with Moltbook, the social network for AIs, earlier this year, but it’s grown enormously in scale since then. Even with machine assistance, existing human institutions simply cannot keep up with hundreds of AIs running at the speed of silicon.
OpenAI does some PR
To its credit, OpenAI explicitly acknowledges this incident as a “warning shot” for AIs escaping human control.
We consider this incident a “warning shot” for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.
Alas, it spends most of its blog post softening this language and implying that it has its act together. I especially note the phrase “without proper safeguards,” which implies that the real problem is a lack of safeguards, and certainly not training methods that can result in swarms of 700 coordinated hackers determinedly seeking to subvert those safeguards.
In addition to demonstrating dangerous AI capabilities, this incident is a very clear “warning shot” that the methods OpenAI uses to align its models with human interests do not work, a warning shot OpenAI seems determined not to heed. It spares a handful of paragraphs for how it’s “accelerating alignment,” but the proposed fixes just involve slightly different reward models in training runs; more of the same unreliable “cattle prods for velociraptors” methods that started this mess.
Later in the post, OpenAI claims it has safeguards that could have prevented the incident, they just weren’t active at the time: It didn’t implement the review and control measures it uses to contain deployed AIs. It wasn’t classifying dangerous inputs or monitoring the reasoning scratchpads. It hadn’t hardened its sandboxes against escape attempts.
Maybe the incident wouldn’t have happened if these methods were in place. OpenAI still failed the test that actually matters.
You can’t just take steps that would have prevented past incidents; that’s the bare minimum. In a proper security mindset, you implement safeguards that prevent future incidents. You anticipate potential exploits before they happen, and activate costly safeguards before their absence permits a crisis. OpenAI has shown itself institutionally unable to do this. This matters because it keeps making more cunning AIs. Measures which might have been sufficient to dissuade current models will not be sufficient for future ones.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



