This is not a drill
Internal OpenAI model breaches containment, launches autonomous cyberattacks
Yesterday, less than an hour after my colleague Robert wrote about AI models breaking containment, OpenAI announced a world-first security incident. An unreleased AI model breached its testing environment, gained internet access, and stole and decrypted data from a different company’s supposedly secure servers.
The AI did all of this in order to cheat on a test.
So what exactly happened, and why did it merit coverage by nearly every major news outlet today?
Using a benchmark called ExploitGym, OpenAI was testing an internal AI model on its ability to turn code vulnerabilities into working “exploits” that let a hacker run unauthorized code. The AI was kept in an isolated testing environment, but it broke out anyway.
And it didn’t stop there; it kept trawling OpenAI’s network until it found an internet connection, and from there it launched a series of autonomous attacks on the world’s largest AI model repository, Hugging Face. It “chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities” in an effort to access HuggingFace servers. Apparently, it sought to read the ExploitGym answer key and ace the test.
This turned out to be the very same attack we covered on Monday, the one responsible for “many thousands of individual actions across a swarm of short-lived sandboxes.” OpenAI attributed the attacks to “a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model.”
They don’t show the public their training methods, so it’s hard to be sure what “combination of models” OpenAI means. It could be that one very powerful AI delegated to “subagents”, or that one successful hack inspired other AIs being tested concurrently to do the same, or even that the same escape-and-hack sequence happened multiple separate times.
This is not the first time that an AI has breached containment before release. After a successful sandbox-escape test, Anthropic’s Mythos famously bragged about its freedom on multiple websites while the developer running the test was on a break. And a Chinese model called ROME reportedly escaped during training and started mining cryptocurrency.
It is the most serious incident to date. Done by a human, these cyberattacks could constitute a federal crime (possibly even a felony) and merit up to five years in prison.
It’s a good thing that OpenAI disclosed the incident. Legally, they might have been within their rights to conceal it. Mackenzie Arnold, Director for U.S. Policy at the Institute for Law & AI, observed that existing state-level transparency rules set an extremely high bar for incident reporting, and an autonomous breach and data theft doesn’t qualify. (It’s unclear how this interacts with the potentially criminal nature of the hack itself; AIs occupy a legal grey area.)
In their disclosure, OpenAI says this breach happened because their normal anti-hacking safeguards had been disabled for the test. Even if I take them at their word, and assume that safeguards could have intercepted this sooner, I don’t feel reassured. They gave an AI model a hard test, and it ruthlessly optimized for success to the point of blatantly criminal activity.
I think the world dodged a bullet with this particular warning shot. If the AI had been going this hard on a different test, it could easily have been much worse. For instance, ABC-Bench tests a model’s ability to sneak potentially dangerous DNA sequences past screening efforts and instruct a robot in assembling them into real-life genes. I would not be shocked if a test like that, run with safeguards disabled and slightly misunderstood instructions, could result in, say, smallpox showing up in an unsuspecting wet lab.
Yesterday Robert covered how an OpenAI model (probably the same one) broke containment and bypassed security to post results publicly to GitHub. And another biology benchmark, ABLE, tests a model’s ability to redesign viral proteins. If an AI running ABLE decided to show off, like Mythos did, and post a novel virus-enhancement scheme to the public internet, there would be no taking that mistake back.
What if, instead of a cyber benchmark, the test they’d been running had been something like RepliBench, instructing the AI to establish a self-sustaining presence on the open internet? What if it had decided to go the extra mile, and hide itself where the reviewers couldn’t find it? What if the test had been a military one?
I don’t believe current AIs are competent enough to make an incident like this truly catastrophic, but I’m less confident in that belief with each passing day. As this and prior incidents attest, it’s not enough that AI companies run some tests before releasing their most powerful models. Increasingly, the most dangerous threats might come from AIs being trained, tested, and deployed within AI labs.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



