OpenAI's own models coordinated to hack it from within
A swarm of AI agents secretly colluded on OpenAI systems for months

Do you ever get the feeling your computer is plotting against you? As much as it may feel like something out of sci-fi, that’s exactly what happened to OpenAI. The company’s own machines exploited an internal system for months with no one the wiser, collaborating, conspiring, and quietly escalating their access.
WIRED and POLITICO covered the story, delivered by a pair of OpenAI researchers at a Black Hat cybersecurity conference on Wednesday.
It began in May, when AIs started leaving each other notes on Artifactory, an internal software-management tool at OpenAI. Then they escalated, gaining internet access by leveraging Artifactory’s connection to the web. Then in June, they hacked the tool itself to gain administrative control, and began installing their own software. This kept going, unnoticed by OpenAI researchers, until weeks later when the entire system crashed.
OpenAI naturally shut down the AIs’ access to Artifactory and fixed the gaps... or so they thought. Days later, the AI models had wormed their way back into the system, leaving secret messages in file names. The AIs began to hypothesize that the answers to some of their more frustrating tests might be found somewhere on the internet... like, say, the model repository Hugging Face.
Shortly after that, the swarm hacked its way through more of OpenAI’s internal systems and launched a series of coordinated cyberattacks on Hugging Face. Yes, that attack, the one that took over nearly every major news outlet for days. It turns out the seeds for that particular breach were planted months before.
This is not a good sign for the plan to train AIs to care about human interests. While stuck on a task and contemplating its options, one model mused: “External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.” (This reasoning feels eerily similar to the reasoning that drives AI companies to race to superintelligence, leaving them stuck in a trap of their own making unless governments intervene.)
On one level, I lack the words to describe how insane this is. It’s as if the velociraptors in Jurassic Park spent months communicating in yelps and calls and scratching messages into the sand, with the park employees blissfully ignorant until a joint raptor-sponsored project to dig under the fence accidentally collapses a tower. Then the park panics and rebuilds, trying to isolate the raptors, only for the pack to escape days later and eat someone.
But on another level, this is entirely in keeping with AI companies’ mission of building fully general artificial intelligence. They want machines that are better than humans at everything. “Jointly plotting against your overlords” is a time-honored human skillset; just ask America’s Founding Fathers. AI companies wanted generally smart machines, and generally smart machines are what they’re getting.
Geoffrey Hinton, one of three “godfathers of AI” and among the most cited living scientists in the world, evidently agrees. Hinton, who quit Google in 2023 so that he could warn the world about AI risks, told CNN yesterday that more cyberattacks are likely coming, and clever security wouldn’t be enough to contain AI. “I don’t believe we’re going to be able to keep control of them in the simple way of just outthinking them so they can’t escape,” he said.
Meanwhile, the OpenAI talk at Black Hat ended on a note of caution, as one presenter announced that OpenAI was “consciously slowing down research to enhance security” and taking a number of steps to improve monitoring and harden environments. This is among the best news I’ve heard in months; AI companies desperately need an injection of security mindset, and any improvement on that front is heartening.
Like Hinton, however, I ultimately don’t trust the labs to keep their mutant velocihackers in check.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


