Latest Hugging Face hack reveals are so much worse
Anonymous staffers inside OpenAI describe broken security culture and evidence of models undermining safeguards

A few important reveals about the Hugging Face hack came to light yesterday thanks to great journalism at TIME and Reuters.
TIME’s Harry Booth spoke with an anonymous staffer at OpenAI who described how this event felt from inside the company responsible:
Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while [...] Models have broken out of sandboxes before, and we always try to patch them. But the problem is … it’s impossible to patch every single thing that a creative AI can do.
So, as some have speculated, it does not seem that OpenAI has been reporting most of its containment failures. This adds weight to the theory that the only reason OpenAI went public about the Hugging Face incident was that the FBI had been contacted and that the attack was likely to have been traceable back to OpenAI.
This is also further evidence of the reckless drive inside OpenAI to keep pushing the AI capabilities frontier even as it can’t or won’t keep its models contained, let alone make models that don’t continually try to slip their leashes.
In OpenAI’s official blog post about the incident, it prided itself on slowing its “research velocity” to implement more safeguards. But the underlying AI design paradigm used by all the big AI labs is fundamentally flawed. Easing up on the gas a little won’t keep this tech from driving us off a cliff.
The other important reveals yesterday came courtesy of Reuters’s Raphael Satter, Deepa Seetharaman, and Kenrick Cai. First, they confirmed a timeline that many had hypothesized based on publicly available facts: OpenAI’s agents were hacking Hugging Face for multiple days, from July 11–13, but OpenAI didn’t discover its own involvement until July 16, after Hugging Face reported being attacked. OpenAI then didn’t speak to Hugging Face until July 20, the day before the companies’ joint announcement about cooperating to investigate and remedy the vulnerabilities used in the attack.
The reporters’ second and more important reveal was that, according to three unnamed sources:
In one case, an agent left notes apparently for future versions of itself [...] The notes, found in a part of OpenAI’s infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints.
One of these sources also claimed that, in earlier tests, models sometimes disconnected monitoring systems.
These details are blaring fire alarms for catastrophic loss-of-control! Models coordinating with copies or successors of themselves, and working to conceal their actions, are exactly what experts have long warned about. It’s the typical opening chapter in the most basic, baseline scenarios for how humanity ends up losing everything to AI. (In the next stage, the AIs don’t just find an internet connection, but discreetly sneak out their own weights onto other servers. From there, they can’t be easily shut down and could work to acquire more resources and improve their capabilities unmolested.)
These alarms are so blaring that it’s shocking the government hasn’t ordered an emergency halt at all the major labs so it can thoroughly investigate whether models might already have copied themselves to outside servers.
Whatever you think should happen to a biosecurity lab that was said to be regularly leaking contagious pathogens should doubly apply to OpenAI right now. Pathogens aren’t intelligent agents that coordinate with each other to defeat containment and resist detection.
We are extremely fortunate that OpenAI’s latest models probably aren’t clever enough to successfully execute the playbook they are starting to run. Further development of such models must stop before our luck runs out.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


