Hidden in the small print
OpenAI updated the Hugging Face incident report and there's an important detail hidden in there
Yesterday, OpenAI published an update to its incident post regarding the Hugging Face hack. The update reveals that the scope of the attacks carried out by the unpublished OpenAI model was more extensive than initially thought. Four other providers were also affected, among them Modal, an AI software infrastructure provider. Hugging Face’s detailed technical timeline of the incident confirms this, as does a statement by Modal to Reuters.
As Modal reports, OpenAI’s AI agent hijacked a Modal customer’s sandbox — which was publicly accessible due to a faulty setting — and used it as a launchpad for its attacks on Hugging Face. Modal insists that its own security mechanisms did not fail.
You can think of it this way: A gang of bank robbers (OpenAI’s rogue AI) scouts out a self-storage facility (Modal) next to a bank (Hugging Face) and discovers that one of the customers has incorrectly configured the locking mechanism for their storage unit. The bank robbers then take advantage of this to set up camp there and drill through the wall into the bank’s vault.
OpenAI stresses that the other providers weren’t affected to the same extent. Yes, the attack on Hugging Face was of a different caliber, and the Modal infrastructure itself was likely not actually compromised. But honestly, I don’t see what’s reassuring about the fact that the AI agent carried out its attack in an even more strategic and indirect manner than originally thought. Especially since we still don’t know who the other three providers were or how the AI agent used them.
One remarkable detail hasn’t been mentioned at all in the coverage so far. On Sunday, we learned that an anonymous OpenAI insider told TIME that this wasn’t the first breach, but that incidents like this have occurred several times before, though presumably had been not as severe. OpenAI now seems to confirm this quite casually, hidden in a parenthesis, in its update:
In our ongoing review of the Hugging Face intrusion and broader activity from our models, we have been finding a small number of cases where the models identified and used publicly exposed credentials at the account level on other publicly available services. This includes four accounts on four services as part of the Hugging Face incident (and a few accounts accessed as part of other evaluations)
To me, this reads as though sandbox escapes during other evaluations had in fact occurred even before the Hugging Face hack, and that in these cases — as in the case of Modal — the AI agents hacked into third-party services at the account level. To do what? To whom? What happened in these cases? Why wasn’t this already reason enough for OpenAI to pull the plug? Because, unlike in the Hugging Face case, the authorities weren’t notified? We don’t know.
I can’t help but feel that OpenAI shouldn’t be the one conducting this investigation. The independent non-profit research organization METR (Model Evaluation & Threat Research) warned in May about something like this happening and now reaches the same conclusion:
For public trust and clarity, this investigation would ideally be conducted or deeply reviewed by independent researchers, who can view evidence that companies would prefer not to share publicly.
Right now, evidence isn’t being independently reviewed. This has to change and it shouldn’t play a role whether OpenAI wants oversight or not. OpenAI wants us to trust them with humanity’s future, but we can’t even trust them to be fully transparent about their mistakes.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



