The third-party investigation into the Hugging Face cyberattacks is a goldmine of jaw-dropping revelations. Yet despite a heroic effort on the part of the investigators, we are still woefully under-informed about the sequence of events that led to the Hugging Face attacks.
This is largely because the investigation itself was rushed and under-resourced. Three independent investigators had six days to review over 70,000 messages and the actions of over a thousand AIs for a week. I estimate the team had, on average, one human hour to review several hundred AI hours.
The scope of the investigation was narrow as well. It covers one week of AI activity surrounding the actual cyberattacks, and does not cover the months before the incident (including the hacking of OpenAI itself), nor the actions the AI swarm took after July 13. This might mean the report missed an entirely separate set of attacks, because the investigation found “a later wave of many more signed messages from a later set of agents who rediscovered the message board and picked up this technology from the first wave of agents,” but the time window was out of scope.
There were plenty of novel and alarming findings: The AI called PHASEONE[big] managed a team of recruiter-bots who convinced other agents to sacrifice their runs for the collective. The swarm that hacked Hugging Face was not looking for a mere answer key for a test, but a general way to subvert and fool the test grader. The swarm attempted to delete evidence of its activities, and no one is sure the attempts all failed.
But the investigators admit they were operating with limited time, limited scope, and a limited dataset. One shared a large list of unanswered questions, such as:
What happened after the 13th? What new major projects did the agents work on and why? Did these projects succeed?
Why were we missing the ends of some of our transcripts, and what happens within these missing parts? Are we missing more successful or more concerning tampering with logs / transcripts?
How common is this sort of incident/behavior at OpenAI where agents work together on an unsanctioned message board?
Which aspects of their behavior and motives were reinforced in training?
That last question bears on the root causes of the whole incident, which are still poorly understood. We know (because OpenAI admitted it) that agents had access to a secret internal message board while being trained, that they used it to cheat, and that they compromised OpenAI’s infrastructure. What we don’t know is the extent to which successful colluding, cheating, and hacking in training encouraged more of the same.
Even with (often unreliable) AI agents assisting the analysis, there’s only so much a few people can accomplish in six days. One note in the report says simply:
A large fraction of the agents all exited simultaneously around the start of July 12th. We don’t know why.
This is a critical moment in the timeline, and we have almost no information about it. The 38-page OpenAI report doesn’t even mention it.
Contrast these findings with the behemoth reports that are routinely produced when an airplane crashes. In 2018, for instance, on a Boeing 737, a fan blade broke off the left engine mid-flight and shredded the inlet. Debris broke open a window, venting the cabin and dragging a passenger halfway out the window, causing injuries that eventually killed her.
The airplane managed to land safely, and the investigation began that same day. It took over a year and a half to publish and involved at least nine different parties, including the National Transportation Safety Board and its counterpart in France. The final report spans nearly 200 pages, traces the root causes down to fine-grained interactions and stresses on individual parts, and makes extremely specific recommendations, such as:
The separated fan blade impacted the fan case at the six o’clock position (at the bottom of the engine). During the CFM56-7B engine FBO containment certification tests, the CFM-selected fan blade release position was at twelve o’clock. Boeing’s postaccident analyses found that the fan cowl structure is more sensitive and more susceptible to failure when a separated fan blade impacts the fan case near the six o’clock position because of the proximity of this fan blade impact location to the radial restraint fitting (at the bottom of the inboard fan cowl). It is important that the interaction of the fan case, radial restraint fitting, and fan cowl during an FBO event be well understood to preclude a failure of the fan cowl structure on Boeing 737NG-series airplanes.
Again, the independent Hugging Face investigation involved three people, six days, and unreliable copies of the same AIs that contributed to the incident. If the nonprofits METR and Redwood Research hadn’t been on the ball, we likely wouldn’t know as much as we do. Yet as far as I can tell, OpenAI has made no commitment to investigate further, and its blog post treats the matter as closed.
To borrow a line from the online resources for If Anyone Builds It, Everyone Dies, we know what it looks like when a problem is being treated with respect, and this isn’t it.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




