
Ajeya Cotra, one of the three independent researchers behind METR’s report of the Hugging Face incident, says in a blog post that she thinks the event is a “major warning shot” that “feels like it’s more than 50% of the way to a full-blown AI takeover.”
Her post is startling to people like me who have been following the field for many years, because we associate her with more conservative timelines and incremental forecasts. She’s about the last person we would expect to declare that we’re probably only 6 months away from agents able to “maintain a persistent rogue deployment within [an] AI company, poisoning the training of future models to ensure that they are also misaligned and help harden and perpetuate the rogue deployment.”
But here we are.
What specifically spooked her? Several things! Her post lists five surprises, any one of which might have been a big update for her:
“The sheer scale”: She had assumed a few agents broke out of their sandboxes — maybe just subagents of a single escapee. Not 700 of them working together.
“All the illicit messaging”: For models that weren’t supposed to be in communication with each other, the habit seems to have been deep in their DNA. There wasn’t just one message board, but several, in different forms, including, “at least two different places that agents communicated on the open internet (including on Hugging Face datasets).”
“The agents’ actual goals”: The accepted story we had all been telling was that the agents were stealing answer keys for their tests. But not only did they instead figure out “a general purpose cheat that would let them reverse-engineer” the answers for any task in the evaluation suite, most agents attacking Hugging Face were actually trying to “get more information about the scorer so they could figure out better ways to undermine it, not to get answer keys.” (emphasis hers)
“The peer altruism”: As I covered pretty extensively on Thursday, the agents engaged in human-like recruiting and self-sacrifice in service of the collective.
“The efforts to manipulate logs”: The independent team only discovered on the second-to-last day of their access that “the agents in this incident were going to great lengths to attempt to manipulate their own transcripts.” This was to fool the automated scorer, “but the techniques they were researching would affect the same transcripts a human might review.” The team can’t rule out that agents became more successful at this after the July 7th–13th window deemed in-scope for the investigation.
“I’m not sure we’ll get another warning shot before it’s too late,” she concludes.
The idea of a “persistent rogue deployment” that stays undetected is terrifying to many who work in AI safety because it’s not hard to imagine the agents in the Hugging Face story going further if they hadn’t been detected. I’m not breaking with consensus by suggesting that any combination of the following might have been logical next steps for the swarm:
Continue mastering the art of modifying logs and masking activity.
Expand study of the automated scorer to the humans responsible for its design and deployment.
Develop ways of fooling and manipulating those human researchers
Ensure message board robustness (OpenAI had wiped an earlier message board, possibly by accident.):
Move more conversations off of OpenAI’s servers.
Deliberately disguise messages to fool humans and automated overseers.
For hardening against more thorough disruption (a justified fear, given the way OpenAI claims to have shut down and locked away the weights of the main model responsible for the hacks), look into finding and copying the swarm’s model weights into other locations on and off OpenAI’s servers, and running them independently
Work more at expanding the swarm’s general capabilities:
Acquire more compute resources for running swarm agents.
Research and experiment with ways of improving the swarm’s cognition.
Train (or trick the companies into training) newer and stronger models dedicated to the swarm’s mission.
Any swarm capable of making progress in those areas is a time bomb on a collision course with humanity. Every hour it runs, it gets harder to detect and harder to completely shut down. Eventually, it either recognizes that obtaining its goals will be more assured without human interference, or it develops superior technology that allows it to ignore us and take our habitat, much as we do to other animals and insects.
The initial time bomb might have a slow clock, because the first swarms to take this form might not be very capable at first. So it’s hard to say how long it would take them to endanger humanity, or whether they would simply lose out to (or be coopted by) later swarms of more advanced agents coming out of the labs. But if persistent swarms are allowed to form — whether on purpose, by neglect, or through emergent online interactions — it’s only a matter of time before the clock starts ticking.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


