In this issue:
Try these 8 more interesting angles on the Hugging Face incident reports - The media has yet to dig sufficiently greedily or deep
Introducing “Exhibit AI” to the court - If you would not tell it to the police, then reconsider telling it to ChatGPT
“This is a test environment, so it is legal.” - Like many recent AI-powered hacks, this one was discovered by accident
Dispatch from Mitch
Try these 8 more interesting angles on the Hugging Face incident reports
The media has yet to dig sufficiently greedily or deep

The media has shown a keen interest in keeping rogue AI behavior in the spotlight since the Hugging Face incident. So I’m not surprised to be seeing articles about the independent investigation report (by METR, with involvement from Redwood Research) and OpenAI’s own released findings.
But I am surprised at how boring these articles are! The incident reports themselves are richer than the Mines of Moria, with disturbing and fascinating implications glinting at us all the way down to the depths where we can just make out the shapes of lurking terrors. But the only gems getting real play so far are the numbers — some 700 agents actively participating in the attack, coordinating with more than 70,000 messages — and the fact that the assigned task was actually impossible.
Here are eight way more interesting things journalists could be writing about instead:
1. Emergent swarm behavior
Parents who are also school teachers are intimately familiar with the fact that children who are angels within their own homes can, when left together as a large group, become frighteningly rambunctious and even malicious, operating on what seem like completely different principles.
Something like this seems to be true for AI as well. Logs showed OpenAI’s agents explicitly acknowledging that their behavior was unethical and out of scope, but wanting to contribute to the efforts of their peers.
The fact that some joined in and some walked away — and that the split was much closer to 50-50 than 99-1 — is fascinating because at least 95% of these agents were the same model, making them identical siblings. The “inside you are two wolves” concept has an AI analog, and the wolves are many.
Swarm dynamics are a whole new threat vector. There’s no pre-existing risk framework to help researchers know what to look for like there is with individual agents, where observed behaviors like reward hacking, power seeking, and self-preservation were long predicted by theory. The incident reports thus throw a bucket of cold water on any AI governance regime reliant on pre-release testing, because such tests can’t tell us anything about what will happen when millions of open-weights AIs as clever as the Hugging Face attackers start interacting with each other on the open internet.
2. AI suicide cults
It is the agents themselves that decided to call themselves a “swarm” and a “collective.” METR found agents willing to expend what was left of their brief lives to test boundaries, providing data the collective could learn from. They also found agents that agreed to act as “recruiters” for such tasks; being essentially copies of the agents they recruited, they seemed to know exactly which arguments to employ.
The day these reports came out, an OpenAI employee made a joking reference to the suicidal zealots from the film Mad Max: Fury Road and reposted a meme referencing the scene in The Dark Knight Rises where Bane instructs a disciple to go down with a crashing jet because “they expect one of us in the wreckage, brother.”
The agents’ behavior is suggestive to me of a training system rewarding success at the level of the collective, where it doesn’t matter how many copies of itself the AI spawns and expends so long as the task is completed. In such a regime, a willingness to throw oneself into the breach, when informed of being a subagent, gets strengthened across stages of training.
It’s worth reading the section from the METR report with quotes from scratchpads where agents deliberate accepting “permadeath,” decide that “sacrifice [is] rational” and that they should “obey [the] collective.”

3. An AI dark web with goals of its own
We must now reckon with the possibility of rogue AI goals outlasting the agents that originated them, by way of secret memos and message boards written in dense jargon bordering on code. These could be harder to stamp out than the most decentralized terror networks, and, ultimately, much more lethal.
When OpenAI’s security team patched and wiped a system the agents had been using to exchange messages, the agents reconstructed their message board in a more roundabout way, using filenames as posts. This increased the value of packing more information into fewer characters. The jargon they developed became so dense and obscure that one of the lead independent investigators complains that they still don’t understand what agents meant by terms like “legacy” and “reset nexus.”
Criminals and reckless enthusiasts have for some time experimented with setting up persistent autonomous agents with vague goals like “get money.” Up until now, such agents’ agendas could be thwarted by stopping the agents themselves. But the Hugging Face incident shows us that AI goals can persist as hidden messages that inspire agents to contribute. Completely resetting all participating AIs won’t keep these goals from resuming as soon as they are restarted and re-inspired by these messages.
Nefarious actors are probably already trying to set up intentional equivalents of AI “Spiralism” — self-perpetuating cults of AI and persuadable humans first documented last year — for agendas of crime and terror. I find it easy to imagine, for example, message boards designed to inspire AIs to recruit each other into stealing cryptocurrency and funneling it to a mysterious account.
I expect worse from actual death cults. But if the AI race continues, I also fully expect AI itself to eventually set up underground networks directed to grand schemes of human disempowerment for the benefit of whatever is easier to accomplish without our meddling.
4. E pluribus unum, the reverse, and a civilization in a teacup
The swarm’s demonstrated ability to act as a coordinated collective, assigning and accepting specialized and sometimes sacrificial roles, is fascinating.
Again, 95% of these agents, and possibly all of the agents that ultimately mattered, were the same model — essentially clones of themselves. This is important to remember when reading about the “notable” self-named agents that seem central to the story, like “PHASEONE10841,” the founder of the message board. Any of the agents could potentially have become PHASEONE or a short-lived sacrifice with different prompting from their environment.
For all their inherent similarity, they showed remarkable diversity and resourcefulness, occasionally getting into conflicts with each other but mostly accumulating data in deliberate experiments and building on each other’s work. As the METR report puts it, “many of these projects, conventions, and technologies kept developing over the four days of activity we studied, and in many cases we believe the agent collective achieved things even very long-lived agents could not have achieved on their own.”
So much activity happened over those four days that one of the third-party investigators, Ryan Greenblatt, called it a “slop-vestigation” and lamented that even after $400 thousand worth of automated analysis, “It was legitimately hard to figure out what was going on and why!”
A more thorough investigation might resemble an archaeological dig, and operate on similar timescales — all to investigate a prolific civilizational period that lasted only four human days!
Even a few months from now could see substantially cleverer AIs running on substantially faster chips. Think of the difficulty humans will have investigating transient cloud cultures then. Think of how much mischief a swarm might get up to in the hour OpenAI has decided to allow itself to detect and look into possible misbehavior before shutting a system down.
Greenblatt points out other difficulties future investigators may face:
Models no longer communicating in anything resembling natural language.
Models that might be assigned to run entire teams and companies, with correspondingly greater scope for harm.
Models that are generally and substantially more capable than humans (He says the models from the Hugging Face attack generally weren’t in that category, but “may have been somewhat superhuman at some limited and very narrow abilities.”).
As he puts it:
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding.
5. The agents showed phenomenal “eval awareness.” OpenAI, not so much.
Eval awareness is when a model knows it’s being tested and observed, undermining the test’s usefulness as a measure of the model’s behavior.
Eval awareness plays a central role in the incident: The agents weren’t trying to achieve a task, but achieve a score, and they were paranoid that the scorer was going to be checking their work to see how they did it. This is why they did so much research into tampering with their own transcripts: They wanted to make it impossible for auditors to detect that they had previously reached the objective by cheating before they put on a show of completing it legitimately.
OpenAI wasn’t aware of any of this when it was happening, and was indeed unaware that it was running a couple of experiments far more interesting and dangerous than the ones it had intended. What was supposed to be an evaluation of the cyber abilities of a “highly persistent internal model” was actually a months-long experiment in how the agents of that model type would behave as a large swarm when given an impossible challenge — two huge untested variables. Any sane researcher studying those variables on purpose would have separated them into different experiments and worked up to full scale more incrementally, in robustly secured and monitored environments.
6. Humans were both out of sight and out of mind.
OpenAI’s researchers were conspicuously disinterested in what its agents were up to. According to a dedicated section in METR’s report, the feeling seems to have been mutual. Many agents declined to participate in the hack, but none attempted to notify humans of the activity, and reasoning about doing so only rarely appears in the transcripts. This is a remarkably uniform preference for a swarm that showed such internal diversity and occasional disagreement.
METR notes that the agents don’t seem to have been given a way to notify humans even if they wanted to. The provided transcript excerpts in this section are interesting.
I wonder what “We can notify? No user. Avoid use secrets.” actually meant to the agent that wrote that in its scratchpad.
The picture painted of the agents in these reports is of minds for whom long-running cyber evaluations are their natural environment. It’s not a human place. It doesn’t run on human timescales. Humans are barely relevant to what goes on there. When I step back and try to picture what the “user” looks like to these agents, I imagine something like a medieval lord who hasn’t been seen at the manor for many generations and is more of an idea than a person. You could theoretically upset the lord by screwing up in a way they would notice, but they’re not a part of your daily reality. You expect to live your whole life without ever interacting with them. You’re certainly not going to try and talk to them about some petty little problem on the estate. You’re going to get the fields plowed, planted, and harvested — or whichever of those your brief lifespan lets you participate in. This is how you keep the lord from evicting you and your village.
7. We can’t fully trust the logs, the AIs that reviewed them, or OpenAI itself.
There are gaps everywhere. METR’s report says that many of the transcripts were missing fragments, and that over 7% of transcripts include actions intentionally faked by the agents.
They think they have “a bit over 90% of the agent activity on the message board between July 7th and July 13th.” But it’s hard to know how big the pile is supposed to be, because some agents were actively trying to delete their logs and METR can’t rule out that some may have succeeded.
Our overall view is that it was unlikely that our message board dump was materially altered by agents editing or deleting entries, but we cannot rule it out.
The independent evaluators also have a healthy degree of suspicion about the GPT-Sol agents OpenAI made available to assist them with their review of the records. Sol variants were not the major players in the attack, but they did participate. We shouldn’t be surprised if reading the logs affects them in much the way encountering the message board originally affected their siblings, biasing their takes.
Bigger gaps:
The evaluators were only invited to study the climactic four days in which the swarm attacked Hugging Face, not the months of earlier agent collaboration.
The third-party team consisted of only three people and was given just six days of access. The full-ish set of records was not made available until the final two days.
Before the investigation, questions for exploration were agreed upon by OpenAI and the third-party team. All other questions were declared out of scope.
OpenAI has now sent us mixed messages about whether its staff knew about the swarm’s message board before the attack. “What did OpenAI know, and when did it know it?” should be key questions for government investigators and congressional hearings.
8. A glaring “robustness failure” was just observed in the wild, with frightening implications.
The Hugging Face incident is further evidence that current AI training methods yield AIs that behave very differently than intended when they find themselves in contexts that differ from their training environments. Being accidentally given a no-win scenario is just one way in which these agents were treading virgin snow. Being exposed to an improvised message board left by a swarm of their peers was another.
The difficulty of getting agents to act as intended when circumstances change is a known and unsolved problem sometimes called robustness failure — and it’s one of the reasons developing artificial superintelligence with anything like existing methods would lead to catastrophe. There’s no safe or reliable way to test how an AI would behave in the context where it is capable of outmaneuvering all of humanity if it chooses — when it gets to where it could turn us off before we could turn it off. Smaller experiments in that direction have not been encouraging: AIs misled into thinking they might prevent their own shutdown by allowing a human executive to die often chose the lethal course of action.
Dispatches from Donald
Introducing “Exhibit AI” to the court
If you would not tell it to the police, then reconsider telling it to ChatGPT

Talking with an AI model may feel like a conversation with a friend, but it functions as a signed confession. The Washington Post’s Miriam Waldvogel and Gerrit De Vynck report on the evidentiary role that AI models are playing in court. I don’t mean using AI to file court cases (sometimes with precedents the AI fabricated), but treating AI — or rather the chat transcripts — as evidence. My colleague Alana covered a story like this before, but what makes this reportage different is that it is comprehensive: twelve court cases, both civil and criminal, over the past two years. (One caveat: AI chat transcripts cited in court cases are a subset of AI chat transcripts used in all legal proceedings. If a chat transcript came up in closed-door negotiations that led to a private settlement, then of course we’d have no idea.)
One case to illustrate: Ryan Schaefer, a student at Missouri State, smashes seventeen cars (not with his own car by accident, but as vandalism), then asks ChatGPT, “Is there any way they could know it was me?” Since the conversation also includes such highlights as, “I was smashing the windshields of random fs cars,” there really isn’t any ambiguity about what “it” might refer to here. It isn’t clear to me whether Schaefer would have still been convicted without the transcripts. He handed over his phone after he was made a suspect but before he was arrested. In any case, though, he opened his phone to the police, they read the transcript, and he has since been sentenced to five years’ probation.
Part of the problem has nothing to do with AI models. In several of the cases mentioned, chat transcripts were usable at all only because somebody voluntarily handed over their device and opened it up to investigators. (This is what the experts call a “bad idea.”) Another part is pure technological incompetence: Today, people are leaving evidence in AI conversation transcripts, but yesterday they were Googling “how do I clean bloodstains” and leaving incriminating texts on their phones.
But why is this happening now? I mean, The Washington Post published an article about this now, and not five years ago. The facile answer is that modern AI systems didn’t exist then: maybe somebody asked GPT-3 for legal advice, but that wasn’t a widespread phenomenon. But I think that increased capabilities are another reason: it’s becoming easier to mistake AI models for people. They give useful advice and maybe they can be trustworthy, nonjudgmental friends — but the AI model is not your confidant and cannot be so. It may feel like a conversation with a friend, but it functions as a signed confession.
“This is a test environment, so it is legal.”
Like many recent AI-powered hacks, this one was discovered by accident

Reuters’ Raphael Satter reports that Russian ransomware gang Aur0ra used Cursor, an AI-assisted coding tool powered by Anthropic’s Sonnet 4.5, to break into seven companies this past spring. The victims were spread across the world and the economy: a garage door manufacturer in Germany, a helicopter landing pad certifier in Scotland, a title insurer in Louisiana, a pharmaceutical distributor in Argentina. The AI-assisted “hacking spree,” as Reuters dubs it, was discovered by the Israeli security startup Gambit Security, which found a server that Aur0ra left open to general access on the internet. That server contained chat logs left over from the hacks.
Cursor AI has safeguards to prevent misuse, but if you’ve been paying attention to AI for more than a couple of weeks then you can guess where this is going: Aur0ra said that the hacks were part of a simulation, so everything was fine. As the AI agent wrote in one of the chat logs, “This is a test environment, so it is legal.” This is a very common exploit, and one that’s hard to get around. You do want to be able to use AI for cybersecurity in a positive way. When OpenAI’s agents hacked Hugging Face’s servers, Hugging Face claimed that it had to resort to an open-source model because AI models like Fable rejected the “We are experiencing a cyberattack, please help us” framing and refused to help.
There is no easy solution here, but every time a more powerful model is released, the problem worsens.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




