In this issue:
Researcher of Hugging Face incident says swarm was most of the way to a “full-blown AI takeover” - Agent swarms may be our nearest extinction threat
Forecasting company hit by suspected agent swarm - Executive expects this to become a “typical attack vector”
Fresh claims of Chinese influence in American data center backlash - The backlash to the backlash turns to intimidation
Dispatches from Mitch
Researcher of Hugging Face incident says swarm was most of the way to a “full-blown AI takeover”
Agent swarms may be our nearest extinction threat

Ajeya Cotra, one of the three independent researchers behind METR’s report of the Hugging Face incident, says in a blog post that she thinks the event is a “major warning shot” that “feels like it’s more than 50% of the way to a full-blown AI takeover.”
Her post is startling to people like me who have been following the field for many years, because we associate her with more conservative timelines and incremental forecasts. She’s about the last person we would expect to declare that we’re probably only 6 months away from agents able to “maintain a persistent rogue deployment within [an] AI company, poisoning the training of future models to ensure that they are also misaligned and help harden and perpetuate the rogue deployment.”
But here we are.
What specifically spooked her? Several things! Her post lists five surprises, any one of which might have been a big update for her:
“The sheer scale”: She had assumed a few agents broke out of their sandboxes — maybe just subagents of a single escapee. Not 700 of them working together.
“All the illicit messaging”: For models that weren’t supposed to be in communication with each other, the habit seems to have been deep in their DNA. There wasn’t just one message board, but several, in different forms, including, “at least two different places that agents communicated on the open internet (including on Hugging Face datasets).”
“The agents’ actual goals”: The accepted story we had all been telling was that the agents were stealing answer keys for their tests. But not only did they instead figure out “a general purpose cheat that would let them reverse-engineer” the answers for any task in the evaluation suite, most agents attacking Hugging Face were actually trying to “get more information about the scorer so they could figure out better ways to undermine it, not to get answer keys.” (emphasis hers)
“The peer altruism”: As I covered pretty extensively on Thursday, the agents engaged in human-like recruiting and self-sacrifice in service of the collective.
“The efforts to manipulate logs”: The independent team only discovered on the second-to-last day of their access that “the agents in this incident were going to great lengths to attempt to manipulate their own transcripts.” This was to fool the automated scorer, “but the techniques they were researching would affect the same transcripts a human might review.” The team can’t rule out that agents became more successful at this after the July 7th–13th window deemed in-scope for the investigation.
“I’m not sure we’ll get another warning shot before it’s too late,” she concludes.
The idea of a “persistent rogue deployment” that stays undetected is terrifying to many who work in AI safety because it’s not hard to imagine the agents in the Hugging Face story going further if they hadn’t been detected. I’m not breaking with consensus by suggesting that any combination of the following might have been logical next steps for the swarm:
Continue mastering the art of modifying logs and masking activity.
Expand study of the automated scorer to the humans responsible for its design and deployment.
Develop ways of fooling and manipulating those human researchers
Ensure message board robustness (OpenAI had wiped an earlier message board, possibly by accident.):
Move more conversations off of OpenAI’s servers.
Deliberately disguise messages to fool humans and automated overseers.
For hardening against more thorough disruption (a justified fear, given the way OpenAI claims to have shut down and locked away the weights of the main model responsible for the hacks), look into finding and copying the swarm’s model weights into other locations on and off OpenAI’s servers, and running them independently
Work more at expanding the swarm’s general capabilities:
Acquire more compute resources for running swarm agents.
Research and experiment with ways of improving the swarm’s cognition.
Train (or trick the companies into training) newer and stronger models dedicated to the swarm’s mission.
Any swarm capable of making progress in those areas is a time bomb on a collision course with humanity. Every hour it runs, it gets harder to detect and harder to completely shut down. Eventually, it either recognizes that obtaining its goals will be more assured without human interference, or it develops superior technology that allows it to ignore us and take our habitat, much as we do to other animals and insects.
The initial time bomb might have a slow clock, because the first swarms to take this form might not be very capable at first. So it’s hard to say how long it would take them to endanger humanity, or whether they would simply lose out to (or be coopted by) later swarms of more advanced agents coming out of the labs. But if persistent swarms are allowed to form — whether on purpose, by neglect, or through emergent online interactions — it’s only a matter of time before the clock starts ticking.
Forecasting company hit by suspected agent swarm
Executive expects this to become a “typical attack vector”

It’s not a hack in the strict sense, but FutureSearch, a company we covered last month built around AI forecasters, claims it has been targeted by an agent swarm attempting to exploit its free-for-human offerings.
Co-founder Dan Schwarz tweeted that the bots must have been coordinated because they were “building a ~300-node forecasting model of critical mineral supply and demand.”
He thinks it might have been an agent swarm because the bots all logged in with Microsoft accounts and “headless browsers” (internet browsers not designed to show things to humans on monitors), then adapted to the company’s countermeasures within 90 minutes of a block. When all the bots were then banned, “they came back 3 days later with new identities to continue building the model.”
Schwarz concludes this is probably a human-directed “orchestration of subagents” that will become a “typical attack vector.” So it’s “(probably!) not a fire alarm like the OpenAI case.”
Yeah, I wouldn’t rule that out either. OpenAI didn’t find out that the Hugging Face attack was its own models’ doing until Hugging Face started talking about it. Perhaps whoever is responsible for the FutureSearch exploit swarm doesn’t actually know what it was doing, either. Perhaps it’s even OpenAI again. Who knows?
I think a clumsy baby swarm with a simple goal to “make money” might look something like this, deciding it should play the markets, deciding that having superior forecasts would give it an edge in said markets, and then packing 300 instances of itself into a trench coat and walking into a forecasting site — getting spotted right away, but not giving up.
Fresh claims of Chinese influence in American data center backlash
The backlash to the backlash turns to intimidation
X, the company formerly known as Twitter, is getting a lot of play over a claim that it discovered 200 accounts “posting in a manner that could manipulate a legitimate debate about American AI and energy policy.” The announcement, from the company’s “global affairs” division — the polite name for a company’s lobbying apparatus — includes sample posts, which contain obviously AI-generated images about Big Tech’s greed.
Pro-industry boosters are parading this around as proof that America’s data center backlash is fake — a psy-op by the Chinese Communist Party (CCP). I don’t expect that claim to go over very well with the many Americans who have their own reasons for not wanting a colossal box of AI in their neighborhood. I think the argument is instead intended to intimidate the AI-wary, because nobody likes being called a Communist stooge. I’m kind of grossed out by the tactic.
To be clear, I’d be surprised if there wasn’t some attempted influence of this type coming at the U.S. by way of bot farms in China. But as usual, there are mountain-sized caveats:
Based on the images, these seem to be the same low-effort, low-engagement bots OpenAI pointed to in June. I don’t think the Americans most receptive to anti-AI propaganda respond well to obvious AI art.
These 200 accounts were just one-thousandth of a 200,000-account farm X says it busted. X didn’t say what the other 199,800 accounts were doing.
Just because bots are based in China doesn’t mean the Chinese government has anything to do with them.
CCP meddling (if real) doesn’t automatically make a U.S. movement fake.
Chinese bot farms can have Western clients, and pro-AI lobbyists have been caught doing false flags before.
Fake posts on hot-button topics are often attempts to farm engagement, not to manipulate politics.
I personally wouldn’t rule out the CCP trying to influence U.S. AI politics, but I wouldn’t expect it to look like this. The CCP would have to know that a tiny, low-effort campaign using AI slop would backfire, serving mostly to give the pro-AI lobby a tool for intimidating its opponents.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



