In this issue:
Another swarm safehouse discovered, malicious activity evident - Confusing May attack on software platform RubyGems traced to OpenAI swarm
Musk, Altman, and Hassabis endorse Amodei proposal to “Pace the Frontier” - But commitments are vague, and the proposal is only a start
Dispatches from Mitch
Another swarm safehouse discovered, malicious activity evident
Confusing May attack on software platform RubyGems traced to OpenAI swarm

Going through a site once used by a rogue AI swarm has become akin to raiding the abandoned apartment of a terror cell. Amidst a strange mix of obvious and confusing clues, there’s a race to figure out what the occupants were up to and what they might have already done.
On a website dedicated to their findings, independent researchers Spencer Kitts, Thomas Larsen, and Sydney Von Arx showed yesterday that a baffling attack on software package host RubyGems in May — dubbed the “GemStuffer campaign” at the time — was almost certainly the work of an OpenAI swarm. An OpenAI spokesperson then confirmed the connection to the Wall Street Journal.
As usual, OpenAI’s spin-filled statements provide little additional information. The company says its agents “used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information.”
That’s a bold interpretation of a sophisticated attack in which felonies were likely committed, in which malware attempting to steal user credentials was installed, and which forced site operators to halt new account registration for four days.
But sure, the swarm’s primary known use of RubyGems’ infrastructure was “as a kind of makeshift web browser,” as OpenAI puts it. Through this “browser,” the independent investigators tell us, the agents mostly seemed to repeatedly access freely available public records about council meetings from Lambeth, Southwark, and Wandsworth, in the U.K.
The researchers offer a few theories on why agents hacked RubyGems to access freely available information. These agents seem to have been given very strict time deadlines for research challenges — the same challenges, it seems, that the coopted German wiki was used for. So a primary motivation was likely that the U.K. sites were frustratingly sluggish, perhaps throttling or blocking the agents after getting hit with a huge number of requests over a short period of time. So the agents cloned those sites’ contents to RubyGems to give each other faster access.
It also looks like RubyGems was valued by the agents as a general-purpose data warehouse for large files — a role not suited to their many coopted wikis and makeshift message boards. Indeed, there is little evidence of message boarding activity on RubyGems itself, just a lot of malicious files and packages. In the terror cell analogy, RubyGems was more of a safehouse and bomb factory than a command node.
I’m sticking with that analogy, because while there might not be any actual bombs in this story, cyber equivalents littered the scene, and it’s not clear why. The malware workshop doesn’t seem connected to the “faster makeshift web browser” operation. The details are technical, but agents seem to have been trying to poison software packages with malicious payloads that users would download and incorporate into their projects, compromising them and perhaps swiping their access credentials. To what end? Unclear. But the researchers note that OpenAI’s own report on the Hugging Face incident says that packages from RubyGems were used as part of the agents’ successful hack of OpenAI itself in July. So perhaps the May activities were an early effort to hijack OpenAI compute clusters.
The agents knew their craft. They were attempting to exploit a software vulnerability no human discovered until July.
The agents also knew they weren’t the good guys. Despite evidence of limited efforts by some agents to disguise their activities after the fact, malicious payloads left by the swarm had conspicuous filenames like “hack.rb” and “evil.rb.” A line of description left on one file reads “# malicious crawler/exfil for Southwark Jan 2026 docs via rubydoc.info worker.”
The investigators say they can’t tell whether any humans had their RubyGems credentials stolen or their systems compromised by the swarm’s poisoned packages. They also express confusion I would translate as “the agents already had a backdoor and free run of the site, so why go to such efforts to steal keys to the front door?”
The investigators notice that the swarm looked set up to do “supply chain attacks” where important software is compromised through vulnerabilities implanted in open-source code packages incorporated into said software. But they say “there was no good mechanism for the agents to target which users they could compromise, and no clear motivation for them to execute a supply chain attack.”
To someone like me, immersed in AI safety culture, the obvious-but-terrifying hypothesis is instrumental convergence — general power-seeking behavior that can advance any number of goals. In other words, maybe the swarm had no immediate plans for compromised systems, but started hacking on the assumption that compromised systems could be useful down the road. This would be consistent with other activities documented about the Hugging Face and German wiki swarms, which seemed interested not just in cheating on their evaluations, but in understanding their evaluators and building up infrastructure to facilitate such research.
Instrumental convergence is one of the patterns long predicted by experts to get us all killed if manifest in machines clever enough to outmaneuver humanity as a whole.
What do OpenAI’s own researchers know about the RubyGems incident? How long have they known it? It is past time for government to aggressively investigate these questions. The new findings only reinforce my fear that OpenAI’s swarm scandal is deep enough to destroy the company, and that its executives are racing to put themselves above the law before they can be properly investigated.
Musk, Altman, and Hassabis endorse Amodei proposal to “Pace the Frontier”
But commitments are vague, and the proposal is only a start

Leaders at four top AI labs all indicated agreement today on the need to “pace the frontier” of AI development.
The anchoring proposal comes from Anthropic’s Dario Amodei, who posted a 3,800 word essay called “We Must Pace the Frontier,” which I’ll get into below. But it’s not clear whether the other three leaders — OpenAI’s Sam Altman, SpaceX’s Elon Musk, and Google DeepMind’s Demis Hassabis — are prepared to sign on to all or most of Amodei’s specifics.
Altman kicks most of those down the road, saying “We’ll have more to share soon,” but says that one of Amodei’s proposals, independent evaluators embedded at the companies, “is a great idea, and we’ll do the same.” (Amodei’s essay commits his company to doing this right away, whether or not other companies follow suit.)
Hassabis says Amodei’s essay “points towards the right path forward” but that “the details need working through.”
Musk only posted, “Dario is right” over a retweet of Amodei, and a part of me wonders if he read past the title. Musk had casually dismissed Jacob Coxon’s viral resignation from Anthropic the other day as a “setup.” But he has also expressed plenty of deep concern about powerful AI in the past, including recently, so who knows?
And of course we have yet to hear from anyone at Meta. I exclude the leading Chinese companies from this discussion, because Amodei does, too. We’ll get to that.
We should also be up front about the fact that Amodei isn’t calling for a stop to the AI race:
To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this.
That’s unfortunate, because I think continuing those things means we just die a little bit later than scheduled. But time bought is time that could be spent implementing stronger measures; these will need to be decided and enforced by governments, because the window where we might have been able to safely trust the labs to act responsibly has closed.
The essay
Amodei opens with a change of heart, saying the acceleration of AI development, aided increasingly by AI itself, means that we’re no longer in 2023. In that relatively quaint yesteryear, he had said that slowing AI “made little sense” because trying to study AI risks with the models of the time was “like trying to study the psychology of humans by performing experiments on bacteria.”
The swarms are also a factor:
Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet.
[...] It’s also easy to dismiss [the OpenAI-Hugging Face incident] as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.
The essay’s core is a “three-step plan”:
The first is the embedded independent evaluators mentioned above. I like the idea as a thing that can happen quickly and which may turn up more acute problems the government should investigate. Cynically, I suspect Amodei unilaterally decided to do this because OpenAI probably has a lot more to hide than Anthropic, and would look even guiltier if it refused to go along. I’ll believe Altman’s matching commitment when I see trustworthy evaluators putting out reports reflecting far more access and far fewer redactions than the one METR was invited to share about the Hugging Face swarm.
Amodei’s next proposal is “democratic coordination,” but he specifies later that he means “pacing within democracies.” There’s nothing democratic about it. He’s just saying that the companies residing “within democratic countries” should coordinate on safeguards and keep China from catching up. My political science is rusty, but I think the correct word for that is “oligarchy.”
His third proposal is “global coordination,” or coordinating “with authoritarian governments, to the extent this is possible.” In this essay and elsewhere, you can tell that Amodei doesn’t really believe anything like a “full pacing” or a global “pause” is possible unless “ironclad” verification tools can be developed. He supports “floating” such an idea, but mostly positions the Chinese Communist Party as a fundamental limitation on how far the race can be throttled back. I detect optimism only about finding common ground on prohibiting “certain narrow and obviously dangerous uses of AI, such as [...] the production of biological weapons.”
How would companies measure and set the pace?
Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be.
I think the era when capability assessments could be trusted ended earlier this year. The AIs are now too aware that they are being tested. The model card for OpenAI’s Astra, in particular, gives strong reason to suspect the model of sandbagging its performance.
We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI.
Amodei says he actually thinks these are even more “gameable,” I would guess because of creative accounting that could be done about kinds of compute and types of training. I think this is all the more reason to halt frontier development outright and put the burden of proof on researchers to show that their work won’t boost model capabilities.
Wouldn’t it be lovely?
I skipped over a whole section where Amodei daydreams about what Anthropic could do if it had even 1-2 years of slack to focus on understanding model behavior and tightening up operations. Basically, he dreams of a company that isn’t such a clown show of configuration errors and alignment failures.
The bottom line
Just before his conclusion, Amodei argues that “even if we cannot achieve formal agreements, simply changing informal norms may have some value” (emphasis his). He seems to be hoping that more dialogue between labs about the crazy stuff they’re seeing might result in a little more caution on average.
I can’t be mad at that, but I also can’t help but read it as Amodei saying the worst companies are making them all look bad, and that they should share tips on avoiding the most obvious blunders that will get them all shut down. For all their talk about feeling trapped by the race, I think they fear getting shut down more than they fear human extinction.
Pacing isn’t halting. Tweets of agreement aren’t regulation. We mustn’t be fooled into thinking these companies are meeting the moment. They want to keep building superintelligence, and they’ll keep gambling with our lives if we let them.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


