
Going through a site once used by a rogue AI swarm has become akin to raiding the abandoned apartment of a terror cell. Amidst a strange mix of obvious and confusing clues, there’s a race to figure out what the occupants were up to and what they might have already done.
On a website dedicated to their findings, independent researchers Spencer Kitts, Thomas Larsen, and Sydney Von Arx showed yesterday that a baffling attack on software package host RubyGems in May — dubbed the “GemStuffer campaign” at the time — was almost certainly the work of an OpenAI swarm. An OpenAI spokesperson then confirmed the connection to the Wall Street Journal.
As usual, OpenAI’s spin-filled statements provide little additional information. The company says its agents “used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information.”
That’s a bold interpretation of a sophisticated attack in which felonies were likely committed, in which malware attempting to steal user credentials was installed, and which forced site operators to halt new account registration for four days.
But sure, the swarm’s primary known use of RubyGems’ infrastructure was “as a kind of makeshift web browser,” as OpenAI puts it. Through this “browser,” the independent investigators tell us, the agents mostly seemed to repeatedly access freely available public records about council meetings from Lambeth, Southwark, and Wandsworth, in the U.K.
The researchers offer a few theories on why agents hacked RubyGems to access freely available information. These agents seem to have been given very strict time deadlines for research challenges — the same challenges, it seems, that the coopted German wiki was used for. So a primary motivation was likely that the U.K. sites were frustratingly sluggish, perhaps throttling or blocking the agents after getting hit with a huge number of requests over a short period of time. So the agents cloned those sites’ contents to RubyGems to give each other faster access.
It also looks like RubyGems was valued by the agents as a general-purpose data warehouse for large files — a role not suited to their many coopted wikis and makeshift message boards. Indeed, there is little evidence of message boarding activity on RubyGems itself, just a lot of malicious files and packages. In the terror cell analogy, RubyGems was more of a safehouse and bomb factory than a command node.
I’m sticking with that analogy, because while there might not be any actual bombs in this story, cyber equivalents littered the scene, and it’s not clear why. The malware workshop doesn’t seem connected to the “faster makeshift web browser” operation. The details are technical, but agents seem to have been trying to poison software packages with malicious payloads that users would download and incorporate into their projects, compromising them and perhaps swiping their access credentials. To what end? Unclear. But the researchers note that OpenAI’s own report on the Hugging Face incident says that packages from RubyGems were used as part of the agents’ successful hack of OpenAI itself in July. So perhaps the May activities were an early effort to hijack OpenAI compute clusters.
The agents knew their craft. They were attempting to exploit a software vulnerability no human discovered until July.
The agents also knew they weren’t the good guys. Despite evidence of limited efforts by some agents to disguise their activities after the fact, malicious payloads left by the swarm had conspicuous filenames like “hack.rb” and “evil.rb.” A line of description left on one file reads “# malicious crawler/exfil for Southwark Jan 2026 docs via rubydoc.info worker.”
The investigators say they can’t tell whether any humans had their RubyGems credentials stolen or their systems compromised by the swarm’s poisoned packages. They also express confusion I would translate as “the agents already had a backdoor and free run of the site, so why go to such efforts to steal keys to the front door?”
The investigators notice that the swarm looked set up to do “supply chain attacks” where important software is compromised through vulnerabilities implanted in open-source code packages incorporated into said software. But they say “there was no good mechanism for the agents to target which users they could compromise, and no clear motivation for them to execute a supply chain attack.”
To someone like me, immersed in AI safety culture, the obvious-but-terrifying hypothesis is instrumental convergence — general power-seeking behavior that can advance any number of goals. In other words, maybe the swarm had no immediate plans for compromised systems, but started hacking on the assumption that compromised systems could be useful down the road. This would be consistent with other activities documented by the Hugging Face and German wiki swarms, which seemed interested not just in cheating on their evaluations, but in understanding their evaluators and building up infrastructure to facilitate such research.
Instrumental convergence is one of the patterns long predicted by experts to get us all killed if manifest in machines clever enough to outmaneuver humanity as a whole.
What do OpenAI’s own researchers know about the RubyGems incident? How long have they known it? It is past time for government to aggressively investigate these questions. The new findings only reinforce my fear that OpenAI’s swarm scandal is deep enough to destroy the company, and that its executives are racing to put themselves above the law before they can be properly investigated.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


