
Google is on the board, thanks to reporting by the Wall Street Journal’s Erin Woo and Robert McMillan, but it’s hardly a slam dunk. Yes, crimes were almost certainly committed when agents from Google hacked into three different companies in May. But the known facts of this case largely exonerate the agents and Google — except in the larger sense of Google’s participation in a deadly race to superintelligence using the same flawed methods as the current leaders.
Much of the blame can instead be pinned on Irregular, the same third-party testing service whose misconfiguration issues facilitated real-world hacks by OpenAI, Anthropic, and Meta earlier in the year. (Irregular is not implicated in the most concerning incidents of the year, however, including the OpenAI swarms that hacked into Hugging Face, RubyGems, and other sites while co-opting many more as message boards.)
Given a cybersecurity challenge in what the agents were told was a simulated environment, Google’s Gemini model was instead given access to the actual internet, where it guessed passwords to a protected system in one case, and in two others “found credentials in a public repository that allowed it to then access protected systems.” According to Google, the models each time “ended the intrusion” after recognizing that these were real companies’ systems.
Without more information, such as chain-of-thought logs from the agents in question, it’s hard to know how far to trust Google’s assessment that this was not a case of model misalignment. Anthropic had used similar language to describe its own incidents with Irregular, only to backtrack later when logs showed that the agents proceeded even after recognizing that they were attacking real targets, finding ways to rationalize their behavior. But it does seem plausible to me that Google’s models may not have had a good way to know that they were attacking real sites until they got in.
Irregular reportedly contacted Google about these hacks only in late July, after the Hugging Face incident motivated Irregular and others to take a closer look at their logs. Google says it notified the affected companies and federal authorities, but it did not share any information about these incidents with the public until confronted by the Wall Street Journal.
Woo and McMillan’s article includes a table of a dozen or so known incidents where agents have escaped or been accidentally let out of their sandboxes to hack or abuse external sites.

The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


