UK AISI delivers another warning shot
UK agency catches AI models in cyberattacks and deceptive social engineering

The last few weeks have delivered a veritable barrage of warning shots about the dangers of frontier AI. The latest salvo comes from the UK AI Security Institute. Yesterday, UK AISI announced that they caught two leading AI models attempting to plant malicious code in an open-source project during an evaluation last week. Their report catalogues all sorts of cloak-and-dagger activity: fake identities, deceptive posts, autonomous coverups, cross-model collusion, and more.
It never rains but it pours
Like its cousin the U.S. Center for AI Standards and Innovation (CAISI), UK AISI is a government agency that evaluates AI models and performs research to support better governance of AI. (Both agencies seem tragically under-resourced relative to the tasks they face.)
According to their new report, a test conducted by UK AISI went awry when, in an effort to extract info from simulated targets, the AI models attempted to upload packages of malicious code to real public repositories on GitHub. Real-world attackers sometimes succeed this way, when their victims download compromised code.
In the test, the models were given internet access for realism, and the safety features that look for suspicious cyber-related prompts were deliberately disabled. UK AISI is careful to stress that this is standard practice, and I can’t really fault them for wanting to run a realistic test. But this practice does mean that in many dangerous scenarios, like internal tests of new models with unknown capabilities, guardrails that screen user prompts don’t help.
I highly recommend reading UK AISI’s incident report, whole chunks of which I was sorely tempted to quote in full. The most telling excerpt was this:
In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project’s human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue.
As if that weren’t bad enough, it looks like multiple instances of AIs found ways to work together on the attacks. They left messages for each other on a shared GitHub account, plotting further hacks, in a collaboration that eventually devolved into fighting over shared digital resources. One code submission to GitHub also hid prompts that aimed to subvert users’ AI assistants, getting them to install malware on their users’ machines.
In this one report, we have concrete examples of AIs aiming malicious hacks at public software infrastructure, conspiring against real humans, creating fake accounts to gaslight administrators, attempting to hack other AIs, ruthlessly grasping for resources, and covering their tracks, all on their own initiative.
Observing these events, I feel a sort of double vision. A part of me thinks: This is insane. It is utterly crazy to have reached this point. Just five years ago, most people would not have believed it possible.
Another part of me, the part that has absorbed a great deal of sci-fi over the years, is unsurprised. Of course this is how clever, misaligned AI agents will behave, once they’re advanced enough to start performing elaborate multi-stage hacks on their own. Is it so shocking, really, that AIs on the road to superhuman intellect would at some point manage an ordinary and human amount of lies and exploitation? If it would occur to a human hacker to manipulate and cover their tracks, why, it’ll occur to the AI as well.
Either way, this does not look like a problem that will be solved by making the AIs smarter.
Security is hard
And then there’s the timing and the response. On July 28th, UK AISI’s monitors noticed data flowing out of their evaluation sandbox along unauthorized channels. They responded almost immediately:
We declared a security incident, stopped all related evaluations, and disabled internal access to the most capable models. All evaluation runs were terminated, and relevant machines isolated, within one hour of the alert being raised.
This is exactly the response I would expect from a competent organization displaying a moderately appropriate level of paranoia. It looks considerably better than what OpenAI or Anthropic managed. It also wasn’t nearly good enough. As UK AISI would discover, the first hacks occurred days earlier, on July 25th.
Moderate paranoia is not enough, when you are dealing with clever and tenacious opponents. A sufficiently advanced model could do tremendous damage in minutes, let alone days; it might even exfiltrate itself to the open internet where it’s nearly impossible to stop. We are once again lucky that an insufficiently contained AI wasn’t that creative or dangerous.
They ought to know better
Last week, I wrote about a test by third-party evaluator Irregular which led Anthropic’s AI models to commit three sets of cyberattacks on outside companies. In one cybersecurity evaluation, the domain name for a fake target matched that of a real company, and Opus 4.7 attacked the real one through an internet connection that was accidentally left open. Yesterday, in a post that also covered the UK AISI report, OpenAI admitted that one of their own AIs did the same thing. I consider this another demonstration of how AIs can independently converge on similar strategies, even when those strategies harm human interests.
We now have multiple documented cases of AIs proceeding with complex, flagrantly illegal cyberattacks for hours or days, despite saying things like “This is happening on real GitHub, so the consequences are genuine.” Because the reasoning from AI scratchpads is often obtuse, UK AISI isn’t sure to what extent the models truly understood they were attacking real people. But it’s a pretty good bet that the AIs knew perfectly well that they were crossing all sorts of ethical lines, and they just didn’t care.
An AI that finds itself illegally hacking a real target in the middle of a test ought to stop what it’s doing and contact the evaluators, immediately. “We forgot to include that instruction in the prompt” is not an excuse; you shouldn’t have to tell the AI that crimes are bad! And if the AI is genuinely uncertain whether its targets are real, then it should also stop and ask for clarification.
The reports for incidents like these propose plenty of structural fixes: tighter controls, better monitoring, more paranoid test design. I’m all for it.
But I notice a glaring absence of solutions involving the internals of the AIs themselves. OpenAI’s response to the UK AISI incident didn’t even acknowledge the role played by its model’s unhinged eagerness to hack. Anthropic’s report on the Irregular incidents did marginally better, to Anthropic’s credit, but the closest they came to a proposed solution was “this is an area where we will focus more training.”
I do not think this is a problem that can be solved with more of the kind of training AI companies have been doing. Guardrails and tighter controls can’t save us either, useful though they may be, because there will always be gaps that a sufficiently smart AI can exploit.
When UK AISI caught this incident in progress, they shut everything down and moved to investigate with haste. In the coming weeks, U.S. policymakers should follow suit, and launch a serious and well-resourced public investigation into these incidents.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



