China (probably) hacked Taiwan
The first documented autonomous cyberattack on a government target is just a prelude
The Financial Times reports that AI systems successfully breached Taiwanese government and public sector accounts in early July.
The tool compromised at least 85 government user accounts, extracting more than 2,500 personnel records before expanding the attack to Taiwan’s nuclear safety agency and at least seven energy companies, the research showed.
I wish they went into more detail about the nuclear safety agency, even if it’s probably not as bad as it sounds. We know it’s possible to do major damage to nuclear facilities with malware; in the late 2000s, Stuxnet did exactly that (to centrifuges, not reactors, but the principle stands). Fortunately Taiwan isn’t a nuclear state, and only recently started looking into restarting its (previously shut down) nuclear power program.
The full details of the attack aren’t public, and “Chinese hackers attacked Taiwan” is an educated guess. The evidence is fairly strong, though: not many governments store their data in Traditional Chinese, and not many hackers write internal messages in Simplified Chinese.
We only know about this attack because an Israeli AI cyberdefense company, Dream, found evidence in an online archive. Importantly, that means there are probably many more such attacks we didn’t hear about. Here’s what Dream discovered:
The archive contained 1,395 files showing the hacking tool used two open-source AI agent systems, Hermes and OpenClaw, which can be downloaded and enable AI models to carry out tasks autonomously.
The researchers could not identify which AI model was used to power the agents. However, the data showed that the underlying model’s safeguards had been bypassed by presenting the hacking activity as an authorised exercise to test for system vulnerabilities.
Sound familiar? The bypass method is a standard trick: you can get many AI models to hack for you by telling them “we’re just testing these defenses.” This is the same exploit that got export controls slapped onto Anthropic’s Fable for a few weeks, and it’s extremely hard to prevent. After all, you want AI models to be willing to help you find vulnerabilities in your systems!
Anthropic crudely patched the problem by screening Fable inputs for anything vaguely cyber-related and rejecting most of them. Other AIs (including open-weight AIs) have looser standards. There’s no shortage of agents that might have powered the probably-Chinese attacks.
The article highlights the degree of persistence and sophistication demonstrated:
The most striking feature of the July attack was how the tool continuously ranked and reprioritised possible attack paths based on available evidence, Dream said.
When one attack path failed, the tool deployed another agent to scour the internet for information and devise a new approach as a human hacker would.
I am reminded of OpenAI’s accidental AI swarm, which showed similar tenacity in its autonomous breakout and subsequent attack on Hugging Face. I think OpenAI learned the wrong lessons from that incident, focusing entirely on how breathtakingly capable the models were and disregarding the failures of ethical reasoning on display. But they did get one thing right: widespread autonomous cyberattacks are coming.
Dream’s chief strategy officer argues that every government on Earth should assume they are under constant siege by AI hackers. I expect the same will soon be true for most companies, from banks to Australian gym websites. This is one price the world is now paying for failing to rein in AI developers.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



