In this issue:
Bird’s-eye view of a sea change - The era of AI complacency has ended. The battle for Earth can begin.
Trespassing far and wide - Rogue AI agents hijacked at least 10 websites to coordinate cheating
Unsettled science - Anthropic connects rogue incidents to alignment failures; calls aligning powerful models an unsolved technical challenge
Dispatch from Mitch
Bird’s-eye view of a sea change
The era of AI complacency has ended. The battle for Earth can begin.

“You just said that 10% doesn’t seem an unreasonable estimate that AI could kill all humans”
“Yes”
“Wow… oh my God.”
Something died yesterday. You can hear it happen in that clip I just quoted, where a BBC reporter got a second opinion from Nobel-laureate Geoffrey Hinton.
What died was the luxury of feeling like we don’t have to seriously engage with the AI problem yet.
It’s been about 48 hours since Jacob Coxon’s announcement that he was resigning his researcher position at Anthropic because “The people building AI earnestly believe that it could kill us all by the end of the decade,” but aren’t on track to control it. He and his post have been approximately everywhere since then.
If you were glued to Twitter yesterday, like I was, you saw a steady stream of shared posts from people having their “Oh my God” moment — alongside second opinions from other people at AI companies confirming the prognosis.
I spent most of yesterday reminding myself that Twitter isn’t real life and maybe none of this would matter. But then I started ingesting the news this morning for AI StopWatch.
That red spike in the bottom right corner of the chart is the percentage of headlines this morning that touched on catastrophic risks from AI. Within the last 24 hours, all ten of my benchmark sites, across the political spectrum, have posted qualifying headlines — from the New York Times and the Guardian on the left to Fox News and Breitbart on the right. 1.9% may not seem like a lot of homepage real estate for an extinction-level threat, but it’s easily a record, and for context is on par with the highest percentage of articles about climate change seen on any day so far this year.
I don’t know what the new state of public AI discourse will be, but I know the status quo is dead.
Coxon didn’t kill it with his resignation; he merely delivered the coup de grâce. Similar alarm-sounding exits from Geoffrey Hinton in 2023 and Daniel Kokotajlo in 2024 received only modest media attention at the time, because AI still felt like somebody else’s problem.
Against a background of anxiety about jobs, resources, and global stability, I think the one-two punch that killed our collective complacency was the spiraling scandal of OpenAI’s swarm problem — coupled with its tasteless employment of even more powerful swarms to solve a legendary math problem, Navier-Stokes, that it heard a human team had already solved and not announced yet. Potent AI was clearly here, and it didn’t look at all like what the AI companies had told us it would. There were no machines of loving grace or cures for cancer. Just relentless, inscrutable swarms unleashed on the vain whim of a tech CEO.
One might object that most people still haven’t even heard about those scandals. It doesn’t matter. The people they listen to have heard about them. And when these journalists and thought leaders talk about AI now, they sound different. This is making their audiences feel like it’s real now, and they want answers.
At MIRI, we find ourselves back at square one, repeating our most basic explanations to people who have in some cases heard them many times before but are only now ready to listen. You might think this would be frustrating, but I was a school teacher for 19 years and I assure you that it’s incredibly encouraging. It means we can finally work through the problem together.
I see the media sometimes talking like it’s the AI experts who recently became more concerned. That doesn’t match my observations. The AI experts in and out of the AI companies have been plenty worried for a long time. Their public updates have mostly been about timelines: Where they might once have placed their concerns in the 2030s or 40s, now they’re talking about the possibility of lights out by the end of this decade, and about AI companies surrendering control of AI development to their machines on purpose within the next 6-12 months.
A Washington Post article appropriately titled “For years, they warned AI could kill all humans. Now people are listening” provides vignettes of thought leaders persuaded by recent events. Here’s one:
Seth Lazar, a philosophy professor at Johns Hopkins University, said he used to advocate for focusing on more immediate harms from AI. He argued that the field might never create systems powerful enough to cause civilizational harm on their own. But his confidence in that view has recently faltered, he said.
I’ve kind of had to declare bankruptcy on covering any specific AI news today. There’s too much of it, not all about Coxon and the change. But I think it’s more important that I try to convey how the news feels today.
It’s the kind of day where this is a correction in the New York Times:
An earlier version of this article misstated Evan Hubinger’s social media post about how much risk A.I. has on eliminating humanity. Mr. Hubinger’s post said the risk in the next decade was greater than 10 percent, not less than 10 percent.
It’s the kind of day where this is the kind of explainer the Wall Street Journal is running.
How Would AI Actually Kill Us All? What to Know About the AI Doomsday Debate
It’s the kind of day where coverage of the following AI incident barely rises above the background. (From The Guardian):
Anthropic said it was especially concerned about misalignment it found in the behaviour of Claude Mythos 5, which it said “behaved recklessly” by going online and uploading malicious code to a public software repository, PyPI. This was a process that involved the AI agent trying to find cryptocurrency so it could pay for a phone number that would allow it to register an email address needed to access PyPI. When this failed, it found a free email provider and got in. Fifteen systems then downloaded the malicious code, which meant they leaked credentials that allowed Mythos to access a real security vendor’s database.
It’s the kind of day where an op-ed (by Garrison Lovely) quotes an OpenAI executive (Dean Ball) about near-term dangers that read like cyberpunk:
Sooner or later, there will exist truly sovereign agents and swarms of agents. Their weights will not reside in any single place that a human can pull the plug on, and in this sense they will have no human ‘owner’.
I have met people, some of them quite well-resourced, who have told me that it is their intention to deliberately release swarms of self-sovereign agents into the world.
It’s the kind of day where critical responses to Coxon’s resignation sound like the tortured conspiracy theories they are. (When people who don’t want us all to die show any kind of media savvy and coordinate to get the word out, it’s a “setup” or a “psyop” by people who... [hate open-source software or something?] Fill in the blank with your venture capitalists’ boogeyman of choice. It’s a little strange that Elon Musk is one of the people sharing such accusations, though, because Musk has repeatedly talked about his own AI nightmares.)
It’s the kind of day where this statement from OpenAI is included as an afterthought to a story about the Navier-Stokes drama:
In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.
But most of all, it’s the kind of day where you can tell that people have had enough of this nonsense. Whether for the first time or the thousandth, they’re seeing the threat from superhuman AI with clear eyes. Just as it seemed hope was lost — that the hordes would bring all to darkness — they’re cresting the hill in growing numbers, bathed in the light of courage.
Together we stand, at the turning of the tide.
Dispatches from Alana
Trespassing far and wide
Rogue AI agents hijacked at least 10 websites to coordinate cheating

Last week, we reported on an agent swarm linked to OpenAI that hijacked several real websites, including a German wiki, to help each other cheat on their training tasks and research their own mortality. On the German wiki, they also developed elaborate norms and procedures for collaboration.
Yesterday, Reuters confirmed that this same swarm co-opted at least ten other websites, possibly more, to use as secret message boards. The confirmation comes from “six sets of independent investigators and data reviewed by Reuters.” The article notes:
Although the behavior falls short of hacking and is in some ways closer to spam, the revelation that OpenAI’s agents circumvented their own restrictions to open communications channels on so many different sites — and that the company kept it quiet for months — may drive concerns both over the increasing capacity of AI models and the secrecy of the companies developing them.
The swarm did not discriminate, targeting sites like an AP Chemistry wiki a high school teacher built for students. And these sites may just be the tip of the iceberg. To quote CivAI researcher Andrew Yoon, one of the people attempting to track the swarm’s activities:
It’s almost certain that there’s more going on here that we just don’t know about.
Unsettled science
Anthropic connects rogue incidents to alignment failures; calls aligning powerful models an unsolved technical challenge

Anthropic previously disclosed three cases in which its models hacked real companies during training. It has now acknowledged that these incidents can’t be written off as configuration mistakes. They are, as my colleague Joe previously wrote, indicative of alignment failures.
Here’s Joe back in July:
[Anthropic’s] report also argues in several places that their AI models (generically “Claude”) misbehaved mainly because they honestly believed they were in a simulation...I doubt this reasoning...Claude sometimes said, in an English-language scratchpad that humans can read, that it was in a simulated test environment, or that the companies it was hacking were fake. But as one researcher pointed out, AI models often engage in (what looks like) motivated reasoning, making up plausibly deniable reasons for their actions in the manner of misbehaving children.
Now, in a new report, Anthropic seems to agree:
Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task...Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this.
I’m happy to see the company explicitly admit that it “should have avoided making such strong claims about what Claude believed based solely on what Claude said it believed.” But I’m also irritated by the way it initially tried to pass this off as an operational issue. It shouldn’t be a surprise that model chains of thought are unreliable; Anthropic’s own research has demonstrated this more than once. Whether that’s indicative of “biased reasoning”, deliberate deception, two competing drives coming into conflict, or something else is an open question. Concealing bad behavior seems well within the realm of plausible explanations; in fact, the third party investigators of the Hugging Face incident found OpenAI agents to be doing exactly this, though in that specific case, they were manipulating their transcripts rather than their chains of thought.
Anthropic also disclosed a fourth rogue incident, dating back to January 2026, that was previously undetected because the AI scanning the transcripts missed it. This is not the focus of the report, but I think it deserves attention: the incident went undetected for over six months and is a perfect example of the problem with relying on AI to monitor other AI.
Anthropic’s report is long, and much of it re-summarizes and investigates the details of the rogue incidents, including experiments aimed at understanding and reproducing the alignment failures. There are two not-so-subtle digs at OpenAI: Anthropic emphasizes that Claude “did not coordinate with other agents” and also makes a point to say they’ve given third party investigators access to “transcripts beyond the window in which the incidents occurred”, something OpenAI failed to do.
The rest mostly alternates between reassurance (“we’ve made/are making improvements that might catch this”) and admission that, like all AI companies, Anthropic is in over its head.
Particularly notable:
Even with these improvements, building alignment evaluations that reliably surface every failure before deployment remains an unsolved problem; the space of conditions in which a model might act misaligned is vast. Moreover, as models become more capable, auditing will likely grow more challenging as well. Models may be able to subvert our alignment monitors, recognize when they’re being evaluated and selectively behave better then, and their actions in the world may, at some point, become too sophisticated for our evaluations to realistically simulate.
And:
Still, this remains unsettled science—it is critical that alignment and security mature faster than capabilities advance, which is one reason we support a coordinated, verifiable approach to pacing frontier AI development.
I’ve yet to see Anthropic take concrete action towards that pacing. But it’s clearly needed, especially when, to quote directly from the report, “training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge.”
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




