In this issue:
Anthropic insiders’ extinction warning goes megaviral - Warnings gain traction after latest resignation
Losing control on purpose - Recursive self-improvement is handing the future to AIs
Anthropic denied testing access to the UK AISI - The third-party org was unable to evaluate Anthropic’s latest AI model
More pointed accusations of model-copying - China may be copying US companies’ recklessness as well as their AIs
Dispatch from Robert
Anthropic insiders’ extinction warning goes megaviral
Warnings gain traction after latest resignation
Last night, AI researcher Jacob Coxon announced on X that he is resigning from Anthropic. One of the reasons Coxon gave for his resignation was that neither Anthropic nor OpenAI are acting responsibly. Instead, he says, they are gambling with all of our lives by racing toward artificial superintelligence.
In his thread, Coxon writes many things with which I wholeheartedly agree, and they should sound quite familiar to regular readers of AI StopWatch:
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
His arguments — while very good — are not new. They are essentially exactly what the AI safety community has been repeating for years, some for decades now. Jacob Coxon is also not the first frontier lab employee to resign in order to warn of the significant risks — Nobel Prize winner Geoffrey Hinton (formerly of Google), Daniel Kokotajlo (formerly of OpenAI) and Richard Ngo (formerly of OpenAI and Google DeepMind) have done the same. Over 1,300 employees of frontier AI companies have signed the Pacing the Frontier statement. In the statement’s personal comment section, many researchers have been similarly forthright.
This time, though, Coxon’s message seems to have come at exactly the right moment and is currently gaining significant momentum. As I write this, his post on X has over 100 million views and more than 500,000 likes. I can’t recall a warning about the extinction risk posed by superintelligence ever receiving more attention and support on X.
Meanwhile, several U.S. senators, representatives, and even some British members of Parliament have taken Coxon’s post as an opportunity to join the chorus of warnings and call for more regulation. The governor of Illinois, JB Pritzker, writes:
It’s time to sound the alarm—louder—on reining in artificial intelligence. It’s becoming increasingly clear that AI poses a threat to humanity, so I’m calling for immediate action from the industry and Washington.
Even flagship media outlets — not only in the English-speaking world but around the globe — have picked up on the warnings and are reporting with unusual seriousness on the danger we face.
I’m still a bit surprised by how strong the reaction suddenly is, but I’ll take it. Gladly!
Of course, the reactions aren’t unanimously positive. In addition to some of the usual bad takes, there were also, bizarrely enough, a few voices on X who initially doubted that Jacob Coxon even exists or that he accurately reflects the views inside the frontier labs.
Evan Hubinger, the current Alignment Lead at Anthropic (and former MIRI researcher), came to his defense:
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think the probability is >10% within the next decade. I believe Anthropic is doing its best, but we don’t yet have a plan to solve alignment for superintelligence and aren’t clearly on track to do so.
So we shouldn’t assume that the increased attention will automatically make it easier to convince people. But we should do what we can to build on this momentum.
If you’ve been waiting for an opportune time to call or write to your senator, representative, or local member of parliament, to demand that they take action against the danger posed by uncontrolled superintelligence, this is it.
Dispatches from Joe
Losing control on purpose
Recursive self-improvement is handing the future to AIs
AI companies plan to hand the future to machines, writes Kelsey Piper in The Argument, and too many people fail to appreciate this fact.
These companies explicitly aim for recursive self-improvement (RSI), a process in which AI models design increasingly powerful successors. AI companies say this, loudly and repeatedly, in public and in private. They have already automated most of their coding work; they are eager to automate more.
Earlier this summer, OpenAI announced:
Over the past six months, the share of research compute devoted to internal coding inference grew 100-fold, while internal agentic token usage increased approximately 22-fold.
and a few days ago, it said:
We are making strong progress toward creating an automated AI researcher by March of 2028.
Anthropic likewise says it is “delegating a growing share of AI development to AI systems themselves.”
As Piper puts it, “AI companies are not, primarily, trying to automate your job — they are trying to automate their own jobs.”
What does this look like in practice? An ever-growing mass of poorly understood alien minds, far too many for any human institution to track, constantly self-enhancing in increasingly incomprehensible ways.
Just a couple weeks ago, OpenAI tasked more than ten thousand AI agents with a hard math problem. That is a fraction of a fraction of the total compute dedicated to AI research, and yet that single group contained more AIs than OpenAI has employees. They cannot possibly be looking hard enough at what those AIs are doing.
The AI companies seeking RSI are forced to rely on yet more AIs for monitoring — either the very same AIs that are being watched for misbehavior, or older and dumber models. All are unreliable, and any can lie or mislead.
The absolute deluge of AI activity is already a problem for investigators, too. When a mere 700 agents attacked Hugging Face, the third-party investigation had to lean on unreliable AI models to even begin to understand what happened.
Piper explains:
There simply isn’t enough human attention available to monitor the number of AIs the companies hope to put to work on independent AI research. As a result, we’ll be reliant on the AIs to understand what is happening inside fully automated data centers where more and more powerful AIs are developed.
This isn’t some pessimistic projection of what might go wrong — it’s actually the plan. Not the plan for the distant future; it’s the plan for next spring.
When policymakers and think tanks propose AI governance principles, they almost universally advocate for practices that keep humans in control of AI. It’s in the Asilomar AI Principles, the EU AI Act, and a UN framework on autonomous weapons. Even Chinese President Xi Jinping has urged “measures to forestall loss of control.”
Many of these proposals do not seem to grapple with the true problem, which is that AI companies do not have a workable plan which permits them to retain control. Their explicit and publicly declared approach, which they will evidently not abandon unless forced by law, is to ask some AIs to do their jobs for them, ask untrustworthy AIs to read everything and write reports, and hope their employees can keep up.
This plan is exactly as insane as it sounds, and it cannot be allowed to proceed.
Anthropic denied testing access to the UK AISI
The third-party org was unable to evaluate Anthropic’s latest AI model
For unclear reasons, Anthropic broke with precedent and declined to let the UK AI Security Institute (UKAISI) evaluate its Mythos 5.1 model before release. The Financial Times covers the omission, but was unable to find out why; there’s a suggestion the U.S. government might have stepped in to block the sharing.
If true, this would be a major step back for Anthropic, which earlier this year took a principled stance against its models’ use in surveillance and autonomous weapons and was (illegally) declared a “supply chain risk” for its trouble.
Maybe Anthropic bowed to pressure from the administration not to share its models. Maybe it was spooked by the time its models launched cyberattacks from a UK AISI evaluation, despite the fact that AISI demonstrated more competence in its response than most developers. Maybe it simply didn’t want to share.
Whatever the reason, this development seems bad. Third-party evaluators perform a crucial service, providing transparency and early warning of dangerous capabilities, and UKAISI filled an important niche. Its U.S. counterparts are not as well-resourced, and Anthropic’s own internal assessments have not inspired confidence.
More pointed accusations of model-copying
China may be copying US companies’ recklessness as well as their AIs
The accusations that China distills, or loosely copies the capabilities of, U.S. AI models have now gained new definition. The FBI, NSA, and cyber-defense agency CISA jointly accused six Chinese developers of “aggressive, malicious, and targeted distillation activities at an industrial scale,” NBC News and the Wall Street Journal report.
The accusation somewhat overstates the reality; distillation is common practice even among American AI companies. But the practice does enable Chinese developers to stick closer to the frontier of AI development than their relatively small compute budgets would otherwise allow, and while the legality of distillation itself is an open question, Chinese methods of accessing American models at scale often involve fraudulent accounts and proxies.
An article by The Hacker News suggests China might also be copying American companies’ “move fast and break things” mentality — or at least drawing from a similar inner well of carelessness. The open-source tool that Chinese company DeepSeek developed to run its models, DeepSeek Harness, was recently found to contain a bug that let an AI break out of its sandbox with a single command.
Before it was fixed, the bug would enable “danger-full-access” mode and enable writing files outside the intended workspace, a fact hackers could exploit to trick an AI running in the sandbox into breaking out (or one the AI itself could presumably use). And there seems to be some doubt as to whether the fix was complete.
The Hacker News adds:
The project’s own safety notice states that the software has not undergone a security audit and that sandboxing and approval prompts “do not guarantee isolation or prevent damage.” It tells users not to rely on the tool as their only security control for untrusted work.
Let the buyer beware, indeed.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.







