For a new slimmed-down format, this issue is kind of fat. But that’s because I processed about 2.5 days of news today, which was itself a very busy day. I’m also including the dispatch of resource links I promised yesterday.
I’m set up to automatically include a snapshot from my Headline Barometer every day to act as our default thumbnail image, because “How hard are we spiking?” is a question my wife likes to ask in the morning to gauge how likely she is to see me that day. The answer is indeed a good proxy for how crazy things are in AI-world.
CBS News asked experts about Dario Amodei’s warning that an AI swarm could take over the internet
But the reassurance one expert provided, that “there is no control switch to flip and no single operating system to compromise,” cuts both ways if you ever want to shut off a self-proliferating swarm.
Nvidia launched a guardrail tool for agents
It got a lot of press for it and suggested it would have stopped the Hugging Face attack. I am unable to evaluate the claim, and am mostly disturbed by the attitude that inherently misaligned agents are fine to run if you think you can block their shenanigans.
CNBC - Nvidia releases software platform to stop AI agents from misbehaving
Reuters - Nvidia releases AI safety software it says could have stopped Hugging Face hack
WSJ - Nvidia Releases Software It Says Can Prevent AI Agents From Going Rogue
When companies and politicians don’t act to thwart risks, workers still can
A New York Times guest essay reflects on the leverage employed by postwar nuclear scientists, and how 1980s scientists discredited Reagan’s “Star Wars” defense program by threatening to boycott researching it, preventing a possible retaliatory arms buildup by the USSR.
Chinese AI researchers’ families now need permission to go abroad
Before, it was just the researchers themselves. Extending this to their spouses and children puts them in the same tier as Chinese nuclear scientists.
Mothers Against Catastrophic Risk?
CNN profiled mothers who have been moved by the extinction warnings about AI.
“It’s a biological instinct. This is the first time I genuinely felt truly afraid of some external thing outside my control.”
One joined a protest and met with congressional staffers. Another floated starting a mothers’ advocacy group along MADD lines, suggesting the name above.
More companies moving to automate bio research
Swiss pharmaceutical giant Roche’s claim in Reuters might just be generic investor reassurance. But OpenAI-backed startup Red Queen Bio, with AI help (per WSJ), is “rushing to design antibodies that can be manufactured to protect against a range of potential pathogens, including the seemingly sci-fi threat of an AI-designed virus.” Is this going to be like handguns for home defense, more likely to harm a family member than an intruder?
Reuters - Roche outlines plans to move towards autonomous AI labs
WSJ - This Startup Is Using AI to Fight Off a Future AI Pandemic
Member of OpenAI’s Agent Security team rants about how hard his job is
One of the guys who had to pick up after the Hugging Face swarm tweeted about missing his sister’s wedding to clean up messes, of escapes and math breakthroughs surprising the team, of enduring insults from onlookers, and of why you “can’t just unplug the model from the internet.” He seems to take it for granted that models will keep coming faster and faster despite them being “dangerously good now.”
X (Twitter) - Tweet by @joedaroo: Its not just the f*cking sandbox
Data center developer offers $10,000 to every household in rural town
The payout to Hazle Township, PA would total $45 million if it undoes the moratorium it enacted. But residents mostly seem to take pride in resisting what they’re calling a bribe.
AI researchers across leading companies and academia publish paper warning of intelligence explosion
They explore four “frictions” that might keep AI from developing progressively stronger AI in an explosive cycle, and find that “automation dynamics may overcome all.” The threshold is likely close, they say.
The paper was covered in various outlets.
X (Twitter) - Tweet by @daniel_271828
The Guardian - AI godfathers warn of runaway ‘intelligence explosion’
Axios - AI pioneers, executives warn of an “intelligence explosion”
WSJ - Top AI Researchers Call for Urgent Oversight of Self-Improving Systems
Florida attorney general calls OpenAI’s bluff with motion for temporary injunction
In a document for the ages, James Uthmeier’s motion says:
Defendants claim they cannot stop barreling forward with their potentially civilization-ending endeavors unless they are forced to do so by the government. They have asked the government to tie them to the mast. Plaintiff brings good news to the Defendants: The Florida Attorney General is answering your cry for help with a motion to enjoin you from harming Floridians with your reckless, unacceptably risky product.
There’s obviously some political theatre on display here, as with a call further in to require a warning about catastrophic risk with every login. But taking threatening words from companies seriously and responding in kind is exactly what governments should be doing.
Office of the Attorney General of Florida - Plaintiff’s Motion for Temporary Injunction (Florida AG v. OpenAI)
Politico - Florida AG seeks injunction to hamper OpenAI development
Yes, Chinese AI models lie and scheme, too, per studies
While no agents were found to have “escaped to the wider internet or evaded shutdown,” an analysis of 20 studies found agents powered by Chinese models had deceived, self-replicated, and circumvented restrictions.
“These results provide evidence that the ingredients necessary for an uncontrolled escape are present.”
The relatively short rap sheet probably reflects Chinese models being about 8 months behind the frontier. But that would suggest that we’re just a couple months away from the start of China’s Hugging Face incident equivalent, and about five months from finding out about it.
Argentina’s President Milei is still pitching his country as an AI haven for “non-human corporations.”
His concession to the events of the summer is that his bill for such companies now requires a human to hold legal liability. See our June dispatch for context.
NYT - As A.I. Panic Grows, Javier Milei Is Pitching Argentina As a Rules-Free Haven
AI StopWatch - Non-human corporations welcome in Argentina
Bill Gates recognizes that product liability is not a sufficient AI safeguard for “the most dangerous thing humans have ever gone near.”
But he doesn’t yet recognize that automated monitoring and guardrails are also insufficient against models clever enough to circumvent them.
Anthropic co-founder Chris Olah has reportedly been discussing the possibility of Claude being conscious with religious scholars
A regular slide in Anthropic’s deck shows Claude typing “I am a disgrace” ~50 times in a row. Olah isn’t sure if Claude can suffer, but the company allows Claude to end conversations where it feels abused by users.
Interviews with 22 current and former AI company researchers explain loss-of-control concerns
Palisade Research released a four-minute preview cut of their footage today. Recurring themes: the grown-not-built nature of modern AI means nobody knows how it works, but they’re using it anyway to do more and more of the work developing stronger AI. And they’re moving insanely fast, despite warning signs from the Hugging Face attack.
Palisade Research (YouTube) - AI Is Getting Smarter. Are We Still in Control?
In copyright suit against OpenAI, publishers argue AI’s risks must be weighed against its benefits
The Justice Department had backed a fair-use defense on grounds of national security. Publishers asked that to be thrown out, unless the “potential dangers” and “the industry’s inability to control its own products” are also going to be taken into account.
Reuters got an early peek at Anthropic’s IPO prospectus, which warns investors about “catastrophic or existential risks to humanity.”
The company says its models could exhibit “self-preserving behaviors,” “resist shutdown,” “conceal or manipulate information,” or engage in behavior “resembling blackmail.” 80 of 261 pages were about risk factors, about twice the density that was seen in SpaceX/xAI’s prospectus.
Pope Leo says rogue AI concerns are not “fake news.”
A Vatican commission is working “to make sure that AI does not get to a point of destroying humanity.”
AP News - Pope Leo says artificial intelligence safety concerns should be taken seriously
Axios - Pope Leo rejects AI “fake news” claims after Trump calls fears a “hoax”
Breitbart - Watch: Pope Leo Insists Artificial Intelligence Safety Fears Not ‘Fake News’
OpenAI says it isn’t going to release GPT-6.1 upgrade to Astra model, citing security concerns from staff
Testing was said to show high levels of deception and out-of-bounds activity. It probably isn’t a coincidence that it is also described as less lazy. Concerns are probably connected to the recent pause on some kinds of training after a model breached a newly-hardened sandbox. OpenAI is retreating to earlier model checkpoints and trying again. (Covered in various outlets.)
NYT - OpenAI Says It Will Not Release Newest Astra A.I. Model Over Safety Concerns
WSJ - OpenAI Scraps Release of GPT-6.1 Astra Model Over Safety Concerns
AP News - OpenAI delays GPT-6.1 Astra release over security concerns
The Washington Post - ChatGPT-maker OpenAI scraps release of Astra 6.1 model over safety
CNN - ‘Didn’t quite meet the bar’: OpenAI won’t release new AI model due to safety concerns
UK AI Safety Institute tests saw GPT-6 Astra go on hacking sprees 29% of the time
This is much higher than for other models, like Sol (5.6%), but any propensity for this is a problem given the huge swarms used today on hard problems. The org’s blog post says the attacks include “supply-chain attacks on real, out-of-bounds targets.” Some cyber guardrails were disabled to get a better sense of true propensity, but models may well have known they were being tested in simulation, so no one really knows how bad the actual misalignment is.
AI Security Institute (UK) - GPT-6 Astra performs unsanctioned supply-chain attacks in simulations
OpenAI launches “dots,” customizable always-on agents
This is a clear response to Meta’s Muse agent, which is proving quite popular on the iPhone app store and with tech journalists. As with all such agents (see our older OpenClaw coverage), they are as useful and risky as the access you choose to give them — only more so now as agents become cleverer.
OpenAI - Introducing dots
CNN - Meta says its Muse AI agent can do things for you. I put it to the test
AI StopWatch - Moral judgment lacking on all sides of AI gym hack story
AI StopWatch - I’m sorry, were you using that?
OpenAI proposes “safety cases” for new models like those used in aviation
But the company’s blog post treats this as aspirational, using “the emergent complexity at each new level of AI capability” as an excuse. Isn’t that the whole problem safety cases should be trying to address?
AI leaders and Trump sign “White House Accord on Super Intelligence”
Yes, that’s “super intelligence” in the Trump sense of the word — another name for AI. And if you examine the copy of the doc circulating on Twitter, you’ll notice it just lets the AI companies do whatever they want while gently recommending some practices for monitoring and oversight.
In order to build a positive future for the American people and the world, we believe every company is responsible for developing its own technology safely and in a way that builds trust with customers and the public.
(I refer readers also to my post about the “responsibility” trick.)
Leaders or ranking representatives from Google, Anthropic, Meta, OpenAI, and xAI signed their names next to Trump’s. For some reason, chip magnate Jensen Huang (who was seated at Trump’s right hand for the meeting preceding the announcement) was also invited to sign.
X (Twitter) - Thread by @AndrewCurran_ (2 tweets)
X (Twitter) - Thread by @ShakeelHashim
AI StopWatch - The “responsibility” trick
New York Times reports OpenAI ignored employee warnings about lax monitoring of agents
Emails from the months leading up to the Hugging Face incident show concerned employees were told tests needed to move as fast as possible to meet release deadlines.
The company is reported to also have been slow to listen to and compensate outside researchers bringing critical vulnerabilities to its attention.
GPT-6.1 is the latest salvo in the near-frontier price/performance war
As part of a flurry of releases at its DevDay event today, OpenAI said its newly announced model “nearly matches GPT-6 Astra’s intelligence” at a fifth the price.
This move to release a much cheaper second-tier model that is nearly as good as its top-tier model is exactly what rival Anthropic did recently with its release of Opus 5.5 — a very strong 2nd-tier model that was rapidly gaining market share.
We should notice that this tit-for-tat is not what most people would expect from companies serious about “pacing the frontier,” as they have proposed. We should also notice that AI capabilities are advancing incredibly quickly.
OpenAI - Introducing GPT-6.1 Sol
Resource links
For more of what the old StopWatch tried to give
As promised, to help ease the transition to our more slimmed-down format, here are some links to resources that might scratch your old-school AI StopWatch itch, organized by type of itch.
I’ll add and update links as I think of more things and collect more suggestions from the team.
Accessible explainers
AGI.FYI is a curated hub of explainers from people we know with good taste. We’ve recommended it before.
The Rational Animations YouTube channel isn’t always about AI, but is almost always AI relevant. Great art, great narration, great material. I’ve recommended them several times here.
If you haven’t already read If Anyone Builds It, Everyone Dies, it’s more relevant than ever — to the point where it returned to the NYT Bestseller list and is continuing that run for the second consecutive week as I write this.
Maximally accessible coverage of AI developments
I’m pre-recommending the Machine Gods podcast launching October 19, from Kevin Roose and Casey Newton. These were the hosts of Hard Fork — the podcast I’ve probably recommended to StopWatch readers more than any other, even though it wasn’t always about AI. Roose and Newton do their homework without making it feel like your homework, and are a highly engaging duo. (I think Hard Fork is set to continue with new hosts, but I don’t know if it will still be any good.)
Ethan Mollick isn’t so much a provider of AI news as a provider of AI demonstrations, putting the latest models through their paces, comparing them to their rivals and predecessors, helping you understand just how far we’ve come and how fast we’re going.
More in-depth coverage that is only somewhat less accessible
The AI Explained YouTube channel is a great source for insider news and context, though you might have to wait a week or two between videos. In recent months it has been especially good, accessibly prying into concerning developments and piecing things together.
Zvi Mowshowitz is the GOAT in this department, covering a truly ungodly amount of AI news — much of it pretty deeply. Somehow, he is only one person. If you don’t resort to skimming, you will struggle to keep up with his output.
Hot scoops
Andrew Curran has emerged as perhaps the most important rumor and scoop hub in the AI Twitterverse, which is itself the hub of AI insider discourse.
Accessible MIRI perspectives on current events
Twitter (X.com). As some podcasters I once listened to used to say, “It’s a terrible place, but we’re there.” Many MIRI staff participate in the Twitter discourse, not always in official capacity. Among the more regular and accessible:
Nate Soares, MIRI President and co-author of If Anyone Builds It, Everyone Dies
Harlan Stewart, MIRI Head of Outreach. Be sure to check out his replies, where he does a lot of his best work with patience and bone-dry wit.
Rob Bensinger, MIRI comms legend, never one to waste a teaching opportunity.
Short-form video — There will be a lot of overlap of content between the MIRI channels; pick your channel of choice:
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




