In this issue:
22 countries call for international cooperation to mitigate AI risks - Ahead of the UN general assembly, a coalition formed to call for an international AI oversight authority
Meta’s Muse was unsafe to use - Zero-day vulnerability allowed malware on a Mac to hijack Muse on multiple devices
Waking up America - New AI developments and more bipartisan calls for urgent action
Dispatch from Robert
22 countries call for international cooperation to mitigate AI risks
Ahead of the UN general assembly, a coalition formed to call for an international AI oversight authority
Finnish President Stubb and Norwegian Prime Minister Støre have launched an initiative to establish an international AI oversight authority, as reported by POLITICO.
As of last night, a total of 22 countries around the world had already signed on to this declaration, including the world’s third-largest economy Germany. The Netherlands, which is particularly important to the chip industry, is also among the signatories. Other key European countries, however, such as the United Kingdom, France, and Italy, were not yet on board.
The declaration addresses, among other things, the latest security incidents and warnings from industry and the scientific community:
Recently we have seen capable AI systems circumventing testing safeguards, exploiting vulnerabilities and gaining unauthorized access to real-world systems. Leading scientists and executives are warning that the pace of development could outpace our ability to manage emerging risks.
The heads of government continue to call for mandatory independent testing for the industry, as well as international cooperation to jointly address the risks posed by artificial intelligence. Signatories also pledge to establish binding international safety standards for AI development, something OpenAI also called for recently.
In an interview with POLITICO, the Finnish president and the Norwegian prime minister called on their British counterpart, Andrew Burnham, to join the initiative and use the upcoming British turn presiding over the influential intergovernmental forum G20 to advance the issue.
Burnham, who had also been pressured by the British opposition to address the issue, announced today that he does indeed intend to use the British G20 presidency to put together an international agreement on AI.
By now, frontier labs, the majority of experts, and a growing coalition of countries around the world are calling for more regulation and international cooperation. However, the signals from the U.S. and China — the two countries that matter most — remain mixed.
On the one hand, Trump has been dismissive of potential risks in recent days, calling warnings about them a “hoax” and rejecting regulation. On the other hand, his administration has certainly taken regulatory action with the restrictions on Anthropic’s model Fable and the executive order regulating frontier model deployment issued in June. Just last week Treasury Secretary Bessent stated that frontier labs must not be exempt from liability.
From China, in turn, we hear time and again encouraging signals from Xi Jinping, but at the same time, Chinese state media is also taking a clear stance against slowing down development.
Hopefully the UN Security Council briefing by Sam Altman, senior representatives from Anthropic, and representatives from the Chinese AI companies DeepSeek and Moonshot will tip the scales in the right direction.
For the growing coalition of countries that recognize the danger, this means they must use their influence to convince the U.S. and China to join international efforts for regulation. Artificial superintelligence is a danger for humanity that can’t be dealt with unilaterally.
Dispatch from Donald
Meta’s Muse was unsafe to use
Zero-day vulnerability allowed malware on a Mac to hijack Muse on multiple devices

Mark Zuckerberg said that Meta’s AI agent Muse was “built from the ground up for privacy and security.” I don’t doubt him, but I think that they stopped building six inches off the ground. In June, I wrote about how Meta’s customer service chatbot would hand over Instagram accounts to anyone who asked nicely — as it turned out, 20,000 accounts were accessed, but Meta rejected the notion that there was a problem with its AI agent. Now Meta has produced a security hole that was decidedly worse.
Yesterday, Mac security researcher Patrick Wardle published a proof-of-concept for a vulnerability that could allow someone to take control of Muse and, through that connection, anything Muse had permission to use on your devices: Did you give Muse permission to use your camera, your calendar, location data...? Then congratulations — anyone who took control of Muse could control those, too: access emails, make purchases, record audio on the microphone, etc. The attacker could also provide instructions to Muse, which would trust those instructions and act on them. The silver lining is that Muse only launched on September 8th — the exploit was glaring enough to be discovered quickly, but the window to exploit it was short.
(As computers and programs get more complicated and LLMs get more advanced, there’s a temptation to hand it all over to ChatGPT or to Claude. I think that this makes basic fluency with computers even more essential, though, so that you have a hope of catching when your AI suggests something really crazy.)
Furthermore, once the attacker got into Muse through one device, they had access to Muse on every device that account was connected to: laptop, cell phone, baby monitor — I don’t know why you would have a smart thermostat, but if you did, the attacker could probably control that, too. That’s probably just annoying in the case of the thermostat, but much more worrying if you’ve got some kind of AI-powered home security system.
Meta shipped a hotfix very early this morning to remove the secret setting that made all this possible. So, you’re safe, at least from this screw-up. (Unless you haven’t updated the app yet. If you have Muse, please update it.) Meta’s David Singleton said on X that because local access was required, “the practical risk to users of the Muse Mac app was therefore quite low.” (Muse is not available on PC.)
It’s true that somebody would need to get malware onto your computer before they could hijack your Muse agent, but that doesn’t mean that the risk is low. One technique for delivering that malware is common enough to have a name, “ClickFix.” (You’re on the internet, you get a popup that says there’s some kind of bug, but if you click this and paste that then you can fix the bug — and presto, you downloaded malware.) Wikipedia has an article on a major ClickFix-enabled attack that just hit the government of Berlin last month. So, if this is a risk for the government of Berlin, I think that it’s a risk for Ma and Pa Facebook User, too.
We recently covered a computer worm — recently built with AI models by the security firm Calif — that could spread through the widely popular Chinese app WeChat from phone to phone without anyone clicking anything. This was a different vulnerability, an insecure AI rather than an insecurity exploited via AI, but the reason is the same: the labs are racing ahead faster than they or anybody else can keep things safe.
Dispatch from Joe
Waking up America
New AI developments and more bipartisan calls for urgent action
The storm of AI news in the wake of Jacob Coxon’s high-profile resignation seems to have subsided to a dull roar, but there’s still more relevant to the extinction threat than can be adequately covered in a day. Here are some of the highlights:
Wake up, America
Republican Greg Murphy of North Carolina warns that concerns over AI dangers are real, nonpartisan, and urgent, Breitbart reports.
These are really smart people, not people with political proclivities that are pushing one way or the other... I do believe it’s something that we need to take the reins of, and I think it needs to be done right now.
Murphy said that AI governance “desperately needs national and international attention,” while still emphasizing the importance of staying ahead of China.
Appropriately, the weekday Newsmax show on which Murphy made these comments was called “Wake Up America Early.” I wouldn’t say Congress is early in addressing AI threats, but I think it’s not yet too late.
Sky nets
The Federal Aviation Administration is experimenting with a new AI-powered system for air traffic control, POLITICO writes. Several airports near Washington, DC (Reagan, Dulles, and BWI) will host the new system for a 90-day trial before it rolls out across the country. As my colleague Mitch pointed out in May, the system does not seem intended as a wholesale replacement for human controllers; it’s more of a tool that analyzes weather and flight plans and recommends time-saving tweaks.
For now, anyway. We have seen the same story play out many times already: first AI is barely able to do a task at all, then it can perform as well as an amateur, then it’s assisting expert humans, then it surpasses them entirely. (Famously, the early game-learning AI AlphaGo Zero passed all of these thresholds in a few days of training, having never once played Go against a human.)
The FAA’s “SMART” system is evidently at the “assist skilled humans” phase today, at least for a subset of routing tasks. That’s probably a good place to be, despite the extra attack surface AI integration offers to would-be infrastructure hackers. But as AI systems improve, they will likely be trusted with more control and less oversight by overwhelmed human staff. That all these systems increasingly depend on the same poorly understood AI creates what assessors might call “correlated risk.”
A rogue or compromised AI with access to a country’s air traffic control systems — and other critical infrastructure — could do an awful lot of damage. We’ve yet to see what can happen when AI really takes off.
Lean machines
Today Anthropic rolled out its latest AI model, Claude Opus 5.5, said to meet or beat prior models in key domains and to have cyber skills on par with those of the limited-access hacking genius Mythos. It’s also claimed to be less verbose and annoying to converse with.
Anthropic emphasizes efficiency gains in Opus 5.5; the headline metric is that the new model “costs 40% less to run than Opus 5” and can accomplish the same tasks as other models much more cheaply.
Hours later, OpenAI announced streamlined siblings for its GPT-6 Astra model, dubbed GPT-6 Sol and Luna. Again, the main focus seems to be on efficiency and cost savings rather than entirely new capabilities.
This efficiency is likely the result of algorithmic improvements, or new ways of arranging AIs so they can do more with less. Along with compute scaling (more operations with more chips), algorithmic gains are one of the major things driving AI capabilities. And at least some of these latest gains were likely developed by the AIs themselves.
In its announcement, Anthropic describes the measures it has taken to reduce the risks, including rerouting dangerous-seeming cyber and bio queries to weaker models like it does for Fable, the consumer-facing version of Mythos. It admits this won’t be enough to mitigate the dangers of later AIs...
For [future] models, we do not assume the measures described above will meet that safety standard on their own.
...but it doesn’t say it’s going to stop. Nor does OpenAI.
This is despite the AIs themselves sometimes warning the companies that what they are doing is dangerous. Anthropic’s latest system card tells us that “Claude Opus 5.5 often reasons that having input into training or deployment may be risky as a result of giving it too much capability or influence.”
As Madison Mills of Axios observed today, AI companies like Anthropic may make concerned noises about safety, but they have trillion-dollar incentives to push the frontier.
Freeze in place
In a personal writeup that also appears in the New York Times, former U.S. ambassador and national security advisor Susan Rice cogently summarizes the urgent need for AI governance:
Dario Amodei, CEO of Anthropic, and leaders of three other frontier AI companies have become so alarmed by rapid AI advances that they now pledge voluntarily to “pace” the development of new AI models, so that safety features can catch up. The firms are also promising to grant independent experts employee-level access to monitor and report on model development.
These are necessary and, if implemented, welcome steps. But they do not go nearly far enough.
“Pacing” amounts to the industry policing itself, while its leading firms simultaneously prepare for blockbuster IPOs. There is no way to ensure frontier companies will slow down sufficiently or invest enough in safety, which they have long subordinated to speedy development. So long as frontier training continues, the risk remains that additional accidents occur with devastating consequences. Moreover, as Amodei stresses, meaningful restraints must also apply to global competitors, particularly in China.
Much more can and must be done to ensure that the march toward super-intelligence does not result in catastrophic harm to humanity.
Rice calls on Presidents Trump and Xi to “freeze in place” AI training in the U.S. and China, and to arrange for scientists from both countries to develop solutions for the threats posed by superhuman AI.
She also suggests that:
In parallel, the U.S. and China should work intensively to reach a bilateral agreement on verifiable safeguards and acceptable uses of AI, then lead efforts to codify these agreements in a binding and verifiable international treaty.
...and she points out that the Chinese Communist Party has a vested interest in maintaining human control over AI systems that would strongly motivate China to cooperate on such an agreement.
With each new voice that calls for a halt to the deadly AI race, I feel my own hope for our future growing.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.






