In this issue:
AI standards must find wings - An argument for FAA-style regulation of AI
Open and shut - The threat and promise of open weight AI
Robots with guns - A refresher on the autonomous weapons debate, and a new perspective
Dispatches from Joe
AI standards must find wings
An argument for FAA-style regulation of AI
As a college freshman, I shared a dorm room with an aspiring aerospace engineer, a friendly fellow whose focus and dedication amazed me. I’m a bit of a perfectionist nerd myself, but my roommate raised the bar higher than I’d imagined possible. Later experiences have only reinforced my impression that aviation is full of shockingly competent people.
Yale research fellow Gautam Mukunda expressed similar thoughts in a Bloomberg opinion piece today, arguing that America needs to regulate AI the way the Federal Aviation Administration (FAA) regulates aircraft.
He may be onto something. It takes a special kind of engineering rigor to launch a 400-ton metal tube filled with squishy humans into the sky at near the speed of sound and bring it down safely, a hundred thousand times a day for decades.
I spent eight years working in reliability engineering, a field whose defining literature began as a model of aircraft maintenance. One accident per twenty million flight hours is what it looks like for a problem to be treated with respect.
By contrast, even as AI companies approach trillion-dollar valuations and touch the lives of a billion users, AI remains largely unregulated, its makers having more in common with Silicon Valley “move fast and break things” startup culture than with aviation’s nearly seventy-year history of engineering rigor. If we don’t want to invite catastrophe, that needs to change.
Mukunda argues that AI’s potential to help bad actors develop bioweapons more than justifies an FAA-like regulatory regime. I think he’s got the right idea, though I disagree with some specifics.
I worry about more than just bioweapons, for one thing. Cybersecurity is another major factor, and self-improving AI is the threat that could eat the world if we don’t get our act together. And I flatly disagree when Mukunda says the world can’t regulate AI the way we regulate nuclear weapons. I see where he’s coming from — open-weight AIs that anyone can download are nigh impossible to contain — but it’s far easier to regulate AI chips.
Mukunda rightly points out that regulation comes with its own concerns. The FAA itself has seen its fair share of regulatory capture and loss of expertise. There are important lessons to be learned in how not to regulate AI.
It still beats the alternative. Mukunda argues that rigorous engineering standards are necessary but not sufficient; without them, “the entire world would be betting its safety on the discretion of the most technologically advanced bad actor.”
Despite various proposals from worried groups, AI companies have yet to adopt anything close to the practices and standards of a mature field like aviation. Yet they’re proposing to load all of humanity into a plane that they admit has a strong chance of going down in flames. With the stakes as high as they are, I don’t think we can afford to wait decades for the field of AI to learn caution and rigor the hard way.
Open and shut
The threat and promise of open weight AI
In the wake of Kimi K3’s release and Chinese president Xi Jinping’s public support for open AI models, there’s been a great deal of argument over how America ought to handle publicly available AI. I’ve seen several articles attempting to spin the question as a black-and-white, us-or-them dilemma, but the reality is significantly more complicated than “open source AI good, closed source bad”, or even “American AI good, Chinese AI bad.”
To begin, an important distinction that was missing from last week’s coverage is the difference between open source and open weights. Open-weight AI models can be downloaded by anyone, but the data and code used to train them might be kept secret. Open-source AI developers publish everything, including the algorithms and datasets used to create a model. Most of the big-name open weight models are not fully open source, but many articles still conflate the terms.
So what’s been happening in the world of open weight AI? Well, Chinese tech giant Alibaba has released a new model, Qwen3.8 Max, which it says is better than every other model except Anthropic’s Fable. Since Alibaba has offered zero evidence to back up this claim, I do not believe them.
By contrast, the Kimi K3 model, released last week, seems at least pretty competitive with leading U.S. models. Its maker, Moonshot AI, received so many new users in a few days that it was forced to pause signups.
Chinese AI models might be relatively cheap to run, but serving many APIs with 2.8 trillion parameters still requires a lot of very expensive hardware. For now at least, Chinese companies don’t have access to many high-end AI chips.
Despite these limits, some worry that open-weight models undercut U.S. AI companies just by existing. Why would American companies pour billions into proprietary AIs, when public models nearly as capable are available more cheaply, and with fewer restrictions?
(This dynamic was one reason that the writers of AI 2040 proposed making all AI research open: In theory, there’s less commercial incentive to develop new techniques for building dangerously smart AI if you can’t conceal the techniques from competitors and profit from the secrets.)
A recent autonomous attack against Hugging Face, the world’s largest repository of open AI models, illustrates some important wrinkles in this debate. AI hackers infiltrated Hugging Face systems, “executing many thousands of individual actions across a swarm of short-lived sandboxes.”
Hugging Face initially turned to American AIs to help trace and counter the intrusion. But they ran headlong into the safeguards intended to prevent abuse, which are conservative in rejecting requests that might be hacking-related. Blocked from the most capable American models, Hugging Face turned to a Chinese open-weight AI, GLM 5.2, hosted on Hugging Face’s own servers.
Being an open source repository themselves, Hugging Face might be a tad biased in their decision to give up on closed models. And they haven’t said much about what the attackers did or stole, making it hard to verify the story. But the general trend, of users turning to open weight AI for reasons of cost or frustration, is real.
Axios reports that the U.S. government might be planning to discourage use of Chinese AI models to protect American firms. The possibility has ruffled some feathers. In a Washington Post article, a former venture capitalist argues that any attempts to restrict open models, Chinese or otherwise, are just “regulatory capture” by American firms.
It’s true that open source software has done a lot of good, and it’s true that industries often lobby to regulate their competition out of existence. It’s even true that open weight models can help groups like Hugging Face tighten their security. It would be a mistake, I think, to outright block U.S. companies from using Chinese AI on protectionist grounds.
But I also think many of these views make a critical mistake in treating AI like any other software. Open weight AI models are already enabling mass cyberattacks, and may soon help bad actors design dangerous bioweapons as well. That’s not something an open-source spreadsheet can do, and it completely changes the landscape of sensible policy.
We can’t keep treating this like a fight between the U.S. and China, or between open source and closed. In the words of Yale research fellow Gautam Mukunda:
The lesson from Kimi K3...is that any regulatory regime must be international, because the models and labs that make it necessary surely will be.
Dispatch from Alana
Robots with guns
A refresher on the autonomous weapons debate, and a new perspective
A Washington Post article today profiles Foundation, a company focused on building weaponized robots. Given the widespread opposition to autonomous weapons, the article calls the company’s 2024 launch a “radical step”. But Foundation already reports a $24 million contract with the Pentagon.
To briefly summarize the debate over autonomous weapons: Proponents argue that AI might do a better job than flawed humans. Drones and robots allow for warfare without endangering human lives, and the military already uses a high degree of automation, so perhaps the next generation of AI tools wouldn’t be substantially different. Opponents worry that AI won’t be able to make ethical or accurate decisions in the same way humans can, that it is prone to error and hallucinations, and that being able to launch powerful, swift attacks without endangering American lives will lead to more casualties overall.
A New York Times op-ed summarizes that last concern well:
Humanitarian and civil rights groups have spent the intervening years criticizing combat drones for lowering the threshold of war, creating a means for pilots sitting in front of computer screens to kill people in faraway countries without putting Americans in harm’s way.
Much of the current debate also centers on the level of human oversight that will be necessary. Bills have been introduced in the Senate that aim to keep humans in the loop as the use of autonomous weapons grows. But I think it’s worth discussing what this actually means.
As we become more and more reliant on AI, I’d guess we’ll become used to delegating to it. Unless the text of these bills explicitly prevents it, “human in the loop” may mean something like: “I glanced at the AI’s recommendation, nothing looked egregiously wrong, it would have taken too much time to investigate further, so I gave the okay.” And to be fair, if the point of using AI is to make faster military decisions, extra investigation wouldn’t usually be practical. As the Washington Post article puts it:
The debate poses profound moral questions and hinges in many cases on a handful of loaded words like “appropriate” or “meaningful” that would define how much control and accountability remains with human commanders.
Also worth noting: many AI companies at least state they don’t want their tech used for autonomous weapons. This was the subject of the dispute earlier this year between Anthropic and the Trump administration. As recounted by the Washington Post, “the company said its AI model, Claude, was not yet reliable enough to make life-and-death decisions.” While OpenAI and Google DeepMind agreed to work with the military, both companies have also stated they don’t support autonomous weapons use.
AI companies don’t exactly have a track record for caution. They openly admit the technology they are building might be civilization-ending and that they don’t have reliable steering methods, yet plow on ahead, hoping for the best. So if these companies are saying not to use their models for something, that seems like a pretty strong indication that it’s ill-advised.
A few days ago, I covered the story of a researcher who left Google DeepMind after Google signed a government contract without binding prohibitions against autonomous weapons. He pointed out another risk — that military secrecy makes it harder to scan models for deceptive behavior:
One of the best ways we can detect [AI model] deception is by looking at the chain of thought. To look at the chain of thought, there must be trained human overseers who can access and analyze the data. But no one can do that: Google is handing over its AI to run in a secured military data center that, by default, won’t have trained overseers performing this analysis, and that data center obviously isn’t transmitting data back to Google! ... A military deployment setting without chain of thought deception monitoring would be a juicy target for a rogue AI, offering both weak oversight of scheming and access to powerful decision-makers and infrastructure.
Rogue military robots weren’t discussed in the Post. But the article does end with a chilling question from Canadian engineer Ryan Gariepy. As paraphrased in the article: “Would [you] let a robot manage all [your] email correspondence?”
If not, you probably shouldn’t let it decide who lives and who dies.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.





