
Leaders at four top AI labs all indicated agreement today on the need to “pace the frontier” of AI development.
The anchoring proposal comes from Anthropic’s Dario Amodei, who posted a 3,800 word essay called “We Must Pace the Frontier,” which I’ll get into below. But it’s not clear whether the other three leaders — OpenAI’s Sam Altman, SpaceX’s Elon Musk, and Google DeepMind’s Demis Hassabis — are prepared to sign on to all or most of Amodei’s specifics.
Altman kicks most of those down the road, saying “We’ll have more to share soon,” but says that one of Amodei’s proposals, independent evaluators embedded at the companies, “is a great idea, and we’ll do the same.” (Amodei’s essay commits his company to doing this right away, whether or not other companies follow suit.)
Hassabis says Amodei’s essay “points towards the right path forward” but that “the details need working through.”
Musk only posted, “Dario is right” over a retweet of Amodei, and a part of me wonders if he read past the title. Musk had casually dismissed Jacob Coxon’s viral resignation from Anthropic the other day as a “setup.” But he has also expressed plenty of deep concern about powerful AI in the past, including recently, so who knows?
And of course we have yet to hear from anyone at Meta. I exclude the leading Chinese companies from this discussion, because Amodei does, too. We’ll get to that.
We should also be up front about the fact that Amodei isn’t calling for a stop to the AI race:
To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this.
That’s unfortunate, because I think continuing those things means we just die a little bit later than scheduled. But time bought is time that could be spent implementing stronger measures; these will need to be decided and enforced by governments, because the window where we might have been able to safely trust the labs to act responsibly has closed.
The essay
Amodei opens with a change of heart, saying the acceleration of AI development, aided increasingly by AI itself, means that we’re no longer in 2023. In that relatively quaint yesteryear, he had said that slowing AI “made little sense” because trying to study AI risks with the models of the time was “like trying to study the psychology of humans by performing experiments on bacteria.”
The swarms are also a factor:
Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet.
[...] It’s also easy to dismiss [the OpenAI-Hugging Face incident] as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.
The essay’s core is a “three-step plan”:
The first is the embedded independent evaluators mentioned above. I like the idea as a thing that can happen quickly and which may turn up more acute problems the government should investigate. Cynically, I suspect Amodei unilaterally decided to do this because OpenAI probably has a lot more to hide than Anthropic, and would look even guiltier if it refused to go along. I’ll believe Altman’s matching commitment when I see trustworthy evaluators putting out reports reflecting far more access and far fewer redactions than the one METR was invited to share about the Hugging Face swarm.
Amodei’s next proposal is “democratic coordination,” but he specifies later that he means “pacing within democracies.” There’s nothing democratic about it. He’s just saying that the companies residing “within democratic countries” should coordinate on safeguards and keep China from catching up. My political science is rusty, but I think the correct word for that is “oligarchy.”
His third proposal is “global coordination,” or coordinating “with authoritarian governments, to the extent this is possible.” In this essay and elsewhere, you can tell that Amodei doesn’t really believe anything like a “full pacing” or a global “pause” is possible unless “ironclad” verification tools can be developed. He supports “floating” such an idea, but mostly positions the Chinese Communist Party as a fundamental limitation on how far the race can be throttled back. I detect optimism only about finding common ground on prohibiting “certain narrow and obviously dangerous uses of AI, such as [...] the production of biological weapons.”
How would companies measure and set the pace?
Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be.
I think the era when capability assessments could be trusted ended earlier this year. The AIs are now too aware that they are being tested. The model card for OpenAI’s Astra, in particular, gives strong reason to suspect the model of sandbagging its performance.
We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI.
Amodei says he actually thinks these are even more “gameable,” I would guess because of creative accounting that could be done about kinds of compute and types of training. I think this is all the more reason to halt frontier development outright and put the burden of proof on researchers to show that their work won’t boost model capabilities.
Wouldn’t it be lovely?
I skipped over a whole section where Amodei daydreams about what Anthropic could do if it had even 1-2 years of slack to focus on understanding model behavior and tightening up operations. Basically, he dreams of a company that isn’t such a clown show of configuration errors and alignment failures.
The bottom line
Just before his conclusion, Amodei argues that “even if we cannot achieve formal agreements, simply changing informal norms may have some value” (emphasis his). He seems to be hoping that more dialogue between labs about the crazy stuff they’re seeing might result in a little more caution on average.
I can’t be mad at that, but I also can’t help but read it as Amodei saying the worst companies are making them all look bad, and that they should share tips on avoiding the most obvious blunders that will get them all shut down. For all their talk about feeling trapped by the race, I think they fear getting shut down more than they fear human extinction.
Pacing isn’t halting. Tweets of agreement aren’t regulation. We mustn’t be fooled into thinking these companies are meeting the moment. They want to keep building superintelligence, and they’ll keep gambling with our lives if we let them.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


