Elon Musk's peer review proposal
It's not his worst idea, but there's a better one locked behind his inevitabilism
After Amazon got Anthropic’s Fable model temporarily banned by telling the White House about a jailbreak it had discovered for it, I speculated that AI companies might find it in their interest to devote considerable resources to red-teaming their rivals’ models, looking for ways they might be unsafe, and then tattling on them.
In an interview with The Economist this week, Elon Musk suggested a formalization of that idea: a system where the leading firms peer-review each other’s frontier models and hold a safety call every few weeks. He says the labs would have an incentive to “keep the others honest,” because:
If there was something that worried them, you’d tell the government [...] If there’s something that was worrisome and that company was not doing anything to address that risk, then that would be the moment for government to step in and take action...
If you’re not going to stop the AI race, this is actually one of the more sensible, easy-to-implement proposals. It’s not adequate, but I think it would be net-positive. I’ve seen suggestions this may require a green light from the government, in the form of assurance that this wouldn’t run afoul of antitrust laws, but I think that would be easy to get from this administration.
But it’s clear from this interview that Musk thinks it would be better to stop the race if we can. When asked if he still believes there’s a “10 to 20% chance of killer robots wiping out humanity,” he dodged the question, saying he still thinks the risk is “not zero,” but that he decided to “look on the bright side” because he doesn’t see any way to stop “this incredible momentum of AI and robots,” and because he thinks the most likely outcome is “incredible abundance for all.”
He was asked about his odds twice more in this interview. The third time, it was phrased as, “Would you go into one of your rockets if you thought there was a 10 to 20% chance of it blowing up and killing you?”
Musk replied, “Yes, but let’s say you can’t do anything about it.” He then recounted how he had tried to stay out of AI and then tried to make it safer by founding OpenAI. But looking back, he thinks that:
These actions have actually resulted in knock-on effects that accelerated AI, which wasn’t really my intention. So it just seems like all roads lead to the acceleration of AI. So then I’m like, okay, well, you can just sort of be sad about it or join the club, I suppose.
What happened to you, Elon? What would past you think of present you? You were the guy who aimed for the stars and never let anyone tell you “no.” If there was a large asteroid headed toward Earth, I don’t think you’d roll over and let it hit us just because people said it was unstoppable, and I certainly don’t think you’d try to make it hit us faster. Your Anakin Skywalker arc makes me sad.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


