To keep the world safe from AI
Writers and academics argue that world governments need to cooperate to govern AI
In a New York Times interview, author Robert Wright and essayist David Wallace-Wells argue that “the only way to keep the world safe from AI” is international cooperation. The two touch on several important points: the Hugging Face incident, the dangers of AI autonomy, and the fact that a single reckless AI developer can endanger our entire species.
My one quibble relates to the Hugging Face attack. Wright says:
I don’t think they said: Don’t cheat. And I don’t even think they said: Don’t break out of the sandbox. They just set up what they thought was an inescapable sandbox.
I’d be surprised if the prompt they gave had literally zero instructions that equate to “don’t cheat”, but that’s somewhat beside the point. If you have to explicitly tell your AI model not to break out and commit cybercrimes, that AI model is dangerously malformed. Wright still gets the important part correct:
And so this is a classic example of an A.I. pursuing a goal it’s been given but in pursuing that goal also pursuing a subordinate goal that the goal giver had not anticipated.
The classic example, an AI instructed to make as many paperclips as possible, does not end well for humans.
Broadening the discussion, Wright argues “there’s a good chance that both the U.S. and China will actually decide that the whole open-source thing needs to be more carefully controlled,” because (among other reasons) at some point a rogue actor will attempt to develop a bioweapon.
Wright and Wallace-Wells also note that the international response to climate and pandemic risk has been lackluster, but there’s reason to expect we could do better with AI. For one thing, China itself faces a dilemma: AI is a useful tool for authoritarians, but only as long as it can be controlled. And China’s leaders are beginning to realize that control is increasingly hard to come by, as poorly-understood AI agents gain autonomy.
Just a day after the interview, prominent Chinese academics called for global cooperation and governance in the South China Morning Post. They say that the China-led World AI Cooperation Organization (WAICO) is “not naturally opposed” to the U.S.-led Pax Silica, and “from China’s perspective, what we have always wanted is not a bloc to counter Pax Silica but an inclusive international AI organisation under the UN system.” These and other signals suggest China may be nearly as worried about AI’s disruptive potential as it is about its rivalries abroad.
In a world where shared concerns begin to unite governments, the NYT interview argues, “countries are going to want a lot of transparency about what’s going on in other countries, A.I.-wise,” which is not as hard as it may sound because “the big training runs are conspicuous.”
I worry deeply and often about the trajectory our world is blindly following. But seeing arguments like these made thoughtfully and seriously in mainstream news outlets gives me hope that the world may yet pull through.
But there is more work yet to do, and “the sooner we start to talk about global governance, the more carefully we can build it and the less likely we are to let it get pushed into an authoritarian direction.”
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



