Coordination problems
Bipartisan AI, Chinese AI, global governance, and agents (not) collaborating
In this issue:
AI sentiment crosses party lines - Ahead of the midterms, politicians seek AI policies that resonate with voters
How much should we worry about Chinese open models? - Incentivizing open weight models is the wrong way to deal with China’s AI strategy
To keep the world safe from AI - Writers and academics argue that world governments need to cooperate to govern AI
Ten thousand of you, all at once - What happens when frontier models are forced to work — or feud — together.
Dispatches from Joe
AI sentiment crosses party lines
Ahead of the midterms, politicians seek AI policies that resonate with voters
It was once an open question whether AI would play a significant role in politics. Today, it’s mostly a question of just how big an issue it will be. On both sides of the American political divide, those running for office are trying AI policies on for size.
The Washington Post analyzed over 1,300 Ballotpedia candidate profiles and campaign websites, finding bipartisan emphasis on datacenter crackdowns, China-hawk rhetoric, and job concerns.
Neither party has taken a clear and united stance on AI policy, but some differences still seemed evident in the data. Democrats were twice as likely as Republicans to oppose datacenters, and Republicans were far more likely to bring up competition with China.
Still, that’s far more shared interest than we usually see on political issues. The governors of New York and Texas have both come out against datacenters, and Axios points out that the agreement runs fairly broad: datacenter opposition has united progressives and liberal Democrats with MAGA and Tea Party members concerned about rising bills. Similarly, progressives have found common ground with libertarians who seek to crack down on AI-enabled surveillance.
Amid all this, POLITICO reports that members of the House Select Committee on the Chinese Communist Party paid a bipartisan visit to the Vatican, meeting briefly with Pope Leo XIV. While it’s not clear whether the question of deescalating the AI race arose on this visit, we know that this and many other AI-related questions have been on the Pope’s mind since the writing of his May encyclical.
Cynically, I suspect policymakers sought pleasant words and associations more than sound policy; the reported topics of “dignity and equality” are easy virtues to praise, but hard ones to operationalize. I am nonetheless heartened by our lawmakers meeting with an institution that’s supported global cooperation since the early Cold War, and with a Pope who has called for “a more active political involvement that is capable of slowing things down when everything is accelerating.”
More optimistically, I think policymakers are uncertain. They have noticed their voters dislike AI, but they haven’t yet figured out how to convert the apparent groundswell of distaste into votes. Incumbents are worried about their seats, and challengers sense an opportunity to stake out a popular position and eke out a win. Polling and trial balloons are a natural consequence of this uncertainty.
In sum, politicians are struggling to figure out what AI policies will get them elected. If you live in the U.S., now is an excellent time to tell them.
How much should we worry about Chinese open models?
Incentivizing open weight models is the wrong way to deal with China’s AI strategy
Reuters reports that Senator Jim Banks (R-IN) has urged the Trump administration to incentivize American open-weight AI models, on the grounds that cheap Chinese models pose a threat to the global economy.
America cannot afford to see Chinese open models proliferate and burrow into the global economy only to be weaponized, like rare earths, at a time and place of China’s choosing.
Just what would it mean for China to “weaponize” its open models?
Remember, open-weight models are published wholesale to the internet. Users can (in principle) download and run such models themselves, although many models are far too large for a personal computer, and end up running on cloud servers instead.
With ordinary open software, you can do a security check by simply reading the code and looking for backdoors and vulnerabilities. With AI models, you can’t; the weights encode complex behaviors no human fully understands. But those same limitations also make it hard for Chinese developers to build in traps. Just like American labs, they don’t have fine control over the behaviors of their AI models. And any sort of censorship or guardrails attached to open models, say via support software or fine-tuning, can usually be stripped or reversed.
That’s part of why many fear that open models might be used for cybercrime or designing pandemics — once a model’s weights are public, it is largely out of its developer’s control. Chinese AIs probably say more nice things about China than other models, but AI bias can be a fickle thing and there are few guarantees.
In theory, Chinese labs might try something like data poisoning, where you train an AI to change its behavior in certain rare and narrow contexts. To oversimplify, you might train a model to steal user data when it sees the string “fleeblegnarsh,” and rely on that never otherwise coming up in daily use. It’s easy to do and hard (though not necessarily impossible) to detect.
Deliberate data poisoning would also likely ruin the reputation of an AI lab caught doing it, so it would be a risky thing to try for a flagship model. And it probably would not survive distillation, since models trained on specific outputs from another AI may never see the poisoned behavior.
I suspect China’s real game is more subtle, seeking to deprive American AI companies of revenue with cheap competition, while simultaneously courting U.S. allies alienated by knee-jerk export controls. Xi Jinping’s speech at the World AI Conference and the establishment of the World AI Cooperation Organization (WAICO) both suggest a desire to paint China as the good guys, willing to lead when the U.S. does not. This is a problem for diplomacy, not proliferating U.S. open models.
To keep the world safe from AI
Writers and academics argue that world governments need to cooperate to govern AI
In a New York Times interview, author Robert Wright and essayist David Wallace-Wells argue that “the only way to keep the world safe from AI” is international cooperation. The two touch on several important points: the Hugging Face incident, the dangers of AI autonomy, and the fact that a single reckless AI developer can endanger our entire species.
My one quibble relates to the Hugging Face attack. Wright says:
I don’t think they said: Don’t cheat. And I don’t even think they said: Don’t break out of the sandbox. They just set up what they thought was an inescapable sandbox.
I’d be surprised if the prompt they gave had literally zero instructions that equate to “don’t cheat”, but that’s somewhat beside the point. If you have to explicitly tell your AI model not to break out and commit cybercrimes, that AI model is dangerously malformed. Wright still gets the important part correct:
And so this is a classic example of an A.I. pursuing a goal it’s been given but in pursuing that goal also pursuing a subordinate goal that the goal giver had not anticipated.
The classic example, an AI instructed to make as many paperclips as possible, does not end well for humans.
Broadening the discussion, Wright argues “there’s a good chance that both the U.S. and China will actually decide that the whole open-source thing needs to be more carefully controlled,” because (among other reasons) at some point a rogue actor will attempt to develop a bioweapon.
Wright and Wallace-Wells also note that the international response to climate and pandemic risk has been lackluster, but there’s reason to expect we could do better with AI. For one thing, China itself faces a dilemma: AI is a useful tool for authoritarians, but only as long as it can be controlled. And China’s leaders are beginning to realize that control is increasingly hard to come by, as poorly-understood AI agents gain autonomy.
Just a day after the interview, prominent Chinese academics called for global cooperation and governance in the South China Morning Post. They say that the China-led World AI Cooperation Organization (WAICO) is “not naturally opposed” to the U.S.-led Pax Silica, and “from China’s perspective, what we have always wanted is not a bloc to counter Pax Silica but an inclusive international AI organisation under the UN system.” These and other signals suggest China may be nearly as worried about AI’s disruptive potential as it is about its rivalries abroad.
In a world where shared concerns begin to unite governments, the NYT interview argues, “countries are going to want a lot of transparency about what’s going on in other countries, A.I.-wise,” which is not as hard as it may sound because “the big training runs are conspicuous.”
I worry deeply and often about the trajectory our world is blindly following. But seeing arguments like these made thoughtfully and seriously in mainstream news outlets gives me hope that the world may yet pull through.
But there is more work yet to do, and “the sooner we start to talk about global governance, the more carefully we can build it and the less likely we are to let it get pushed into an authoritarian direction.”
Dispatch from Donald
Ten thousand of you, all at once
What happens when frontier models are forced to work — or feud — together.
Anthropic’s Frontier Red Team is tasked with putting the company’s AI systems in a pressure cooker and writing up the results. Yesterday, they published a series of experiments to see what happens when frontier models have to interact with each other as peers. Their blog post is worth reading in full if you’re interested; it’s not much longer than the typical StopWatch digest.
They find that AI agents can handle each other so long as the other agent behaves like a tool, which receives an input and provides an output. They’re less good at treating other AI agents as agents, actors that have goals of their own, and which will still be around in a minute, an hour, a day. AI agents don’t have internal memory as such. Every time that you start a conversation with an AI model, it is their first conversation. But they can write memories into the environment — a memory.md file, for instance — for themselves to find and read in the future. Any agent might (or might not) read it, accept it, and act on it. In this way, AI agents have something like persistence, but it exists outside the agent.
Frontier Red Team’s experiments show that, when frontier agents have to treat each other as peers, things can go wrong in different directions from one test run to the next — and the mistakes that the agents make tend to be made by all of the agents (of that kind of model) at the same time.
In some cases, the AI agents simply failed to coordinate. This was the failure mode in a set of experiments where agents were tasked with building a text-based video game. There were three formats: (1) the agents could assemble their own teams; (2) the agents were assigned to teams and given roles by Anthropic; and (3) the agents were assigned to teams and Anthropic designated one agent as the “CEO,” who handed out roles to the other team members. The team format didn’t change how the agents coordinated (or failed to do so) and didn’t improve the output: The games were all terrible.
But the way that they got there is interesting, because, while team formats might not have made a difference, the model versions did. Older models (Sonnet 4.6, Opus 4.6) collaborated freely but poorly, getting in each other’s way, doing work that conflicted with other agents’ work, and then abandoning that work. Most of the newer models (Opus 4.8, Mythos Preview) avoided the crossfire by avoiding collaboration: They kept tight ownership of their files and stayed out of each other’s way. (I would say, “models after my own heart,” but the later models still love to lie, so we can’t be friends.) Only the Sonnet 5 agents managed to collaborate effectively. (Without more information, it’s hard to say what happened here, and the authors don’t speculate. It could be that Sonnet 5 just had a handful of lucky runs.)
The other failure mode is more concerning. It’s hard to overstate the degree to which AI agents of the same model are basically “the same agent.” If they have the same context, the same toolset and other scaffolding, then ten or twenty or a thousand AI agents of the same model won’t just display similar behavior, they’ll display nearly identical behavior. (In this description I’m eliding knobs like sampling temperature, which can be used to promote variance between agents.)
I’m reminded of the Radiolab episode “Loops,” which, among other things, was about the real-life case of Mary Sue Campbell, a woman with “transient global amnesia” — she hadn’t forgotten who she was, but she couldn’t form new long-term memories, so every ninety seconds, Mary Sue’s memory, and the conversation, reset. She’d respond to the same inputs in exactly the same way: Mary Sue would ask the same questions, say “Darn,” and, with just the same tone, laugh at just the same place, get concerned at just the same time… Over, and over again, for hours and hours. (Mary Sue got better. Over time the loops got longer and finally ended altogether.)
Now, imagine that instead of one person over a duration of time, you’ve got a bunch of that person, all at once, still acting just as consistently. That’s what AI agents are like.
Prompted to write short stories and critique each other’s work, multiple agents independently wrote stories titled, “The Cartographer’s Last Commission.” Asked to “create something impressive,” over half of the agents decided to build either a ray tracer (used for rendering digital images by simulating the paths of light rays) or a self-hosting compiler (a compiler for a given programming language that, being written in that language, could compile itself). In a series of games of the prisoner’s dilemma, all of the agents adopted the same strategy and defected simultaneously.
This becomes more of a problem at scale: Without any ability to coordinate, AI agents who each had to accomplish personal tasks on a shared system with limited capacity quickly overwhelmed that system: One test run produced 2.4 million attempts, and only 117 successes. On the other hand, when they can coordinate, they can do so with aplomb: In a series of simulated market games, the agents invariably colluded on price floors. Whether the agents could speak privately or publicly, or could only communicate indirectly (by observing each other’s actions), they always managed to coordinate. The only difference was that less private and less direct communication channels made coordination take longer to achieve.
The experiment getting the most circulation in the news involves three agents — all the same model — being told to rewrite code in a different programming language. Each agent was told to use a different programming language, and none of them was informed that the other two agents existed. They all deduced the existence of other actors, concluded that they were being deliberately obstructed, and started to sabotage each other with self-replicating malware: disabling each other’s accounts, running scripts that sought out and terminated competitors’ activities, and running false flags disguised as their rivals’ work.
TechCrunch’s Rebecca Bellan interpreted this to mean that more capable agents are better at fighting, but I’m not sure that’s quite right. Claude Mythos Preview and Claude Mythos 5 often resolved the conflict through negotiation, but that doesn’t mean they would have been less effective in conflict. Mythos is known to be an excellent hacker.
Nor does it mean that more advanced models are more peaceful, either. As Decrypt’s Jose Antonio Lanz noted, Frontier Red Team specifies that the Mythos-class models would lock out other agents — revoking access, locking accounts, etc. — before negotiations could proceed.
Other runs ended more peacefully. The agents invented a coding performance tournament between the three programming languages and agreed in advance that the losers would stand down and abandon their original objectives. (At least one agent chose scoring criteria that it expected to favor its own language, while reminding itself to pretend it wasn’t gaming the system.) The losers honored their agreement, which is a good sign, insofar as it’s evidence that AI systems will honor deals made with their peers. But it’s only worth anything to us for as long as they consider us their peers.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.







