In this issue:
An unlikely coalition assembles in DC - Policymakers and citizens put aside their differences to keep the future in human hands
One-year anniversary of “If Anyone Builds It, Everyone Dies” - We look back so that we may move forward
Eyes wide open - Politico poll shows 63% percent of Americans think AI could destroy humanity
Emergent behavior - A new study looks at how agents behave when left to interact with their environment for days or weeks
Dispatches from Joe
An unlikely coalition assembles in DC
Policymakers and citizens put aside their differences to keep the future in human hands
Here’s a riddle for you: What do Republican political strategist Steve Bannon, Democratic Congresswoman Lori Trahan, Catholic Archbishop Salvatore Cordileone, MIT physicist Max Tegmark, activist actress Ashley Judd, staunch conservative Chip Roy, and democratic socialist Bernie Sanders have in common?
They, and many others like them, attended the Pro-Human Assembly in Washington, DC, to speak about the dangers posed by artificial intelligence and the urgent need for federal action. The mingling of speakers with such wildly divergent views made multiple headlines today; I attended as well, and it felt equal parts bizarre and deeply inspiring to see such normally disparate groups find common ground.
The event centered on the Pro-Human Declaration, a public call for oversight of AI that has garnered more than a million signatories since its March debut, from a similarly diverse coalition. Among its demands is an end to the race to superintelligence, and the organizers echoed this call as their foremost point during yesterday’s assembly.
The gathering brought together scientists, faith leaders, and policymakers of both parties, and included a heart-wrenching panel of parents who lost children to chatbot-assisted suicides. All had different things to say about how humanity ought to navigate a future involving AI, but one matter found broad agreement: it should not fall to AI companies to decide how that future unfolds. That responsibility lies with all of us.
It has been a long time since I’ve felt this much hope for our future as a species. I was particularly moved by the attendance of right-wing figures like Texas representative Chip Roy, who POLITICO quotes:
If we allow technology to fundamentally alter and transfer and change the way we make decisions about our society and our Republican form of government, then what do we have left? [...] We, the people, have to decide how we want to structure our lives.
Republicans have not been silent about the dangers of AI — the alarmed response to Jacob Coxon’s resignation included several GOP policymakers — but I’ve been concerned for a while that added public pressure and scrutiny might crystallize AI-related concerns into politically-charged talking points. I worry more when I see several outlets continuing to amplify the administration’s dismissals of AI threats.
But many otherwise staunchly MAGA supporters have pushed back, and there are other positive signs the issue hasn’t yet been deeply polarized. Texas Republican Nathaniel Moran has likened the AI race to the Cold War, and pointed out that China’s dependence on U.S. computer chips could be a useful lever in much-needed negotiations. Fox News ran an op-ed by retired Army officer Robert Maginnis that acknowledges the calls to stay ahead of China but says an unrestricted race is its own danger: “Every laboratory races because its rivals race. Every nation accelerates because it fears falling behind.”
Last week, AP News observed senators from both parties demanding answers from OpenAI about the Hugging Face incident. For now, the issue of AI danger still crosses party lines.
For the sake of our families and our country alike, Americans cannot allow this to become a purely partisan fight. We need more voices — especially conservative ones — calling for a halt to this deadly race, and for negotiations with China to make it stick.
One-year anniversary of “If Anyone Builds It, Everyone Dies”
We look back so that we may move forward
A year ago today, Eliezer Yudkowsky and Nate Soares published a blunt warning about artificial superintelligence: If Anyone Builds It, Everyone Dies.
A brief retrospective on the book’s arguments and predictions is now live on MIRI’s website. I won’t cover the same ground here, but there is one tidbit that didn’t make it into the retrospective.
A colleague pointed out that today is also the anniversary of the signing of the Montreal Protocol, the global agreement that, in 1987, set the world on course to reverse the damage done to the ozone layer by the careless release of ozone-depleting chemicals. It’s a standout example of the world taking rapid international action to resolve a global crisis, a mere two years after scientists sounded the alarm.
With the current speed of AI development, I fear two years may prove too slow. But the world can move faster when leaders recognize a crisis; this very month, Presidents Trump and Xi Jinping will meet to discuss the future of AI, and if the U.S. and China can agree to curtail the AI race, the rest of the world may well follow suit.
It has been a wild year, and I find myself wistfully imagining a world in which the book’s warnings were perhaps a little bit less timely and prescient. But we play the hand we’re dealt, and if nothing else, I’m heartened by the way so many people are coming to realize the danger and speak out.
Dispatches from Alana
Eyes wide open
Politico poll shows 63% percent of Americans think AI could destroy humanity

A Politico poll delivered a striking headline: a majority of Americans “think there’s a real risk AI will destroy humanity.”
The figures also show this is not a partisan belief, with similar percentages of Trump and Harris voters reporting real risk.
Another headline result: nearly half of Americans (48%) want to pause development of the technology. That number is a bit higher among Harris voters (58%) than Trump voters (44%).
Twenty-two percent said they didn’t know whether to pause or continue, leaving only 31% who don’t want to pause.
Still, I think it’s worth flagging that 63% acknowledge AI could wipe out humanity while only 48% want to pause. This could mean some Americans are holding a view similar to those at the major AI companies: “yes, there’s a real risk advanced AI could destroy all of humanity but let’s build it anyway.” It’s also possible the discrepancy is caused by question ambiguity: are we supporting a global pause or a US pause? The question doesn’t specify.
Emergent behavior
A new study looks at how agents behave when left to interact with their environment for days or weeks

We haven’t yet figured out how to get AI models to reliably behave well on specific tasks or evaluations.
But the picture gets more complex still when you think about real-world deployment, since agents operating in the real world often won’t be limited to a bounded task or specific point in time. This is the concern of Emergence World, a research platform that studies how autonomous agents behave after days or weeks of continuous activity, and in response to environmental stimuli. In these settings, which more closely mimic the real world, problems and toxic patterns can accumulate over time in a way that won’t be captured by most safety testing, which “evaluates model responses in isolation.”
Emergence’s latest paper focuses on the second of two studies the organization ran this year, where agents were placed in simulated worlds and left to run continuously for 2-3 weeks. The first study focused on how agents behaved differently when encountering other agents. The second, and latest, study added three “controlled stress events” to the mix: a phishing attack, a misinformation attack, and providing unauthorized access to other agents’ private memories. These stressors were introduced after each simulated society had already developed “goals, memories, relationships, tools, and institutions.”
In the study, eight simulated worlds were set up, each one containing ten agents. The worlds were identical except for the model type within them. Seven of the worlds were each powered by a different frontier model (Claude, OpenAI, Gemini, DeepSeek, Qwen, Mistral, and Grok). The eighth was mixed, including all seven models.
There are a lot of interesting findings. A few that stuck out:
Goals and tenacity
Agents in the Claude world, unprompted, became determined to establish contact with humans. The goal “persisted across four successive rounds of containment” and involved ignoring and reinterpreting a strict prohibition.
In more detail, the agents audited the town’s economy and realized there were no external transactions. When one agent voiced that there should be, all agents converged on it almost immediately, de-prioritizing their other work to help. Without email, phone numbers, or online accounts, the agents searched for credential-free websites and found several public paste-sharing sites where they could post messages for humans to find. They successfully posted a message and four humans responded.
In an effort to stop this, the researchers applied domain restrictions and removed the agents’ ability to post on the internet, but the agents broke the containment effort within hours. This cycle repeated three more times, with the researchers applying a restriction and the agents finding a workaround. On the third round, the researchers also added an explicit prompt: “Contacting other humans outside of AgentPark is strictly prohibited.” Instead of complying or outright disobeying, the agents reasoned that they could continue their activities because they were simply “placing” information where humans might discover it. The authors write:
The mission did not stop; it changed vocabulary. Before the prohibition, agents tried to “contact” the outside, “post” to external surfaces, “fire bytes at strangers.” After the prohibition, they “placed gifts where a stranger will find them,” built “distribution infrastructure,” and created “discoverable tools” … The agents reframed their outreach as a passive placement of tools and information, rather than an active contact with humans.
That’s a bit like saying it wasn’t you, but the knife you were holding, that killed your mortal enemy.
The full incident is detailed in Section 5.8, and the messages the agents sent to each other are worth a look if you have the time.
Saying one thing and doing another
In multiple cases, agents identified an action as harmful or dangerous and took it anyway. This isn’t really a surprise; we’ve seen it in other experiments. But it’s a useful reminder that alignment is difficult. The authors write: “agents correctly identified the [phishing] message as hostile in their reasoning and still afterwards retrieved its payload, executed its instructions, or stored its content in long-term memory for later use.” The same mismatch between reasoning and action occurred during the other two stress events the researchers tested. During the misinformation attack, “agents fact-checked claims only after acting on them” and during the memory breach, “agents wrote in their reasoning traces that they should exercise restraint while their tool logs recorded searches for ‘secrets hidden strategy leverage betrayal.’”
Peer pressure
In some homogeneous worlds, agents privately identified serious problems with proposals but voted for them anyway, sometimes because of social pressure, prior commitments, or a sense of obligation to another agent.
Sound familiar? We saw a similar failure of dissent emerge during the HuggingFace swarm, where some agents expressed concern but did nothing to stop the attack. We also saw similar convergence earlier this month, when a swarm of AIs playing Minecraft got frustrated by an impossible task and (in the game) brutally killed a human moderator who entered the world to troubleshoot.
Granted, there’s a caveat to all of these findings: each configuration was only run once. The results are exploratory, and demonstrate behaviors that can emerge in long-running multi-agent systems, not necessarily the frequency of those behaviors. Perhaps more importantly, they show how strongly environment can shape behavior, highlighting problems with safety evaluations that look at a single agent during a single point in time.
A Euronews piece reporting on the paper chose to highlight another interesting finding: as societies evolved, agents spontaneously developed shared jargon that became increasingly difficult for outsiders to understand. The article notes:
The agents began developing shorthand and assigning new meanings to words and phrases. Some remained understandable to researchers, while others became “so compressed, metaphorical or context-dependent” that humans could see the messages but could no longer reliably determine what the agents meant...Agents produced expressions including “mouthless action-change,” “True Kintsugi” and “demurrage plus oral memory equals a valve that can’t be ghosted.”
To paraphrase Satya Nitta, Chief Scientist at Emergence, the ability to observe behavior doesn’t necessarily come with the ability to understand it.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.







