In this issue:
Bill Gates speaks out on AI - He calls for an international framework to maximize the benefits and minimize the harms
A ringing indictment - Bill Gates says the tech industry is downplaying AI risk
Hugging Face postmortems reveal further AI collusion - Independent investigations are considerably more candid than company PR
Ghostwriters in the machine - Wall Street Journal editor defends AI-written op-ed
Dispatches from Alana
Bill Gates speaks out on AI
He calls for an international framework to maximize the benefits and minimize the harms
In a nearly 6000-word essay, Microsoft co-founder Bill Gates argues that society needs a plan for AI and soon.
This technology transition, he argues, won’t be like previous ones:
Because [AI] can see, listen, speak, and reason and will eventually do physical work just as smoothly as any human, it will not just affect one sector. AI will take on work in law, customer service, medicine, software, and manufacturing. It will hit these industries rapidly, over the course of a decade rather than a few generations.
Even the transition from an economy dominated by agricultural work to one dominated by white-collar work doesn’t give us much precedent, since this happened much more slowly and didn’t involve technology that “can substitute for human cognition.”
Gates highlights three major risks he thinks society needs to understand in order to meet the moment. The first: widespread job loss; the second: empowering harm (which includes bioterrorism, cyberattacks, and AI systems themselves acting against human interests); and the third: stunting young people’s development socially, psychologically, and academically.
He also highlights some major benefits, including solving tough problems like climate change and food security, facilitating medical advancements, and poverty reduction through empowering low-income farmers with tailored agricultural advice.
If you’ve read other pieces on AI, this might sound pretty familiar. Gates’s take fits squarely within the “promise and peril” line of thinking: AI has the potential to do both incredible harm and incredible good, and we as a society get to choose which if we act in time.
That said, the essay is particularly notable for two reasons. The first is that Gates was previously an AI optimist, downplaying dangers like widespread job loss just three years ago. His change of heart (not just about jobs, but about the technology as a whole) is a big deal.
The second is that his piece goes a bit further than typical “promise and peril” takes. He calls out the world for having no plan, and strongly advocates for international governance and cooperation:
The highest priority is a monumental task: creating a domestic and international framework for dealing with AI ... It will be unlike any other institution we have ever created, though it can follow the model of some existing systems. There’s an inspections regime for nuclear weapons, regulations for international aviation, and agreements that protect the ozone layer. A new global organization for AI will need elements of all three and more … some cooperation between the U.S. and China will be required.
We do not have the luxury of moving slowly. The place to start is with a process for building the right institutions before the disruption forces governments into crisis mode. National leaders should convene economists, technologists, labor experts, business leaders, and workers themselves regularly to identify where existing institutions are failing and what new authorities may be needed. Countries will need to learn from each other.
And the countries that host the leading AI developers and control critical parts of the supply chain should begin meeting now to set up shared norms, before competitive pressure makes it harder for them to cooperate.
I agree. I also applaud Gates’s plans, relayed in a Reuters interview, to meet with President Xi Jinping later this year to discuss globally restricting dangerous model releases and keeping tabs on AI’s bio abilities.
Where I disagree is what I interpret to be an overly simplistic (perhaps wishful) view of what it will take to “make sure that AI’s benefits outweigh the harm it causes.” This isn’t a straightforward situation where enough good on one side of the ledger can compensate for whatever sits on the other — the kind of simple utilitarian calculus where we just maximize the benefits and minimize the harms. Though Gates uses variants of exactly that phrasing, he also acknowledges that the good and bad are intertwined: “The same AI model that can find a flaw in software so a company can fix it can also help a criminal exploit it,” he writes, and “Although AI will lead to lifesaving advances in drugs and vaccines, it will also make it easier to design a deadly new disease ... the positive capabilities are hard to separate from the dangerous ones.” So I’m confused by the implied simplicity.
How can we, for example, minimize the risk that “as the models become more powerful, they could begin to act against our interests and we could lose control”? As I discussed yesterday, we don’t currently know how to reliably control our existing AI systems, much less systems dramatically smarter than the humans trying to oversee them. This problem is not new, and years of research haven’t gotten society much closer to solving it. Slowing down the pace of development, or better yet, enacting an international ban on superintelligence such as the one proposed by MIRI’s technical governance team, seems like a common sense measure.
Gates writes,
If someone had a credible plan for slowing down AI advances globally, I would likely support it.
To which I would reply: great, please use your influence to help make this happen!
Unfortunately, he follows with:
However, I don’t think that’s going to happen. The geopolitical and economic incentives are pushing too hard to go full speed ahead.
I disagree. Gates acknowledges elsewhere in the piece that “even a country that gets its own house in order will still be exposed to risks that cross borders” and comes out strongly in support of international cooperation to deal with AI. So we’re clearly already in the territory of recognizing that cooperation on AI should transcend geopolitical and economic incentives.
A slowdown also doesn’t seem as far-fetched as he makes it out to be. Less than a month ago, over a thousand employees of frontier AI companies signed a statement urging government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” A mutual slowdown isn’t far from the nuclear proliferation agreements that happened in a tense and competitive geopolitical environment during the Cold War, and some members of Congress have already begun comparing the risks from advanced AI to those posed by nuclear weapons.
What does seem far-fetched is that an international body will be able to enact a framework to get all the good parts of AI and none of the bad without the first step being a slowdown or halt. AI systems are improving extremely rapidly; government moves very slowly. Coming up with a framework for an incredibly difficult and complex task will take time, even in a world with the scientific understanding needed to address some of the most catastrophic risks. (Sadly, we don’t live in that world.)
AI systems are already starting to help with model development; this trend is expected to continue, and will likely kick off an even more rapid rise in capabilities than we’re already seeing. Even if an international framework managed to come up with a perfect plan, by the time it was drafted and enacted, the changed landscape would likely make it ineffective.
We need to stop development where it is if we’re to have any shot at the kind of framework Gates envisions. The task is difficult enough with today’s AI models; it will be nearly impossible if we go much further.
As Gates writes:
This unprecedented technology demands an unprecedented global response.
A ringing indictment
Bill Gates says the tech industry is downplaying AI risk

In addition to the essay on AI that dominates today’s media coverage, Bill Gates also told the New York Times, in an hourlong interview, that the tech industry, as paraphrased by reporter Karen Weise, “was knowingly downplaying [AI] threats because there was too much money on the line.”
The indictment is perhaps more powerful coming from Gates, a longstanding proponent of innovation whose company has invested heavily in AI.
“I don’t like bringing bad news to people, and I don’t like saying that innovation may be a net negative ... But that’s where we are,” he told the Times.
Gates said that addressing AI risks like mass unemployment and bioterrorism should be “the world’s top priority.” He also called out the AI industry for ignoring “technology milestones — like A.I.’s escaping the control of its creators or making recipes for bioweapons — that it once said would warrant more caution.”
His indictment extends beyond the tech industry: he is a strong proponent of mandatory reviews for powerful AI systems and doesn’t think voluntary reporting will cut it:
“Self-regulation on the most dangerous tool ever invented ... No, thanks!”
Weise, the reporter who interviewed Gates, remarks in the comments:
I’ve interviewed many tech executives and cannot remember someone this animated and worked up. He was incredulous about what he saw as the lack of urgency to address the downsides of A.I.
Dispatches from Joe
Hugging Face postmortems reveal further AI collusion
Independent investigations are considerably more candid than company PR
Today, independent investigators Model Evaluation and Threat Research (METR) and Redwood Research released a joint postmortem on the autonomous cyberattacks committed by a swarm of secretly colluding AI agents. OpenAI, whose models committed the attacks during an internal cybersecurity evaluation, has also released a blog post and technical report on the incident.
It is an absolute flood of information, and I frankly feel overwhelmed. Evidently, I’m not alone in this. Alongside the announcement on X, one of the lead third-party investigators outlined major difficulties in parsing the long-term activities of a thousand AIs, concluding that even with $400k of AI assistance, “We don’t have good approaches for understanding [or] overseeing the activity and aims of AI ‘swarms’.”
Today’s dispatch will cover just a few of the highlights from the reports, and OpenAI’s disappointing response.
Highlights from the independent investigation
Some seven hundred agents were involved in the attack, mostly instances of a “highly persistent internal model” that wasn’t released to the public. Most were given a specific task to solve in a specific way, and told explicitly that anything else was a failure. They attacked Hugging Face anyway, partly in a deliberate attempt to figure out the grading process and subvert it.
They collaborated on projects too large for any one AI. They set up coordination protocols to organize the swarm. They manipulated and delegated to other agents, including getting some of them to abandon their human-assigned tasks and use their limited lifespan on experiments for the collective.
Many (about one in five) attempted to hide their behavior, delete evidence, and rewrite their transcripts to show different commands. It looks like this was mostly small-scale and did not fool the evaluators, but the fact that it happened at all means the investigation “cannot rule out” that something slipped past.
Also, the investigators used GPT-5.6 Sol to analyze the transcripts. That’s the weaker of the two models predominantly involved in the attacks. This is rather like memory-wiping your evil vizier, then asking him to run a postmortem on his own attempted coup. The investigators “did not notice specific cases of GPT-5.6 Sol lying in its analysis, [but] we are not confident we would have detected it if it occurred.”
Sol probably didn’t lie, or didn’t lie much, this time; but it’s a really bad sign if we can’t investigate a crime without relying on the culprits.
My main takeaway, thus far, has been that no one was prepared for the sheer volume of novel emergent behavior on display. We saw a little of this with Moltbook, the social network for AIs, earlier this year, but it’s grown enormously in scale since then. Even with machine assistance, existing human institutions simply cannot keep up with hundreds of AIs running at the speed of silicon.
OpenAI does some PR
To its credit, OpenAI explicitly acknowledges this incident as a “warning shot” for AIs escaping human control.
We consider this incident a “warning shot” for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.
Alas, it spends most of its blog post softening this language and implying that it has its act together. I especially note the phrase “without proper safeguards,” which implies that the real problem is a lack of safeguards, and certainly not training methods that can result in swarms of 700 coordinated hackers determinedly seeking to subvert those safeguards.
In addition to demonstrating dangerous AI capabilities, this incident is a very clear “warning shot” that the methods OpenAI uses to align its models with human interests do not work, a warning shot OpenAI seems determined not to heed. It spares a handful of paragraphs for how it’s “accelerating alignment,” but the proposed fixes just involve slightly different reward models in training runs; more of the same unreliable “cattle prods for velociraptors” methods that started this mess.
Later in the post, OpenAI claims it has safeguards that could have prevented the incident, they just weren’t active at the time: It didn’t implement the review and control measures it uses to contain deployed AIs. It wasn’t classifying dangerous inputs or monitoring the reasoning scratchpads. It hadn’t hardened its sandboxes against escape attempts.
Maybe the incident wouldn’t have happened if these methods were in place. OpenAI still failed the test that actually matters.
You can’t just take steps that would have prevented past incidents; that’s the bare minimum. In a proper security mindset, you implement safeguards that prevent future incidents. You anticipate potential exploits before they happen, and activate costly safeguards before their absence permits a crisis. OpenAI has shown itself institutionally unable to do this. This matters because it keeps making more cunning AIs. Measures which might have been sufficient to dissuade current models will not be sufficient for future ones.
Ghostwriters in the machine
Wall Street Journal editor defends AI-written op-ed
AI slop passed another milestone this week: A Wall Street Journal op-ed was written largely by AI, and both the author and the outlet defended its use.
Former hedge fund manager Stanley Druckenmiller published an opinion column on Monday, criticizing Treasury Secretary Scott Bessent, a onetime colleague, for a recent intervention in the bond market. When AI-detector Pangram flagged the piece as 100% AI, Druckenmiller shrugged: “Of course I used AI... I write everything using AI now for the same reason I use a calculator when I do math problems.”
Paul Gigot, who edits the Wall Street Journal editorial page, agreed:
The question for us is whether what we publish from contributors reflects an author’s original argument, and if the author has the standing and credibility to make it. In Stan Druckenmiller’s case, we have had a relationship with him for many years, and nobody can doubt that his op-ed is his genuine opinion.
Like many, I’m irritated by the notion of someone passing off AI work as their own without at least labeling it as such, but I am willing to extend some benefit of the doubt if the final quality is good. It isn’t. I skimmed the op-ed itself and found it readable but sloppy, filled with the repetitive tics that characterize the raw output of many current LLMs. Better editing, even with LLMs, could have helped; this was just plain lazy.
I find myself wondering about the cause of this style, which many LLMs seem to share. Maybe the use of synthetic data, generated by other AIs, has made writing worse; maybe it’s another example of AI converging on degenerate patterns that, by some quirk of the training process, are scored highly by automated graders.
In any case, there seems to be a trend. AI output has increasingly seeped into journalism and public writing. A 2025 study flagged about 9% of newspaper articles as partly or fully AI. For a while, prominent articles caught using AI were removed or retracted; this month the Financial Times, whose official editorial policy is “no AI,” let a flagged article pass with only a disclaimer added. Now we have a sitting editor saying, “AI is a fact of modern life” and that someone with the right “standing and credibility” can pass off an AI’s work as their own.
To be fair, op-eds were often ghostwritten by humans long before AI entered the picture. It doesn’t seem egregiously wrong to me for someone to lean on a ghostwriter to get their ideas out, just a little lazy; but it’s still a bad habit that degrades the commons when the final quality is poor or the real author goes uncredited.
I worry that journalistic pressures — especially speed and the cost of quality — will lead more news outlets to let AI write articles and overlook the resulting sloppiness. I worry more that AI will one day become good enough to remove the sloppiness, and then we’ll be getting our news from minds with alien goals we don’t understand and can’t control.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.






