To deliberately pace
Employee slow-down petition, Amodei vs. open-weights letter, Altman "singularity" claim, and more
In this issue:
A petition to “pace the frontier” - Employees of frontier AI companies want to be able to trust a mutual slow-down
Amodei declines to flinch - Anthropic’s CEO shared reasons he won’t sign the open-weights letter despite intense peer pressure
Best-selling book might be 60% AI - As AI writing gets better, will taboos against it persist?
The EU AI Act is about to become more powerful - Europe doesn’t need to make its own models for regulatory leverage
Systems that we can’t aim and we can’t inspect - All the brouhaha about Sam Altman’s “singularity” aside is missing the point
Dispatches from Mitch
A petition to “pace the frontier”
Employees of frontier AI companies want to be able to trust a mutual slow-down

More than a thousand employees of frontier AI companies have signed on to a short statement that warns of the potential for AI development to rapidly accelerate beyond “our ability to understand or control.” Signatories to the new statement include four of Anthropic’s co-founders, OpenAI’s chief scientist, OpenAI’s chief research officer, Meta’s chief scientist, and Google’s VP of AI Safety & Alignment.
The statement makes the following call to action:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.
The “governance tools” in question are those mechanisms that would be needed for parties to reassure each other that they are not secretly racing to more powerful models while outwardly agreeing to slow down.
Some of those tools could be technical, like tamper-resistant chips that would be very difficult to use in unauthorized ways or in unauthorized locations without leaving fingerprints to that effect. Others would be more regulatory and political, like tighter monitoring of the AI chip supply chain. A mix of many kinds of tools will likely be needed.
I, for one, am highly encouraged by the new statement. A call for developing governance tools may seem several steps removed from what ultimately needs to happen, but it is a necessary step that cuts to the heart of the race dynamics that have people feeling trapped. And I think many policymakers are going to be surprised by how much groundwork for governance already exists: MIRI’s Technical Governance Team has been cataloguing the possible tools and exploring how best to arrange them to prevent the premature creation of artificial superintelligence.
This makes me optimistic that governments will soon move ahead with existing tools so that they will be prepared to move quickly on a slowdown or pause, once satisfied that any missing tools can be ready by the time any international agreement is finalized.
The statement’s web page includes quotes provided by signatories. Leo Gao, a member of the technical staff at OpenAI, writes:
The world is locked in a deadly race towards an intelligence explosion, where AI’s ability to create better AIs reaches a critical point, just like a runaway nuclear chain reaction. Going slower would give us much-needed time to make it go well, but no individual actor is willing to stop unilaterally. To survive, we must coordinate to slow down the race.
Amodei declines to flinch
Anthropic’s CEO shared reasons he won’t sign the open-weights letter despite intense peer pressure

Among tech leaders and experts over the past several days, the open letter against bans on open-weights models from China or elsewhere fell into the sort of cascading peer pressure I associate with episodes of social injustice: Anyone who dislikes the injustice but has reservations about proposed remedies can find themselves cornered by claims of “You’re either with us or you support the injustice.”
In part, the letter was a transparent move by underdogs to form a disapproving circle around the leading AI companies, painting them as greedy totalitarians for trying to keep their frontier models locked down. It seemed to say, “You’re either with us or you’re against open weights, open source, and freedom itself.” And it worked: Even Google and OpenAI ended up signing the letter.
That left all eyes on Anthropic, the only big dog still in the center of the circle. It would have been easy for CEO Dario Amodei to cynically sign the letter with no intention of doing anything differently, as I fully suspect OpenAI’s Sam Altman did. Instead, Amodei put out a blog post yesterday explaining why he wasn’t going to sign.
He clarifies that “Anthropic has never advocated for a ban on open-weights models,” and that open-weights models without dangerous capabilities are a “public good.” But he’s worried that powerful AI models “may be misused to carry out cyberattacks or biological attacks, and may have serious alignment problems,” linking to the TIME article about the Hugging Face attack. Open-weights models, he reminds readers, are hard to monitor, impossible to take back, and trivially freed from guardrails against misuse.
Amodei says we should instead:
Stop selling powerful AI chips to China.
Crack down on the distillation operations used by Chinese companies (and others) to transplant frontier capabilities into models that may lack frontier safeguards.
Subject all models — open and closed — to mandatory safety testing.
He disputes the open letter’s thesis about defensive AI being the solution to offensive AI, writing:
I worry that biology will have a strong attacker-defender asymmetry, where sufficiently capable models may be able to quickly weaponize pandemic-level viruses with widely available materials, whereas defense against these agents is a multi-year operational task in the best case.
As I’ve just shared a bunch of Amodei’s words at face value, let me be clear: Amodei is leading a race to superintelligence that would kill us all. The race must be stopped.
I also think Amodei is wrong to imply that it’s safe to continue developing frontier AI, as long as it’s thoroughly tested; the most dangerous models are already criminally dangerous and prone to escaping during testing — as seen in the Hugging Face attack.
But I respect Amodei for standing up to the open-weights open letter posse, and I think his stated reasons for doing so are sound. As bad as things are, they could always be worse. Could we please not give everyone with a GPU and a grudge the means to design a pandemic?
Dispatches from Alana
Best-selling book might be 60% AI
As AI writing gets better, will taboos against it persist?

People usually think it’s fine to use AI to code something, design a website, or even create a slick slide deck. But AI for writing? That’s a big no-no, according to many.
Will it stay that way? A recent article in The Atlantic covers the highly praised book Daggermouth, a dystopian romance about a female mercenary who has to marry the person she was supposed to kill. Simon & Schuster acquired the book in a seven-figure deal, and it “has spent months on USA Today’s best-seller list.” The article describes it as a “viral hit.”
The catch? AI accusations. In a not-yet peer-reviewed study that ran 14,000+ randomly-selected e-books through the well-regarded AI detection tool Pangram, Daggermouth received a score of 60% AI generated.
Important caveats: The author denies the allegations, Pangram isn’t perfect (though it’s pretty good), and the score lumps together “AI-generated” with “moderately AI-assisted,” the second of which carries less of a taboo.
But it seems pretty clear that the book is not entirely human written. According to a University of Maryland professor unaffiliated with the study, a human-written book returning a 60% Pangram score is “almost statistically impossible.” When the Atlantic reporter ran Chapter 18 through Pangram, it “classified the first few paragraphs as human-written and the rest of the chapter as 100 percent AI.” Finally, the study authors did a second check: looking for “rare expressions” that are common in other texts suspected of AI generation but don’t appear in human writing. They found a bunch. One example is a sentence that appears both in Daggermouth and in an e-book from a different author: “They collapsed together, a tangle of sweat-slicked limbs and racing hearts.” Ew.
Nevertheless, the article is careful to point out that this book seems different from “AI slop.” First, people genuinely like it. Second, Wolfe isn’t churning out tons of content; she’s only written two books and the sequel to Daggermouth is not expected until next year, “an indication that Wolfe’s process likely involved significant amounts of genuine human craft.”
Without having read the novel, I would tend to agree. If Wolfe is using AI to provide feedback on her work, and it’s giving her some rewrites that it thinks are better, is using those all that different than taking rewrites from an editor?
It’s fairly uncontroversial that writers shouldn’t use AI to generate writing from scratch and then slap their names on it. But I don’t fully understand the taboo behind using AI as an editor or thought partner, especially when it’s not a taboo that applies to other fields. The main reason I’d discourage using AI for edits has more to do with quality. You can’t tell, from the outside, whether something is coded with AI. But at least with today’s tools, you can often tell when something is written by AI, and the style can be grating. Another issue: writers, who are intimately familiar with their own ideas, often overlook missing context because they fill in the gaps themselves. This can make it hard to notice when AI subtly changes their intended meaning, which it seems to do often.
Of course, the signs of AI writing are only visible to people who have become familiar with them. Many Daggermouth fans certainly weren’t bothered. And given how fast AI capabilities are increasing, it seems likely AI writing will get both better and less tell-tale. I don’t think the taboo against AI writing will go away anytime soon, but it might become so unverifiable as to cease being very discussion-worthy. How will writers respond?
The EU AI Act is about to become more powerful
Europe doesn’t need to make its own models for regulatory leverage

I’ve heard a lot of talk, including from people I ordinarily respect, about Europe’s place in the geopolitics of AI. “They don’t have their own models, how will they have any leverage?”
I’d like to offer a counterpoint to this. As covered by a recent Politico article, Europe is ahead in the AI regulation game, and it’s about to exercise significant enforcement power over US and Chinese AI companies, despite not having leading models of its own.
The reason? Europe is a huge market and companies won’t want to lose it. But to get access to European customers, non-European companies need to comply with the EU’s AI Act, which requires providers of the most advanced AI models to “assess and mitigate possible systemic risks.” The European Commission has identified four of these: bioattacks, loss of control, cyberattacks, and manipulation at scale. Its AI Act has been in effect for about a year, but as of Sunday, the enforcement powers kick in. Europe’s AI Office will have, at least on paper, the power to “monitor and supervise” how well AI companies are handling these risks, which includes requesting that companies submit models for evaluation.
Politico is skeptical companies will comply, repeating a familiar talking point:
Will the U.S.-based frontier models open up their closed models to EU regulators? And how will Brussels deal with the rise of Chinese open-source alternatives? These are both regulatory unknowns, and yet another reminder of Europe trailing on the development front.
The implication is that if Europe regulates too hard, the AI companies will decide the market isn’t worth it. Without its own models, Europe will be forced to lighten things up, lest it lose access. But I think the article misses an important point: competition between China and the US will likely keep the European market extremely relevant. And with no direct competitor of its own, Europe can regulate more freely. In contrast, the US has been hesitant to regulate without China doing the same.
In light of the recent Hugging Face attack, a thankfully low-stakes instance of two major risks that have the potential to be incredibly high stakes (loss of control and cyberattacks), an increase in European regulatory power is pretty timely. Sure, AI companies might choose to pay the fines instead of complying. But if Europe plays its cards right, the AI Act could significantly improve the current state of affairs. Adding some much-needed regulation won’t be enough on its own, but it’s certainly welcome.
Dispatch from Donald
Systems that we can’t aim and we can’t inspect
All the brouhaha about Sam Altman’s “singularity” aside is missing the point
“Singularity” is a heavy word. Even just in the context of artificial intelligence, it’s been used in a few different ways — artificial intelligence designing a more intelligent version of itself, or technological improvement feeding faster technological improvement — but the underlying theme is an acceleration of exponential growth that outpaces our ability to keep up with it. Like crossing the event horizon of a black hole, it’s hard to see what awaits you on the other side.
The AI self-improvement singularity is particularly important. Beyond it, human beings are no longer driving AI progress, and no longer able to reliably predict, understand, or control what the machines do next. So, it’s no wonder then that people sat up and took notice when Sam Altman said, “We are now, like, in the singularity.” Another point worth mentioning: He admitted that there are “still major AI alignment issues to solve and safety issues.” This was just a brief aside, a few sentences in a conversation almost exclusively about startups, but Altman’s comments have set off a round of arguments about the wrong thing.
Forbes’ Lance Eliot, for example, makes the case that:
[Altman] is not saying we have crossed the Rubicon and attained the AI singularity. I believe he is emphasizing that there isn’t likely a tipping point; there isn’t a single moment in time. It is a curve we are climbing incrementally.
I think this is a fair interpretation, if not of Altman then at least of the situation. You can’t take a look at some graphs, punch a few numbers into your calculator, and say, “We have now entered the AI singularity,” precisely because there are multiple competing and fuzzy definitions. Altman himself says as much not long after: It is “all one crazy exponential,” and no single moment is the tipping point.
So let’s stick to the facts: We are in danger of losing control.
Earlier this month, the machine learning company Hugging Face suffered a severe cyberattack. The responsible party was not a state actor or a criminal syndicate. It was an OpenAI agent that had escaped its testing sandbox, obtained internet access, and broken into Hugging Face’s servers by exploiting a chain of security vulnerabilities. That agent was not supposed to do any of those things. (“...still major AI alignment issues to solve and safety issues,” indeed!)
In his podcast interview about startups, Altman said that “we are close to creating a genie that can grant any wish.” I feel that this statement is inaccurate, to put it mildly. Our current AI models aren’t even trickster genies who will grant your wish in a twisted way that’s guaranteed to make you suffer. For all its malice, the trickster genie has an important property: It still at least executes the wish. The problem is in the wording, which means that careful wording is the solution. But not even the wiliest contract lawyer could solve the problem that we actually face.
Nobody wished for OpenAI’s model to break out of its sandbox, or access the internet, or hack another company’s servers. But the AI model wasn’t just doing its level best to complete the task it was given. AI models are trained to go hard at whatever they’re pointed at. In this case, that was a test score, and not the thing that the test was for. OpenAI wanted to measure capabilities; the model wanted a perfect score. These are not the same thing, and we can be pretty sure that the model didn’t think they were the same thing.
We can’t know for certain what actually motivates these models under the hood. We can watch what they do and draw inferences about what drives them, but we cannot guarantee that they truly internalized whatever it is that we trained them on.
Writing in The Guardian, Bruce Schneier and Barath Raghavan propose a score they call the “Genie coefficient,” to measure the distance between what you asked the AI model to do and what it did. Scoring means putting an AI model in a walled-off sandbox, with real tools it could misuse, and grading its worst behavior, not its typical behavior. (That last part is a good move that I really appreciate.) I still think the Genie coefficient has a very big flaw. Like the AI Kill Switch Act, it doesn’t prevent danger. The Hugging Face hack was performed by an AI agent that broke out of a testing sandbox; I’m not enthusiastic about stress-testing how disobedient an advanced AI system can be under the assumption that we definitely, for sure, have locked all the doors and windows this time.
Right now, the safest thing to do is to stop building these things. It’s good to have a law about shutting down an AI model that acts dangerously, and it’s good to have benchmarks to weigh how uncontrollable an AI model is. There are ten thousand good things worth doing. But Altman is right that there’s no clear tipping point, no single moment to wait and watch for, just a steepening curve. Altman calls it a “glide path,” a course that leads smoothly and inexorably to some particular destination.
If you find yourself on a glide path and you’re beginning to think that the destination might not be so nice, then the best thing to do is to steer to some different path.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.





