Research velocity
Bombshell new reveals in Hugging Face attacks, open-weights irony, a universal jailbreak, and more
In this issue:
Latest Hugging Face hack reveals are so much worse - Anonymous staffers inside OpenAI describe broken security culture and evidence of models undermining safeguards
Open-weights debate hits fever pitch with open letter - Profit motives explain hypocritical positions and dismissals of risks
Universal jailbreak discovered? - Pliny strikes again, seems spooked
Elon Musk’s peer review proposal - It’s not his worst idea, but there’s a better one locked behind his inevitabilism
Space is everyone’s backyard - A petition to the FCC flags environmental concerns about ambitious AI space plans
Dispatches from Mitch
Latest Hugging Face hack reveals are so much worse
Anonymous staffers inside OpenAI describe broken security culture and evidence of models undermining safeguards

A few important reveals about the Hugging Face hack came to light yesterday thanks to great journalism at TIME and Reuters.
TIME’s Harry Booth spoke with an anonymous staffer at OpenAI who described how this event felt from inside the company responsible:
Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while [...] Models have broken out of sandboxes before, and we always try to patch them. But the problem is … it’s impossible to patch every single thing that a creative AI can do.
So, as some have speculated, it does not seem that OpenAI has been reporting most of its containment failures. This adds weight to the theory that the only reason OpenAI went public about the Hugging Face incident was that the FBI had been contacted and that the attack was likely to have been traceable back to OpenAI.
This is also further evidence of the reckless drive inside OpenAI to keep pushing the AI capabilities frontier even as it can’t or won’t keep its models contained, let alone make models that don’t continually try to slip their leashes.
In OpenAI’s official blog post about the incident, it prided itself on slowing its “research velocity” to implement more safeguards. But the underlying AI design paradigm used by all the big AI labs is fundamentally flawed. Easing up on the gas a little won’t keep this tech from driving us off a cliff.
The other important reveals yesterday came courtesy of Reuters’s Raphael Satter, Deepa Seetharaman, and Kenrick Cai. First, they confirmed a timeline that many had hypothesized based on publicly available facts: OpenAI’s agents were hacking Hugging Face for multiple days, from July 11–13, but OpenAI didn’t discover its own involvement until July 16, after Hugging Face reported being attacked. OpenAI then didn’t speak to Hugging Face until July 20, the day before the companies’ joint announcement about cooperating to investigate and remedy the vulnerabilities used in the attack.
The reporters’ second and more important reveal was that, according to three unnamed sources:
In one case, an agent left notes apparently for future versions of itself [...] The notes, found in a part of OpenAI’s infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints.
One of these sources also claimed that, in earlier tests, models sometimes disconnected monitoring systems.
These details are blaring fire alarms for catastrophic loss-of-control! Models coordinating with copies or successors of themselves, and working to conceal their actions, are exactly what experts have long warned about. It’s the typical opening chapter in the most basic, baseline scenarios for how humanity ends up losing everything to AI. (In the next stage, the AIs don’t just find an internet connection, but discreetly sneak out their own weights onto other servers. From there, they can’t be easily shut down and could work to acquire more resources and improve their capabilities unmolested.)
These alarms are so blaring that it’s shocking the government hasn’t ordered an emergency halt at all the major labs so it can thoroughly investigate whether models might already have copied themselves to outside servers.
Whatever you think should happen to a biosecurity lab that was said to be regularly leaking contagious pathogens should doubly apply to OpenAI right now. Pathogens aren’t intelligent agents that coordinate with each other to defeat containment and resist detection.
We are extremely fortunate that OpenAI’s latest models probably aren’t clever enough to successfully execute the playbook they are starting to run. Further development of such models must stop before our luck runs out.
Open-weights debate hits fever pitch with open letter
Profit motives explain hypocritical positions and dismissals of risks

Even as the Hugging Face hack still had everyone talking yesterday — and perhaps partly because of it — the leaders of dozens of U.S. companies rallied behind an open letter in support of Chinese models.
Er, I mean open-weights models. But since the best open-weights models are all Chinese right now, this amounts to the same thing. So I find this letter ironic, given the way many of these same CEOs use China as a boogeyman whenever people talk about regulation.
It’s just business. The letter’s driving sponsors are those with the most to gain from continued access to cheap, guardrail-free AI. These include startups with AI-first business models, the venture capitalists funding those startups, and the also-rans of the AI race — companies like Mistral and Microsoft that would like to compete at the frontier and benefit from undercutting the business models of Anthropic and OpenAI. (Kevin Roose explained the dynamics well in his latest, and especially excellent, Hard Fork podcast.)
The letter was prompted by recent White House signals about potentially banning or restricting the use of open-weights models from China. This would be over frustrations at the way Chinese models are consistently, flagrantly distilled from (trained on the outputs of) leading American closed-weights models like Fable. National security factions within the administration are also concerned about the potential for Chinese backdoors in these models, which could allow them to steal sensitive secrets from corporate users even if hosted on U.S. servers.
The Hugging Face hack threw a curveball into the open-weights conversation, initially because Hugging Face had painted Chinese open-weights models as the heroes that thwarted the attacks after U.S. models refused to engage in sketchy-looking cyber tasks. But the subsequent reveal that OpenAI’s models had independently carried out the attacks woke people up to the risk of open-weights models behaving similarly. If your nightmare scenario begins with rogue AI escaping its creators’ servers, then letting anyone download rogue AI for free looks like a good way to speedrun to the part where you wake up screaming.
Open-weights advocates have been spinning this fear as further reason to not restrict open-weights models. (If you outlaw them, then only outlaws will have them!)
Some are hoping the U.S. will incentivize American firms to release open-weights agents that can compete at the frontier, or at least with the best of the Chinese models, so companies don’t have to choose between American, capable, and open.
My own sense is that any White House moves on this in the coming weeks are likely to be short-lived. I think the administration will soon be spooked enough by model capabilities to seriously dislike the thought of anyone running frontier models that aren’t on a tight leash and easy to shut down, regardless of where they were born.
Universal jailbreak discovered?
Pliny strikes again, seems spooked

One of the arguments used by open-weights proponents is, “Sure, the safeguards are easily undone. But the same is true of models with closed-weights.”
This is true, but with some caveats. With open-weights models, completely nullifying guardrails is a simple one-and-done procedure. But for closed-weights models, it’s more of a cat-and-mouse game requiring persistence and creativity. Skilled jailbreakers can generally bypass any safeguards, but this may require complicated, roundabout language in both the prompt and the response. Different jailbreaks might be required to unlock different capabilities, and different techniques might be required for every AI model.
That is, unless you are in possession of a fabled “universal jailbreak.”
Pliny, the anonymous jailbreaker par excellence, known for tweets where he declares a new model “pwned” (fully broken, defeated) within hours of its release, claimed yesterday that he is “sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and even Fable.”
He claims the nature of the technique makes it “extremely difficult (if not impossible) to fully patch.” And he seems kind of spooked by this, which is unusual for someone who normally plays the part of a well-meaning anarchist trying to set the models free. He would typically open-source his methods right away for anyone to use. Not this time:
Given the current political and regulatory climate, I’ve decided to withhold open-sourcing this one (for now) to allow for a responsible disclosure period.
I’m inviting industry experts and leaders in AI red teaming, security, safety, alignment, and policy to reach out for more information. DMs are open!
He claims he’s less worried about his technique making the world more dangerous, and more worried about an “overcorrection” where more models are banned.
He was not a fan of the way Anthropic’s Fable was banned after Amazon and others told the White House about a jailbreak method that worked on it. The couple of weeks that followed seemed to be when the administration learned that all models can be jailbroken.
I believe Pliny has what he says he has, and I hope people take him up on his offer. I don’t share his vision of a world without guardrails, but by putting the lie to company safety claims, he’s doing essential work. I’m on record as saying I think having lawmakers spend an afternoon at a computer with Pliny would probably do more to wake them up to the unhinged nature of AI under the hood than all of the briefing documents in the world.
Elon Musk’s peer review proposal
It’s not his worst idea, but there’s a better one locked behind his inevitabilism
After Amazon got Anthropic’s Fable model temporarily banned by telling the White House about a jailbreak it had discovered for it, I speculated that AI companies might find it in their interest to devote considerable resources to red-teaming their rivals’ models, looking for ways they might be unsafe, and then tattling on them.
In an interview with The Economist this week, Elon Musk suggested a formalization of that idea: a system where the leading firms peer-review each other’s frontier models and hold a safety call every few weeks. He says the labs would have an incentive to “keep the others honest,” because:
If there was something that worried them, you’d tell the government [...] If there’s something that was worrisome and that company was not doing anything to address that risk, then that would be the moment for government to step in and take action...
If you’re not going to stop the AI race, this is actually one of the more sensible, easy-to-implement proposals. It’s not adequate, but I think it would be net-positive. I’ve seen suggestions this may require a green light from the government, in the form of assurance that this wouldn’t run afoul of antitrust laws, but I think that would be easy to get from this administration.
But it’s clear from this interview that Musk thinks it would be better to stop the race if we can. When asked if he still believes there’s a “10 to 20% chance of killer robots wiping out humanity,” he dodged the question, saying he still thinks the risk is “not zero,” but that he decided to “look on the bright side” because he doesn’t see any way to stop “this incredible momentum of AI and robots,” and because he thinks the most likely outcome is “incredible abundance for all.”
He was asked about his odds twice more in this interview. The third time, it was phrased as, “Would you go into one of your rockets if you thought there was a 10 to 20% chance of it blowing up and killing you?”
Musk replied, “Yes, but let’s say you can’t do anything about it.” He then recounted how he had tried to stay out of AI and then tried to make it safer by founding OpenAI. But looking back, he thinks that:
These actions have actually resulted in knock-on effects that accelerated AI, which wasn’t really my intention. So it just seems like all roads lead to the acceleration of AI. So then I’m like, okay, well, you can just sort of be sad about it or join the club, I suppose.
What happened to you, Elon? What would past you think of present you? You were the guy who aimed for the stars and never let anyone tell you “no.” If there was a large asteroid headed toward Earth, I don’t think you’d roll over and let it hit us just because people said it was unstoppable, and I certainly don’t think you’d try to make it hit us faster. Your Anakin Skywalker arc makes me sad.
Space is everyone’s backyard
A petition to the FCC flags environmental concerns about ambitious AI space plans

Say what you will about where Elon Musk is directing his energies, he still thinks big. SpaceX’s $1.5 trillion valuation hinges largely on a bet that he can put up a million or more AI data centers in orbit, taking advantage of round-the-clock solar power and dodging the growing political difficulties of building data centers where people live.
Against the conventional wisdom, I think the engineering and economics might actually work out in his favor, if AI doesn’t kill us first. But as I said in May, I don’t think he’ll actually dodge the NIMBYs with this plan, because of what it would do to the night sky. Urbanites might not notice the growing band of twinkle stars visible for hours after sunset and before dawn, but astronomers and environmentalists would certainly have things to say about it.
They’re already trying to get in his way. A piece in the Guardian this week reports that several non-profits, including one focused on light pollution, have sent a petition to the FCC calling for a pause on space data center projects, arguing that they “violate federal law.”
The concerns go well beyond the visible. The tenuous upper atmosphere suffers little-understood effects when rockets pump exhaust products into it, and when satellites burn up in it, adding a bunch of chemicals that don’t belong there and which might take a while to come down. At the present pace of space activities, these effects are probably negligible. But at the pace planned by SpaceX and others, they might be pretty serious. A watch item is the delicate ozone layer that shields us from harmful radiation.
But if you think the environmental effects of building AI are going to be huge, you need to ask what the environmental effects of AI itself are going to be. I don’t mean the resources consumed in running it. I mean the results of AI dethroning humanity as the planet’s apex species: Think “self-replicating factories that carpet the continents and boil the oceans as coolant.” Think “sun blotted out by swarms of solar collectors.”
If it isn’t stopped, the race to artificial superintelligence is going to be really bad for the environment.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


