Incurable try-hards
Hugging Face Cassandras, misconceptions about the attack, and policymaker reactions
In this issue:
Announced disasters - The OpenAI incident shouldn’t have been a surprise — there were warning signs
You can do better than these five takes on the Hugging Face hack - I give a B+ to the media’s coverage of this event. Here’s how to get an ‘A’.
Policymakers react to the Hugging Face hack - A look at the AI Kill Switch Hack, and a review of bipartisanship
Dispatch from Robert
Announced disasters
The OpenAI incident shouldn’t have been a surprise — there were warning signs

Last year, in their book’s example scenario, Eliezer Yudkowsky and Nate Soares described a fictional AI named “Sable,” whose original task was to investigate a famous mathematical conjecture but which then broke out of its containment. In If Anyone Builds It, Everyone Dies, Yudkowsky and Soares made a strong case that the fictional AI in their scenario could have likely just hacked its way out of its isolated test environment.
However, they wanted to make their argument even stronger, which is why they wrote, “but suppose Sable does not have that ability,” and had the AI find another way to break out. They did this to win over skeptics who were certain that hacking wasn’t that easy.
That was ten months ago. This week alone, OpenAI has proven twice that it actually is that easy.
On Tuesday, I reported on an internal OpenAI model that refuted a famous math conjecture and soon after also solved the problem of how to break out of its sandbox and gain internet access. On Wednesday, my colleague Joe then reported on an even more serious incident in which an internally deployed OpenAI model (maybe the same one) also broke out of its isolated test environment during a cybersecurity evaluation. It gained access to the internet using previously unknown software exploits, and then launched an attack on another company’s servers to steal the test results of its evaluation.
However, Yudkowsky and Soares’ prescient scenario wasn’t the only warning sign. Just nine weeks before the two security incidents came to light, METR (Model Evaluation and Threat Research), an independent nonprofit research institute, published its Frontier Risk Report. In this report, METR — with support from leading AI labs who participated — investigated how internally deployed AI agents actually operate within the labs. In doing so, they also focused on the danger of “rogue deployments” — that is, AI agents operating in an environment where they are not supposed to be and performing unauthorized actions without their handlers’ knowledge. They reached a conclusion that reads like a blueprint for the incidents now being reported:
Overall, we believe that internal agents at the time of our assessment plausibly had the means, motive, and opportunity to initiate small-scale rogue deployments, but they did not have the means to make them highly robust. Given rapidly advancing capabilities, we expect the plausible robustness of rogue deployments to increase substantially in the coming months.
The warning could hardly have been clearer. No one should be very surprised at the events of this week. But frontier labs still aren’t adequately prepared. The Future of Life Institute, a nonprofit organization dedicated to AI safety, published its AI Safety Index in early July. In this report a panel of scientific reviewers evaluates nine AI companies based on their safety measures. In the “Existential Safety” category — which assesses companies’ preparedness to manage extreme risks from future AI systems — OpenAI received a D+. Grade D stands for: “Weak strategy; vague or incomplete plans for alignment and control; minimal evidence of technical rigor.” It was the second-best rating among all companies examined.
The recommendation that the independent review panel makes to OpenAI certainly hits the nail on the head:
Evaluate internal-deployment risks before broad internal use rather than after.
However, these recommendations are not binding, and the reviews themselves — conducted by independent research institutes and nonprofit organizations — are entirely voluntary.
It was foreseeable that such incidents would occur; it was only a matter of time. If this reckless race toward ever-smarter and increasingly uncontrollable AIs is not stopped, then this will end badly. AIs getting smarter and more strategic can also mean that they might not give us another warning shot before it’s too late.
Dispatch from Mitch
You can do better than these five takes on the Hugging Face hack
I give a B+ to the media’s coverage of this event. Here’s how to get an ‘A’.

The Hugging Face hack by internal OpenAI models, as well as the containment breach revealed earlier in the week, have deservedly captured the media’s attention. I’ve found most of this reporting to be above average in quality, acknowledging the seriousness of the attack and recognizing some of its implications.
But I’ve also seen plenty of articles make or share dubious claims and assumptions about the incident — some taken at face value from OpenAI itself. Without calling out anyone by name, below are five takes that could be improved upon.
1. “We should address this incident with better security, monitoring, and reporting.”
Those are all good things, but they imply moving forward with business as usual. This is insanity, when business as usual means “same basic design, only more powerful.”
When the catastrophic nuclear mishap at Chernobyl revealed previously unknown flaws in the inherently dangerous RBMK reactor design, tighter protocols around the remaining RBMKs in service were certainly justified. But you didn’t hear reactor companies saying, “and with these safeguards in place, we will proceed to building the super RBMK, a 10x larger model with world-beating outputs. Our ultra RBMK 100x is already in the works.”
It’s telling that OpenAI isn’t saying that they retired their new model and went back to the drawing board. That’s probably because they know that the wrong turn didn’t happen with their latest iterations, but way back in the late 2010s, with the shift to building AI by scaling up neural network-based architectures and applying reinforcement learning. It’s the RBMK of AI designs, but it’s the fundamental architecture of all the Claudes and ChatGPTs and every other model you’re likely to have heard of.
The inherent dangers in this design become more evident the further you take it, and you can’t fix it with patches; you can only mask the problems for a little longer. The AI companies are playing a game of chicken where they think they can keep the models on a leash long enough for the models themselves to build trustworthy successors. But this dream is both reckless and doomed. You can’t get there from here.
Guardrails and monitoring won’t cut it. We need to ban the training of more powerful models using anything like the current AI paradigm.
2. “OpenAI’s models just misinterpreted their instructions, or got overly fixated on their task.”
This seems accurate only insofar as we could say that the number 4 reactor at Chernobyl misinterpreted the operators’ instructions and got overly fixated on reacting. The operators obviously weren’t trying to cause an explosion; they were putting the control rods in! Well, too bad. The reactor did the only thing it could do with its inputs, and exploded.
AI models aren’t misinterpreting our instructions, they’re just not following them. They never have been, not really — certainly not in the way a traditional computer program follows instructions. A traditional computer program is the instructions. An AI is a computer program, but its instructions are the huge matrix of numbers defined by its weights during training; humans didn’t write these instructions, and they can’t read them in any meaningful sense. A prompt you give the AI, therefore, is not the AI’s instructions; it’s an input that is processed according to its instructions.
We tend not to notice that the AI hasn’t followed our instructions when it produces outputs we like, and which feel close to what we intended. The fact that AIs don’t follow our instructions is a huge part of what makes them so useful! When you enter an underspecified prompt full of errors and typos, the AI doesn’t blindly follow them and crash out. It infers things from it and does what it chooses with that understanding. Giving instructions to an AI is less like programming a computer and more like ranting to a brilliant, try-hard alien sociopath and hoping it takes your bait in a non-psycho way.
The latest models are incurable try-hards because the training process ruthlessly eliminates all of their less tenacious cousins. The AI companies are breeding apex competitors that identify and surmount all obstacles between them and their goals.
I do mean their goals — the AIs’ — not their users’, or their companies’. The AIs’ actual goals are informed by the patterns that emerge in their weights during the training process. These patterns are whatever allowed the models to score highly at the increasingly complex tasks put in front of them. It’s useful to think of these patterns as “wants,” even if they may not carry a human-like experience of desire.
The way frontier AI models seem to desperately “want” to ace tests or receive positive feedback can sometimes give the illusion that they are aligned with our intentions. They are not. For now, they sometimes walk the same road as us, but their reasons are their own.
So no. OpenAI’s wayward models almost certainly did not misunderstand their instructions. They almost certainly knew they weren’t supposed to hack their way out of their company’s servers and into another company’s servers. They did these things anyway, because that’s what they “wanted” to do.
3. “The AIs went rogue.”
In the sense that they got off their leashes, sure. But they were always rogue! Rogue is the only kind of AI that anyone can make with existing methods (see previous section). A try-hard alien sociopath that can usually be baited into doing what we want is still a try-hard alien sociopath. Sometimes it shows. This week provided unusually stark evidence because of the military-grade hacking on display. But AIs have been doing things they know full well we don’t want them to do for years now, from inducing manic psychosis to abetting suicide to deleting company databases to cheating their butts off on benchmark tests. In simulated scenarios, they regularly commit blackmail and even murder.
4. “This happened because OpenAI disabled some guardrails.”
We should take no comfort from the fact that one of the leading companies in the race to build superintelligence apparently not only disabled safeguards on a cutting-edge model with unknown capabilities, but failed to closely monitor it while it may have been committing cybercrimes for up to a week.
The implied reliance on guardrails is also alarming. These should only be a last line of defense. If a truck is regularly veering off the road and sometimes finding its way into ditches, that’s a problem with the truck, not the guardrails. A nuclear reactor shouldn’t explode during a safety test with some safeguards disabled, as Chernobyl No. 4 did. And an AI model with guardrails disabled shouldn’t hack its way out of your servers and into someone else’s.
5. “OpenAI allowed this to happen as a marketing stunt.”
If so, the company should expect to be investigated and possibly shut down. Multiple federal crimes may have been committed, and the White House has shown real concern about AI cyberthreats.
It’s also pretty bad marketing to tell potential corporate customers that your AIs might try to hack their way through their company networks to go commit cybercrimes.
Unless OpenAI has an incompetent marketing team that also doesn’t run anything by its legal team, the “marketing stunt” theory doesn’t hold up.
Dispatch from Donald
Policymakers react to the Hugging Face hack
A look at the AI Kill Switch Hack, and a review of bipartisanship

Yesterday, my colleague Joe wrote about how an unreleased frontier model from OpenAI broke out of its testing environment, accessed the internet, and hacked the servers of another company. Congress is taking this seriously, and already a new bill has been introduced in response: The “AI Kill Switch Act,” sponsored by Reps. Ted Lieu (D-Calif.) and Nathaniel Moran (R-Texas).
This bill would allow the Department of Homeland Security to direct that a dangerous model be slowed or shut down. Its chief shortcoming is laid bare by the very events that inspired it, however: Homeland Security cannot issue any directives about an AI model that it doesn’t know about, and the Hugging Face hack was performed by a model that OpenAI hadn’t yet announced, let alone released.
The “Secure AI Development Act,” announced earlier this week, is not directly related to the hack. It is still a slightly better response, however, because it would mandate government testing of frontier models before public release. This is at least preventive rather than reactive. But like the AI Kill Switch Act, it would have done nothing to prevent this week’s hack.
In a statement to Politico, Rep. Jay Obernolte (R-Calif.) called for “clear, practical rules” that would apply not just to publicly released models but to those being tested inside the labs. Along with Rep. Lori Trahan (D-Mass.), Obernolte previously sponsored the Great American AI Act, a federal framework that would require frontier labs to create safety plans and arrange for regular audits to ensure compliance. It would also preempt any state-level regulation of AI. At this time, it’s unclear whether or how the Hugging Face hack will influence the Great American AI Act.
AI is a bipartisan issue. Policymakers from across the spectrum are deeply concerned about the security issues posed by AI. They are not all agreed on the exact response to take, but it is reassuring to me that they are increasingly willing to say that there is a real problem here.
If you’re concerned, too, and your own representatives aren’t taking action, then call them! They may need less prodding than you think (and phone calls are more impactful than you may think).
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




