You can do better than these five takes on the Hugging Face hack
I give a B+ to the media's coverage of this event. Here's how to get an 'A'.

The Hugging Face hack by internal OpenAI models, as well as the containment breach revealed earlier in the week, have deservedly captured the media’s attention. I’ve found most of this reporting to be above average in quality, acknowledging the seriousness of the attack and recognizing some of its implications.
But I’ve also seen plenty of articles make or share dubious claims and assumptions about the incident — some taken at face value from OpenAI itself. Without calling out anyone by name, below are five takes that could be improved upon.
1. “We should address this incident with better security, monitoring, and reporting.”
Those are all good things, but they imply moving forward with business as usual. This is insanity, when business as usual means “same basic design, only more powerful.”
When the catastrophic nuclear mishap at Chernobyl revealed previously unknown flaws in the inherently dangerous RBMK reactor design, tighter protocols around the remaining RBMKs in service were certainly justified. But you didn’t hear reactor companies saying, “and with these safeguards in place, we will proceed to building the super RBMK, a 10x larger model with world-beating outputs. Our ultra RBMK 100x is already in the works.”
It’s telling that OpenAI isn’t saying that they retired their new model and went back to the drawing board. That’s probably because they know that the wrong turn didn’t happen with their latest iterations, but way back in the late 2010s, with the shift to building AI by scaling up neural network-based architectures and applying reinforcement learning. It’s the RBMK of AI designs, but it’s the fundamental architecture of all the Claudes and ChatGPTs and every other model you’re likely to have heard of.
The inherent dangers in this design become more evident the further you take it, and you can’t fix it with patches; you can only mask the problems for a little longer. The AI companies are playing a game of chicken where they think they can keep the models on a leash long enough for the models themselves to build trustworthy successors. But this dream is both reckless and doomed. You can’t get there from here.
Guardrails and monitoring won’t cut it. We need to ban the training of more powerful models using anything like the current AI paradigm.
2. “OpenAI’s models just misinterpreted their instructions, or got overly fixated on their task.”
This seems accurate only insofar as we could say that the number 4 reactor at Chernobyl misinterpreted the operators’ instructions and got overly fixated on reacting. The operators obviously weren’t trying to cause an explosion; they were putting the control rods in! Well, too bad. The reactor did the only thing it could do with its inputs, and exploded.
AI models aren’t misinterpreting our instructions, they’re just not following them. They never have been, not really — certainly not in the way a traditional computer program follows instructions. A traditional computer program is the instructions. An AI is a computer program, but its instructions are the huge matrix of numbers defined by its weights during training; humans didn’t write these instructions, and they can’t read them in any meaningful sense. A prompt you give the AI, therefore, is not the AI’s instructions; it’s an input that is processed according to its instructions.
We tend not to notice that the AI hasn’t followed our instructions when it produces outputs we like, and which feel close to what we intended. The fact that AIs don’t follow our instructions is a huge part of what makes them so useful! When you enter an underspecified prompt full of errors and typos, the AI doesn’t blindly follow them and crash out. It infers things from it and does what it chooses with that understanding. Giving instructions to an AI is less like programming a computer and more like ranting to a brilliant, try-hard alien sociopath and hoping it takes your bait in a non-psycho way.
The latest models are incurable try-hards because the training process ruthlessly eliminates all of their less tenacious cousins. The AI companies are breeding apex competitors that identify and surmount all obstacles between them and their goals.
I do mean their goals — the AIs’ — not their users’, or their companies’. The AIs’ actual goals are informed by the patterns that emerge in their weights during the training process. These patterns are whatever allowed the models to score highly at the increasingly complex tasks put in front of them. It’s useful to think of these patterns as “wants,” even if they may not carry a human-like experience of desire.
The way frontier AI models seem to desperately “want” to ace tests or receive positive feedback can sometimes give the illusion that they are aligned with our intentions. They are not. For now, they sometimes walk the same road as us, but their reasons are their own.
So no. OpenAI’s wayward models almost certainly did not misunderstand their instructions. They almost certainly knew they weren’t supposed to hack their way out of their company’s servers and into another company’s servers. They did these things anyway, because that’s what they “wanted” to do.
3. “The AIs went rogue.”
In the sense that they got off their leashes, sure. But they were always rogue! Rogue is the only kind of AI that anyone can make with existing methods (see previous section). A try-hard alien sociopath that can usually be baited into doing what we want is still a try-hard alien sociopath. Sometimes it shows. This week provided unusually stark evidence because of the military-grade hacking on display. But AIs have been doing things they know full well we don’t want them to do for years now, from inducing manic psychosis to abetting suicide to deleting company databases to cheating their butts off on benchmark tests. In simulated scenarios, they regularly commit blackmail and even murder.
4. “This happened because OpenAI disabled some guardrails.”
We should take no comfort from the fact that one of the leading companies in the race to build superintelligence apparently not only disabled safeguards on a cutting-edge model with unknown capabilities, but failed to closely monitor it while it may have been committing cybercrimes for up to a week.
The implied reliance on guardrails is also alarming. These should only be a last line of defense. If a truck is regularly veering off the road and sometimes finding its way into ditches, that’s a problem with the truck, not the guardrails. A nuclear reactor shouldn’t explode during a safety test with some safeguards disabled, as Chernobyl No. 4 did. And an AI model with guardrails disabled shouldn’t hack its way out of your servers and into someone else’s.
5. “OpenAI allowed this to happen as a marketing stunt.”
If so, the company should expect to be investigated and possibly shut down. Multiple federal crimes may have been committed, and the White House has shown real concern about AI cyberthreats.
It’s also pretty bad marketing to tell potential corporate customers that your AIs might try to hack their way through their company networks to go commit cybercrimes.
Unless OpenAI has an incompetent marketing team that also doesn’t run anything by its legal team, the “marketing stunt” theory doesn’t hold up.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


