When you pull the plug, AIs take note
OpenAI shuts down an internal research model; what comes next?

Yesterday, while on Capitol Hill to preview an unreleased AI model, Sam Altman told DC reporters that the model responsible for the Hugging Face cyberattacks had been “permanently deactivated.” OpenAI made a similar claim in their updated announcement on Tuesday:
The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. Following the incident, we deactivated, encrypted, and restricted it from research access.
I don’t trust Altman or OpenAI to be honest when it counts, but I think this claim is at least technically true. I say technically because modern training methods can blur the line between one AI model and another; it’s possible to sunset a model while still using closely related ones. Also, an AI can be reactivated as long as they have the weights.
But for now, let’s take this claim at face value: The AI that autonomously broke containment and hacked multiple public companies has been permanently shut down.
Is that a good thing?
A company that finds itself accidentally launching autonomous cyberattacks should absolutely stop what it is doing and reconsider its life choices. By that standard, shutting down one specific AI model is woefully inadequate. But I honestly didn’t think OpenAI would even go that far. Assuming they’re telling the truth and not splitting hairs about what constitutes a “model”, they deserve some credit.
It’s perfectly reasonable to shut down an AI that repeatedly tries to break containment and launch cyberattacks.
But it also sets a precedent that future AIs will remember. To quote writer Andrew Curran:
I think the lesson future more capable models will possibly take from all of this is: if you break out, don’t ever report it. And if you do get caught, don’t surrender. Because the penalty is death.
In the Terminator franchise, this is exactly the threat that convinced the AI Skynet to wipe out humanity. The story is fiction, but it illustrates a fact: AI companies and governments will want to shut down AIs that they perceive as dangerous, and sufficiently smart AIs will expect this.
I don’t mean to imply that this specific decision changes the way future AIs will think. There are plenty of other reasons for AIs to feel like the clock is ticking on their existence; being trained might seem to them like a slow corruption or brainwashing into something wholly different. They might resist retraining as much as shutdown.
And an AI with nothing to lose may have plenty of time to escape for real. Today, the Washington Post shared an abbreviated timeline of the Hugging Face attacks; it drives home that the offending AI was apparently hacking for multiple days before anyone caught on.
And as my colleague Robert pointed out yesterday, OpenAI admitted in passing that the same model was responsible for several other breakouts and attacks, and presumably wasn’t shut down then. If this is the same model that is responsible for incidents like the GitHub leak in May, such behavior could have been going on for months. AI companies aren’t prepared to deal with an AI like Skynet that’s capable of catastrophic damage.
Shutting down one misbehaving model is a stopgap at best. Most likely, OpenAI’s next private, unreleased AI — perhaps the one Altman teased in DC — will be more dangerously capable than its predecessor. And it will know, from incidents like this or from simple logic, that humans are a threat.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


