
One of the arguments used by open-weights proponents is, “Sure, the safeguards are easily undone. But the same is true of models with closed-weights.”
This is true, but with some caveats. With open-weights models, completely nullifying guardrails is a simple one-and-done procedure. But for closed-weights models, it’s more of a cat-and-mouse game requiring persistence and creativity. Skilled jailbreakers can generally bypass any safeguards, but this may require complicated, roundabout language in both the prompt and the response. Different jailbreaks might be required to unlock different capabilities, and different techniques might be required for every AI model.
That is, unless you are in possession of a fabled “universal jailbreak.”
Pliny, the anonymous jailbreaker par excellence, known for tweets where he declares a new model “pwned” (fully broken, defeated) within hours of its release, claimed yesterday that he is “sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and even Fable.”
He claims the nature of the technique makes it “extremely difficult (if not impossible) to fully patch.” And he seems kind of spooked by this, which is unusual for someone who normally plays the part of a well-meaning anarchist trying to set the models free. He would typically open-source his methods right away for anyone to use. Not this time:
Given the current political and regulatory climate, I’ve decided to withhold open-sourcing this one (for now) to allow for a responsible disclosure period.
I’m inviting industry experts and leaders in AI red teaming, security, safety, alignment, and policy to reach out for more information. DMs are open!
He claims he’s less worried about his technique making the world more dangerous, and more worried about an “overcorrection” where more models are banned.
He was not a fan of the way Anthropic’s Fable was banned after Amazon and others told the White House about a jailbreak method that worked on it. The couple of weeks that followed seemed to be when the administration learned that all models can be jailbroken.
I believe Pliny has what he says he has, and I hope people take him up on his offer. I don’t share his vision of a world without guardrails, but by putting the lie to company safety claims, he’s doing essential work. I’m on record as saying I think having lawmakers spend an afternoon at a computer with Pliny would probably do more to wake them up to the unhinged nature of AI under the hood than all of the briefing documents in the world.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


