Open models are uncontrollable models
Frontier models are dangerous enough without making their weights public

When the machine learning company Hugging Face suffered a severe cyberattack last week — before it learned that the responsible party was an OpenAI agent that had escaped a sandbox — it turned to a Chinese open model to contain the attack. Hugging Face did so because the leading US models’ safeguards were too strict; there were (valid) concerns that dishonest parties could launch cyberattacks by telling a model that they were trying to prevent one. (Anthropic believes that this technique was attempted in at least one major cyberattack already.)
But open models carry their own dangers. A model with public weights cannot be recalled, and its safeguards can be trivially broken. Open models are free models — not just free as in “costless,” but free as in “uncaged, loose, uncontained and uncontainable — and perhaps someday casually hacking a $4.5 billion company.” (In that sense, open models are extremely costly.)
Last week, Yoshua Bengio, the world’s most-cited living scientist, warned that the deployment of open models is an irreversible decision. “The logical thing to do,” he said, “is to find a good evaluation of these models, share the models that are not too dangerous, and not share those above the threshold of risk.”
That said, whether the models are American or Chinese, private or open, this research remains dangerous. Open models aren’t the risky option, they’re the riskier option. OpenAI is a private frontier lab and they failed to keep their model locked up.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.
You can receive emails of dispatches as we write them, or subscribe to our Daily Digest for a once-a-day compilation.


