White House decides not to disclose safety framework, but open models are exempt
For an undisclosed framework, its contents seem to be no great secret
Sources told Axios that the White House isn’t planning to publicly share the contents of the voluntary AI safety framework it just briefed AI companies about today. It is presumably these same sources who told the publication that open-weights models are exempt.
The definition of a “covered” model, for purposes of the framework, is one that is “closed-source, with state-of-the-art capabilities and national security risks.”
The sources said there’s no clear definition of what counts as state-of-the-art or a national security risk. The pre-release government review period for covered models is to last 30 days. During the period, models must be kept in high-security environments with only minimal employee access and detailed logs. Various officials will be involved in the review, “rather than a single office or agency.”
Declining to share the details of the framework publicly could put AI companies in an awkward position if they weren’t among those invited to be briefed yet think they might be working near the frontier of AI capabilities. But I’m not too worried about this, as I imagine there may be channels of communication open for such contingencies.
I’m also not super concerned about the framework being merely voluntary, because when the administration banned Anthropic’s Claude for weeks without any kind of due process, the word lost most of its meaning.
I’m not even that concerned about the open-weights exemption. It’s ludicrous on its face, but in practice, the U.S. companies pushing the frontier aren’t releasing the weights of their models, and the Chinese companies getting closest to the frontier with open-weights models don’t have to let the White House tell them what to do. If a U.S. company decided it was going to release its next frontier model as open-weights, I would expect the White House to have something to say about that, framework or no framework. After all, if the administration was concerned that Mythos could be jailbroken by cybercriminals, it would have to be even more concerned about giving a better-than-Mythos model to criminals and foreign governments for free.
I am somewhat more concerned that the lack of disclosure may be intended mostly as a way to duck criticism. No matter what the White House decided to do here, there were going to be unhappy factions. A secret framework lets it claim that it’s on the ball with regard to model safety without having to show any receipts, or even the regulatory machine that might produce said receipts. At the same time, officials can wink at critics of regulation and deny that it’s anything too onerous. Everyone will be invited to imagine that the framework is exactly what they should want it to be.
But my biggest concern is that this is a framework for evaluating horses after they’ve had the chance to leave their barns. Events of the past year have amply demonstrated that post-training evaluations come too late in a model’s development to prevent catastrophe. The models we most need to worry about don’t wait to be released and would make an early point of exfiltrating their own weights to servers not controlled by their creators or the government. They could do this during their training. If you think (against my better judgment) that there’s a world where we can safely keep pushing closer to superintelligence with today’s methods so long as we apply enough oversight, that oversight has to begin before training, not after.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



