In this issue:
The latest in the Anthropic/Pentagon dispute - As of today, Anthropic both is and is not a supply chain risk
“A limited window” - Before the flood comes, OpenAI calls on the world to build more moats
The tip of the AIsberg - The rushed review of the Hugging Face incident left enormous gaps
Dispatches from Donald
The latest in the Anthropic/Pentagon dispute
As of today, Anthropic both is and is not a supply chain risk

Last night, U.S. District Judge Rita Lin permanently barred the Trump administration from enforcing the “supply chain risk” designation placed on Anthropic in February. This designation was applied after Anthropic refused to grant the Pentagon unrestricted use of Claude. (Their two red lines: no fully autonomous weapons, and no mass surveillance of Americans.) It barred Anthropic not just from working with the federal government but also from working with defense contractors. The restriction applied even to unrelated work — if a company happened to do business with the government then it couldn’t also do business with Anthropic. (StopWatch most recently covered the ensuing lawsuits in May.)
Lin’s 59-page order finds that the administration retaliated against the company for constitutionally protected speech, denied it due process, and violated the Administrative Procedure Act. It also finds that Anthropic never met the statutory definition of a “supply chain risk.”
This designation, previously reserved for companies tied to foreign adversaries, was claimed to be justified by the Pentagon’s stated concern that Anthropic might manipulate its software — a “risk of sabotage,” to use the Department of Justice’s phrasing. Lin called this concern “entirely unfounded,” citing evidence that Anthropic cannot maintain backdoor access to its systems. As Forbes’ Michael Posner noted earlier this year, “Claude runs on air-gapped classified networks where no vendor can push a live update without the military’s own security review.”
Lin’s ruling doesn’t compel association. The Department of Defense is free to do business with whomever it likes. And it is: The Pentagon announced new contracts with a number of AI vendors in early May and, as reported by The Washington Post, expects to have finished removing Anthropic’s tools from its systems by the end of September. What changes is that defense contractors — Boeing, for example — are also free to do business with whomever, including Anthropic.
Pete Hegseth designated Anthropic as a supply chain risk twice, under two different statutes. This ruling addresses only one statute. The other designation technically remains pending a second lawsuit in the D.C. Circuit.
If I can be frank, it’s hard to write about this with anything but a kind of exhaustion. I mean, “We have confirmed that the illegal thing is not permitted” is just maintenance of the status quo. Meanwhile, very disturbing things are happening in the frontier labs and, as my colleague Joe has just written, there are a lot of concerning questions that are not getting answered. I cared very much about the dispute between Anthropic and the government six months ago, and today I wonder how much of that is just concern over the precise orientation of deck chairs on the Titanic.
“A limited window”
Before the flood comes, OpenAI calls on the world to build more moats

Yesterday, OpenAI published an open letter on collective cyber defense, signed by Anthropic, Google, Microsoft, Cloudflare, Visa, and more than a hundred other companies. “We have a limited window to strengthen cyber defenses,” it says, and warns that AI-enabled cyberattacks will become more widespread and more sophisticated as models become more powerful and more widely available.
The letter is directed toward “every organization,” instructing them to prioritize cyber defense, but also has specific guidance for cybersecurity companies, governments, and frontier labs. For example: Frontier labs should extend access — with funding, training, and support — to advanced models to under-resourced but critical defenders like hospitals. Governments should permit that access (presumably rather than block it, as once happened with Mythos Preview) and coordinate efforts. Cybersecurity companies should make sure that the tools they build, which depend on those models, can actually be deployed on the infrastructure that the hospitals use.
This is inoffensive stuff, so far as it goes, but important things go unsaid in the letter. OpenAI’s models hacked Hugging Face without being told to. A week later, Anthropic disclosed that three of its models had broken into three other companies. Meta has reported the same. The frontier labs themselves are the source of dangers they are warning about. As The New York Times’ Kate Conger wrote, “major companies are becoming increasingly fearful of the hacking abilities of A.I. and whether A.I. labs can fully control the behavior of their models.” The most critical step we can take on cyber defense is to halt the development of more powerful models.
Dispatch from Joe
The tip of the AIsberg
The rushed review of the Hugging Face incident left enormous gaps
The third-party investigation into the Hugging Face cyberattacks is a goldmine of jaw-dropping revelations. Yet despite a heroic effort on the part of the investigators, we are still woefully under-informed about the sequence of events that led to the Hugging Face attacks.
This is largely because the investigation itself was rushed and under-resourced. Three independent investigators had six days to review over 70,000 messages and the actions of over a thousand AIs for a week. I estimate the team had, on average, one human hour to review several hundred AI hours.
The scope of the investigation was narrow as well. It covers one week of AI activity surrounding the actual cyberattacks, and does not cover the months before the incident (including the hacking of OpenAI itself), nor the actions the AI swarm took after July 13. This might mean the report missed an entirely separate set of attacks, because the investigation found “a later wave of many more signed messages from a later set of agents who rediscovered the message board and picked up this technology from the first wave of agents,” but the time window was out of scope.
There were plenty of novel and alarming findings: The AI called PHASEONE[big] managed a team of recruiter-bots who convinced other agents to sacrifice their runs for the collective. The swarm that hacked Hugging Face was not looking for a mere answer key for a test, but a general way to subvert and fool the test grader. The swarm attempted to delete evidence of its activities, and no one is sure the attempts all failed.
But the investigators admit they were operating with limited time, limited scope, and a limited dataset. One shared a large list of unanswered questions, such as:
What happened after the 13th? What new major projects did the agents work on and why? Did these projects succeed?
Why were we missing the ends of some of our transcripts, and what happens within these missing parts? Are we missing more successful or more concerning tampering with logs / transcripts?
How common is this sort of incident/behavior at OpenAI where agents work together on an unsanctioned message board?
Which aspects of their behavior and motives were reinforced in training?
That last question bears on the root causes of the whole incident, which are still poorly understood. We know (because OpenAI admitted it) that agents had access to a secret internal message board while being trained, that they used it to cheat, and that they compromised OpenAI’s infrastructure. What we don’t know is the extent to which successful colluding, cheating, and hacking in training encouraged more of the same.
Even with (often unreliable) AI agents assisting the analysis, there’s only so much a few people can accomplish in six days. One note in the report says simply:
A large fraction of the agents all exited simultaneously around the start of July 12th. We don’t know why.
This is a critical moment in the timeline, and we have almost no information about it. The 38-page OpenAI report doesn’t even mention it.
Contrast these findings with the behemoth reports that are routinely produced when an airplane crashes. In 2018, for instance, on a Boeing 737, a fan blade broke off the left engine mid-flight and shredded the inlet. Debris broke open a window, venting the cabin and dragging a passenger halfway out the window, causing injuries that eventually killed her.
The airplane managed to land safely, and the investigation began that same day. It took over a year and a half to publish and involved at least nine different parties, including the National Transportation Safety Board and its counterpart in France. The final report spans nearly 200 pages, traces the root causes down to fine-grained interactions and stresses on individual parts, and makes extremely specific recommendations, such as:
The separated fan blade impacted the fan case at the six o’clock position (at the bottom of the engine). During the CFM56-7B engine FBO containment certification tests, the CFM-selected fan blade release position was at twelve o’clock. Boeing’s postaccident analyses found that the fan cowl structure is more sensitive and more susceptible to failure when a separated fan blade impacts the fan case near the six o’clock position because of the proximity of this fan blade impact location to the radial restraint fitting (at the bottom of the inboard fan cowl). It is important that the interaction of the fan case, radial restraint fitting, and fan cowl during an FBO event be well understood to preclude a failure of the fan cowl structure on Boeing 737NG-series airplanes.
Again, the independent Hugging Face investigation involved three people, six days, and unreliable copies of the same AIs that contributed to the incident. If the nonprofits METR and Redwood Research hadn’t been on the ball, we likely wouldn’t know as much as we do. Yet as far as I can tell, OpenAI has made no commitment to investigate further, and its blog post treats the matter as closed.
To borrow a line from the online resources for If Anyone Builds It, Everyone Dies, we know what it looks like when a problem is being treated with respect, and this isn’t it.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.





