For unclear reasons, Anthropic broke with precedent and declined to let the UK AI Security Institute (UKAISI) evaluate its Mythos 5.1 model before release. The Financial Times covers the omission, but was unable to find out why; there’s a suggestion the U.S. government might have stepped in to block the sharing.
If true, this would be a major step back for Anthropic, which earlier this year took a principled stance against its models’ use in surveillance and autonomous weapons and was (illegally) declared a “supply chain risk” for its trouble.
Maybe Anthropic bowed to pressure from the administration not to share its models. Maybe it was spooked by the time its models launched cyberattacks from a UK AISI evaluation, despite the fact that AISI demonstrated more competence in its response than most developers. Maybe it simply didn’t want to share.
Whatever the reason, this development seems bad. Third-party evaluators perform a crucial service, providing transparency and early warning of dangerous capabilities, and UKAISI filled an important niche. Its U.S. counterparts are not as well-resourced, and Anthropic’s own internal assessments have not inspired confidence.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



