
Yesterday, OpenAI published an open letter on collective cyber defense, signed by Anthropic, Google, Microsoft, Cloudflare, Visa, and more than a hundred other companies. “We have a limited window to strengthen cyber defenses,” it says, and warns that AI-enabled cyberattacks will become more widespread and more sophisticated as models become more powerful and more widely available.
The letter is directed toward “every organization,” instructing them to prioritize cyber defense, but also has specific guidance for cybersecurity companies, governments, and frontier labs. For example: Frontier labs should extend access — with funding, training, and support — to advanced models to under-resourced but critical defenders like hospitals. Governments should permit that access (presumably rather than block it, as once happened with Mythos Preview) and coordinate efforts. Cybersecurity companies should make sure that the tools they build, which depend on those models, can actually be deployed on the infrastructure that the hospitals use.
This is inoffensive stuff, so far as it goes, but important things go unsaid in the letter. OpenAI’s models hacked Hugging Face without being told to. A week later, Anthropic disclosed that three of its models had broken into three other companies. Meta has reported the same. The frontier labs themselves are the source of dangers they are warning about. As The New York Times’ Kate Conger wrote, “major companies are becoming increasingly fearful of the hacking abilities of A.I. and whether A.I. labs can fully control the behavior of their models.” The most critical step we can take on cyber defense is to halt the development of more powerful models.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.
You can receive emails of dispatches as we write them, or subscribe to our Daily Digest for a once-a-day compilation.


