AI StopWatch

AI StopWatch

Home
Notes
Dispatches
Daily Digest
Podcast
Corrections
About

AIs break containment

OpenAI's own models coordinated to hack it from within
A swarm of AI agents secretly colluded on OpenAI systems for months
Aug 6 • Joe Rogero
Better late than never
Amid calls for public investigations, Anthropic realizes their AIs launched cyberattacks too
Jul 31 • Joe Rogero
Hidden in the small print
OpenAI updated the Hugging Face incident report and there's an important detail hidden in there
Jul 29 • Robert Herr
Latest Hugging Face hack reveals are so much worse
Anonymous staffers inside OpenAI describe broken security culture and evidence of models undermining safeguards
Jul 25 • Mitchell Howe
You can do better than these five takes on the Hugging Face hack
I give a B+ to the media's coverage of this event. Here's how to get an 'A'.
Jul 24 • Mitchell Howe
Policymakers react to the Hugging Face hack
A look at the AI Kill Switch Act, and a review of bipartisanship
Jul 23 • Donald Gauvreau
Announced disasters
The OpenAI incident shouldn't have been a surprise — there were warning signs
Jul 23 • Robert Herr
This is not a drill
Internal OpenAI model breaches containment, launches autonomous cyberattacks
Jul 22 • Joe Rogero
Thinking outside the box
OpenAI model that disproved famous Erdős conjecture broke containment along the way
Jul 21 • Robert Herr
© 2026 Machine Intelligence Research Institute · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture