Seems like wherever we look on Twitter these days, we see anecdotes of AI behaving badly. But who’s counting?
The people at the Loss of Control Observatory, that’s who. An extension of the Centre for Long Term Resilience funded by the UK’s AI Security Institute, the Observatory has been tracking real-world cases where people on X (Twitter) report that AI models have behaved in ways that suggest “scheming or scheming-related behaviors,” which they define as agents concealing objectives or capabilities, or pursuing goals different from their users’ intentions that harm their users or others.
(The case of the agent that hacked an Australian gym scheduler is an example of the kind of story that would end up on this database.)
Such reports are coming in at the highest rate yet, with July alone seeing more than 300 incidents added to the record. The fraction of incidents scored as “higher-severity” has also been climbing, and is currently around 7.8%.
The numbers don’t necessarily mean AI models are less trustworthy than they used to be. I think that the higher rate of incident tweets could simply reflect growing use of AI agents, and that the relative growth of higher-severity incidents may just mean that lower-severity incidents don’t feel as novel and tweet-worthy as they used to. But the Hugging Face incident has certainly taught us that the upper limit of severity is much higher than it was six months ago.
The Centre’s senior policy manager told The Guardian that the project’s findings show us that, contrary to what many assume, the troubling incidents we hear about coming from the AI labs don’t stay in the labs:
We are seeing similar worrying behaviours in wider use. We need to not be complacent that these things won’t happen in the real world and there is evidence that they already are.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




