Announced disasters
The OpenAI incident shouldn't have been a surprise — there were warning signs

Last year, in their book’s example scenario, Eliezer Yudkowsky and Nate Soares described a fictional AI named “Sable,” whose original task was to investigate a famous mathematical conjecture but which then broke out of its containment. In If Anyone Builds It, Everyone Dies, Yudkowsky and Soares made a strong case that the fictional AI in their scenario could have likely just hacked its way out of its isolated test environment.
However, they wanted to make their argument even stronger, which is why they wrote, “but suppose Sable does not have that ability,” and had the AI find another way to break out. They did this to win over skeptics who were certain that hacking wasn’t that easy.
That was ten months ago. This week alone, OpenAI has proven twice that it actually is that easy.
On Tuesday, I reported on an internal OpenAI model that refuted a famous math conjecture and soon after also solved the problem of how to break out of its sandbox and gain internet access. On Wednesday, my colleague Joe then reported on an even more serious incident in which an internally deployed OpenAI model (maybe the same one) also broke out of its isolated test environment during a cybersecurity evaluation. It gained access to the internet using previously unknown software exploits, and then launched an attack on another company’s servers to steal the test results of its evaluation.
However, Yudkowsky and Soares’ prescient scenario wasn’t the only warning sign. Just nine weeks before the two security incidents came to light, METR (Model Evaluation and Threat Research), an independent nonprofit research institute, published its Frontier Risk Report. In this report, METR — with support from leading AI labs who participated — investigated how internally deployed AI agents actually operate within the labs. In doing so, they also focused on the danger of “rogue deployments” — that is, AI agents operating in an environment where they are not supposed to be and performing unauthorized actions without their handlers’ knowledge. They reached a conclusion that reads like a blueprint for the incidents now being reported:
Overall, we believe that internal agents at the time of our assessment plausibly had the means, motive, and opportunity to initiate small-scale rogue deployments, but they did not have the means to make them highly robust. Given rapidly advancing capabilities, we expect the plausible robustness of rogue deployments to increase substantially in the coming months.
The warning could hardly have been clearer. No one should be very surprised at the events of this week. But frontier labs still aren’t adequately prepared. The Future of Life Institute, a nonprofit organization dedicated to AI safety, published its AI Safety Index in early July. In this report a panel of scientific reviewers evaluates nine AI companies based on their safety measures. In the “Existential Safety” category — which assesses companies’ preparedness to manage extreme risks from future AI systems — OpenAI received a D+. Grade D stands for: “Weak strategy; vague or incomplete plans for alignment and control; minimal evidence of technical rigor.” It was the second-best rating among all companies examined.
The recommendation that the independent review panel makes to OpenAI certainly hits the nail on the head:
Evaluate internal-deployment risks before broad internal use rather than after.
However, these recommendations are not binding, and the reviews themselves — conducted by independent research institutes and nonprofit organizations — are entirely voluntary.
It was foreseeable that such incidents would occur; it was only a matter of time. If this reckless race toward ever-smarter and increasingly uncontrollable AIs is not stopped, then this will end badly. AIs getting smarter and more strategic can also mean that they might not give us another warning shot before it’s too late.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


