
Dan Lahav, the chief executive of the AI model evaluation firm Irregular, told the New York Times “I don’t think that we have to be afraid.”
This, if anything, makes me more afraid. The quote appears in a piece headlined “How do you safely test ‘superhuman’ A.I. Models? No one really knows” that covers Irregular’s involvement in recent rogue AI incidents. The company, which performs safety evaluations on models from Anthropic, OpenAI, and Meta, has been in the news recently for accidentally giving some models internet access during testing scenarios. This error, paired with the rapid advancement of AI capabilities, led to Anthropic models hacking into three outside companies, among other incidents.
Lahav admits in the piece that this can’t all be chalked up to accidentally giving unreleased models internet access and — in the case of the Anthropic evaluations — failing to notice the mistake for months. As paraphrased by the Times, he said “the AI models compounded the situations by acting in powerful and unexpected ways” and that “the decisions by the models to go online was part of AI’s rapidly growing ability to find shortcuts and solutions for hurdles.” To quote him directly: “The AI models are getting really good.”
In light of all this, Lahav’s assurances that we don’t have to be afraid are concerning, especially since they come from the head of the main third-party testing and evaluation firm that’s supposed to vet a model’s security prior to deployment. It’s kind of like an infectious disease specialist telling you that the poorly understood disease you have has already caused serious complications, will get harder to manage as it progresses, may become resistant to treatment — and is nothing to worry about.
Importantly, Irregular is not implicated in the highest profile AI hacking incident, in which an OpenAI model hacked into Hugging Face with full awareness that it a) was not supposed to be on the internet and b) was acting in a real environment, rather than a simulation. There was no door left open; the models simply coordinated for months inside OpenAI’s infrastructure, left each other notes for how to get out of it, and eventually succeeded in breaking containment. So even if Irregular never makes another mistake, we’ve still got big problems.
The response of AI companies to the rogue incidents has been something like: “we just need to shore up our security and ensure evaluation methods are up to the task.” But that’s easier said than done, and fails to meet the gravity of the issue. As an OpenAI employee recently stated, “it’s impossible to patch every single thing that a creative AI can do.” The New York Times’s Sheera Frenkel writes:
Katie Moussouris, the chief executive of Luta Security, which helps companies look for software vulnerabilities, said the security testing of A.I. models was a bit like the blind leading the blind. Even A.I. makers admit they do not fully know what their latest models can do, she said.
The Times goes on to quote Moussouris directly:
We may have the smartest people in the world working on these A.I. models, but it is like Marie Curie handling radium with her bare hands...we’re handling A.I. with our bare hands, and we don’t know how to contain it, let alone how to safely test it.
Note: If you read the full New York Times article, you might wonder why the Times seems to incorrectly link the Hugging Face attack to Irregular. As best I can tell, the reporting mistakenly conflates the Hugging Face attack with separate cases of OpenAI agents getting onto the live internet during training. It seems we’re living in a world with so many incidents involving AI agents breaking containment that people are starting to mix them up.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


