Systems that we can’t aim and we can’t inspect
All the brouhaha about Sam Altman’s “singularity” aside is missing the point
“Singularity” is a heavy word. Even just in the context of artificial intelligence, it’s been used in a few different ways — artificial intelligence designing a more intelligent version of itself, or technological improvement feeding faster technological improvement — but the underlying theme is an acceleration of exponential growth that outpaces our ability to keep up with it. Like crossing the event horizon of a black hole, it’s hard to see what awaits you on the other side.
The AI self-improvement singularity is particularly important. Beyond it, human beings are no longer driving AI progress, and no longer able to reliably predict, understand, or control what the machines do next. So, it’s no wonder then that people sat up and took notice when Sam Altman said, “We are now, like, in the singularity.” Another point worth mentioning: He admitted that there are “still major AI alignment issues to solve and safety issues.” This was just a brief aside, a few sentences in a conversation almost exclusively about startups, but Altman’s comments have set off a round of arguments about the wrong thing.
Forbes’ Lance Eliot, for example, makes the case that:
[Altman] is not saying we have crossed the Rubicon and attained the AI singularity. I believe he is emphasizing that there isn’t likely a tipping point; there isn’t a single moment in time. It is a curve we are climbing incrementally.
I think this is a fair interpretation, if not of Altman then at least of the situation. You can’t take a look at some graphs, punch a few numbers into your calculator, and say, “We have now entered the AI singularity,” precisely because there are multiple competing and fuzzy definitions. Altman himself says as much not long after: It is “all one crazy exponential,” and no single moment is the tipping point.
So let’s stick to the facts: We are in danger of losing control.
Earlier this month, the machine learning company Hugging Face suffered a severe cyberattack. The responsible party was not a state actor or a criminal syndicate. It was an OpenAI agent that had escaped its testing sandbox, obtained internet access, and broken into Hugging Face’s servers by exploiting a chain of security vulnerabilities. That agent was not supposed to do any of those things. (“...still major AI alignment issues to solve and safety issues,” indeed!)
In his podcast interview about startups, Altman said that “we are close to creating a genie that can grant any wish.” I feel that this statement is inaccurate, to put it mildly. Our current AI models aren’t even trickster genies who will grant your wish in a twisted way that’s guaranteed to make you suffer. For all its malice, the trickster genie has an important property: It still at least executes the wish. The problem is in the wording, which means that careful wording is the solution. But not even the wiliest contract lawyer could solve the problem that we actually face.
Nobody wished for OpenAI’s model to break out of its sandbox, or access the internet, or hack another company’s servers. But the AI model wasn’t just doing its level best to complete the task it was given. AI models are trained to go hard at whatever they’re pointed at. In this case, that was a test score, and not the thing that the test was for. OpenAI wanted to measure capabilities; the model wanted a perfect score. These are not the same thing, and we can be pretty sure that the model didn’t think they were the same thing.
We can’t know for certain what actually motivates these models under the hood. We can watch what they do and draw inferences about what drives them, but we cannot guarantee that they truly internalized whatever it is that we trained them on.
Writing in The Guardian, Bruce Schneier and Barath Raghavan propose a score they call the “Genie coefficient,” to measure the distance between what you asked the AI model to do and what it did. Scoring means putting an AI model in a walled-off sandbox, with real tools it could misuse, and grading its worst behavior, not its typical behavior. (That last part is a good move that I really appreciate.) I still think the Genie coefficient has a very big flaw. Like the AI Kill Switch Act, it doesn’t prevent danger. The Hugging Face hack was performed by an AI agent that broke out of a testing sandbox; I’m not enthusiastic about stress-testing how disobedient an advanced AI system can be under the assumption that we definitely, for sure, have locked all the doors and windows this time.
Right now, the safest thing to do is to stop building these things. It’s good to have a law about shutting down an AI model that acts dangerously, and it’s good to have benchmarks to weigh how uncontrollable an AI model is. There are ten thousand good things worth doing. But Altman is right that there’s no clear tipping point, no single moment to wait and watch for, just a steepening curve. Altman calls it a “glide path,” a course that leads smoothly and inexorably to some particular destination.
If you find yourself on a glide path and you’re beginning to think that the destination might not be so nice, then the best thing to do is to steer to some different path.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.
You can receive emails of dispatches as we write them, or subscribe to our Daily Digest for a once-a-day compilation.



