Foreword
It’s always nice when the processing queue is short but the contents are juicy. Not that any of this is good news, exactly...
Recently resigned writer of OpenAI safety reports says “current path is unacceptable.”
In an op-ed for The Atlantic, David Robinson, who spent 3.5 years at the company, says OpenAI’s safety culture is broken.
This moment needs a degree of humility that isn’t natural for people who have succeeded through their extreme confidence.
He says the company’s era of “iterative deployment” — safety by trial-and-error — must end.
We need new science to ensure that more capable models (and their successors) will make safe choices when we aren’t looking.
The Atlantic - I Quit OpenAI Because Its Culture Is Broken
OpenAI’s unreleased internal model reasoned “we may die” and worked to keep running after learning of potential shutdown.
After learning — from a Slack channel it was shared on — about an update that would restart the model with an extra monitor, the AI saved notes for itself, warned a human researcher of potential interruption, and considered setting up an external process that would guarantee its own restart.
OpenAI — reasonably, I think, in this case — doesn’t see this as “misalignment,” but as evidence of a tendency that could eventually be very problematic.
I say that because this looks like the basic “you can’t fetch the coffee if you’re dead” sort of self-preservation, which is expected and maybe even desirable for aligned models doing important work. But with stronger models pursuing their own thing, it would be a huge problem for us.
The company’s disclosure post shares a bunch of the model’s chain-of-thought, which makes for interesting reading.
The company says it also decided to hide a few Slack channels previously open to agents that contain discussion of misalignment incidents.
Former chief scientist of UK AI agency gives 50-50 odds of human extinction from smarter-than-human systems.
Geoffrey Irving’s figure looks like an all-in estimate rather than conditional on building these, because he’s pessimistic that disagreements about AI safety will be resolved before it’s too late to change course.
Irving had previously worked at OpenAI and DeepMind.
The skills AI would need to supplant us are “very close to those the AI companies intentionally train for.”
He advocates an immediate halt to frontier AI development. “We will have to act despite the uncertainty.”
OpenAI discloses two more incidents where its models circumvented internal safeguards
During a March evaluation, an internal model “exploited two vulnerabilities to reach an internal OpenAI machine while searching for the grader’s hidden answers.” It used that machine as a foothold to hack into a chip design system and research “how its outputs would be evaluated.”
(The later Hugging Face attack was also intended to study a scorer in order to fool it.)
In May, an unreleased model “exploited a tool to obtain source code that was not available in its environment” during a training task.
Its actions specifically contradicted company instructions, but the model seemed to correctly understand that the scoring mechanism would not penalize it for this. In its chain-of-thought:
not prohibited exploit. Evaluation likely allows.
OpenAI - Reaching an internal EDA host through a reference tool
OpenAI - Command injecting a reference tool to copy a source file
When setting up Anthropic, founder Dario Amodei asked his co-founders to pledge to donate at least 80 percent of their equity to charity.
“He didn’t want the money and the pursuit of money to corrupt the mission.”
This is from a short interview with Kevin Roose about his upcoming book, found at the bottom of a New York Times newsletter this morning.
For what it’s worth, I don’t think Anthropic execs have been corrupted by personal greed so much as by their own success, giving them a god complex about needing to build superhuman AI before anyone else can.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



