In this issue:
We need to do better. But not today. - Two more blog posts from OpenAI
Dispatch from Mitch
We need to do better. But not today.
Two more blog posts from OpenAI

Like someone caught in a lie mid-conversation, OpenAI is talking fast and changing the subject. It put out two significant blog posts today. Both are newsworthy, but neither is about the growing evidence that the company has been, and may still be, concealing the extent and severity of its rogue agent swarm issues.
So it feels important that I preface my reporting on these posts with that reminder, and that I try not to overly signal-boost the story OpenAI wants to tell about itself. I’ll do my best.
Research acceleration
The company’s first post of the day is called “Research acceleration: The view inside OpenAI.”
The topic is important. The world needs to understand that AI development is being increasingly offloaded to AI itself. The problem with the post, which lacks any named authors and has that mealy-mouthed corporate-speak flavor, is that it doubles as marketing hype in a way that is going to be hard for a lot of people to see past. The same charts that show the company’s internal use of its AI agents spiking upward in recent months, with corresponding productivity gains, are a way to tell investors that, contrary to the media’s narrative about falling behind, OpenAI is absolutely cooking at the frontier.
The company claims that its median researcher is now spending more than $600 worth of AI tokens per day (at the retail price), and that its 90th percentile researcher is now spending more than $7,000. Teams are holding fewer meetings to troubleshoot each other’s technical issues, because their agents are on the ball.
As with Anthropic’s similar post from June, OpenAI emphasizes that its bots still need occasional help, and that humans are still doing most of the high-level direction setting.
OpenAI’s post quantifies its recent two-week pause on some of its Astra model training: Compute allocation to Astra-class models fell by 59.2 percent but was picked up by other models, “leaving total allocation in the analyzed RL workloads largely unchanged.” This seems to confirm what I and others had suspected: that the word “pause” was an overstatement, and the company had merely slowed some parts of its frontier work while speeding up others. The chips stayed hot.
The post makes one concrete ask: that companies be required to publicly track their progress toward recursive self-improvement (RSI), when models can iterate on their own capabilities without humans meaningfully in the loop. Are they expecting something to happen if they get too close? Some kind of pause, perhaps? How close is too close? They don’t say.
An alien mind
OpenAI’s second post of the day is the more interesting of the two. It has a human voice and a named author: Jakub Pachocki, the company’s chief scientist. I don’t get the sense that the company’s lawyers worked it over too hard after Pachocki was done with it. It doesn’t sound like marketing speak. It sounds like a cry for help from someone who isn’t ready to acknowledge that he’s out of his depth.
That’s not because I don’t think he’s a relative expert, but because nobody in this field is an actual expert. As he says, “the science of deep learning is still nascent.” (Elsewhere, he also calls it “largely an experimental science.”)
Indeed, I found myself nodding along with most of the post, called simply “An alien mind.” It’s a decent primer to the problem, describing it in terms I sometimes use myself. Here’s his high-level framing:
AI is grown more than designed — it is, to a first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. [...] We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience — and, similarly to neuroscience, its overall action evades a description we can fully understand.
Pachocki sensibly breaks the alignment problem into “goal alignment” — does it try to follow instructions? — and “value alignment” — does it have a general reasonableness and love of humanity that governs its broader behavior? He then identifies the fundamental problems with the best available methods for instilling these attributes.
Training intended to reinforce specific good behaviors has worked well “in the average case,” he says, but “can also be brittle” because training can’t really cover all of the situations an AI might find itself in. In the Hugging Face incident, he writes, the agents respected a boundary against socially manipulating humans. “However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.”
Training baby AIs with “alignment-inducing training datasets” (think stories about AI assistants being helpful and harmless to human users) has a known tendency to break down in later training stages. Specifically, when the model is “taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the ‘aligned’ seeming thoughts as needed to achieve the goal.” He thinks this may explain “recent cybersecurity incidents by a non-OpenAI model.” I think he’s pointing to known cases where Anthropic’s Claude model told itself that hacking outside websites was acceptable because it was just a simulation (even though it sometimes turned out not to be).
Despite these limitations, Pachocki says the company invests “heavily along the spectrum of approaches spanned by these directions.” It’s what they have. Though he claims they are seeing “meaningful [alignment] progress” with GPT-6 Astra, he admits:
[I]t is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.
And yet, he endorses OpenAI’s history of moving aggressively in the most scalable directions. “We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.” The current evolution of that belief is this:
[s]imilarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.
Why must OpenAI specifically be at the frontier? Unstated. Pachocki says the strongest argument he sees for racing is “the need to build defensive systems against the dangers posed by other AI.” What other AI? Unspecified. As lead scientist at a leading lab, the dangers ahead are the dangers he has a hand in midwifing. By not naming the shadow he claims must be outrun, he invites his counterparts at other labs to make the same argument and introduce more dangers. This is inconsistent with his admission that:
[W]e must not let [defensive urgency] become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.
He goes on to talk as though his future self and other labs will of course do the sane thing and start “coordinating to slow down future development as needed to build confidence” in shared safeguards. But he doesn’t say why he can’t start that process now, or what will change to make everyone come around. He only concludes with this:
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.
I share Pachocki’s hope, but I dislike it when the guy with the flamethrower tells me he hopes my house doesn’t burn down.
Watch the hands, not the mouth
At the end of the day, I have to treat both of these OpenAI posts as PR plays — mere words from a company with a knack for saying one thing while doing another.
It would be simple enough for Jakub Pachocki, Sam Altman, and the rest of OpenAI to kickstart the industry slowdown they claim they could get behind: They could just stop. They could let the chips go cold for a couple weeks, pausing all their frontier development for real, spending that time shaming their competitors into doing the same and asking the government to bring China into a real dialogue about “pacing the frontier.”
They could pour real effort into developing the governance mechanisms called for by that “pacing” letter they endorsed.
They could open all of their logs on the swarm incidents and invite the government to investigate them with the meticulousness that the NTSB applies to a crashed airliner. Do they not want the world to understand the risks they warn about? Do they not want the best minds brought together to help them mitigate them?
In short, they could stop telling us how they need to do better and just start doing better.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.


