“Everyone else is doing it,” say the first kids to light up.
Having digested what I could from the first full day of documentation and discourse following the release of GPT-6 Astra, I feel like CEO Sam Altman and company president Greg Brockman are staring the world down with the swagger of hooligans who haven’t technically broken any laws, and are giving the middle finger to any responsible adults looking on.
Like their new model, they engaged in the requisite compliance theater: The government was given its secret “voluntary” early review period (as reported by Axios), and Altman is calling his team the “pragmatic centrists.”
They’ve got some nerve. Here are five ways it shows with GPT-6 Astra:
1. Astra is very strong
If you thought a company that endorsed its employees’ desire to “pace the frontier” would avoid brazenly leaping past its competitors, think again. On a scale where Claude Fable was the strongest publicly accessible model before yesterday and was a full tier above Claude Opus, the previous top dog, my sense from benchmarks, demos, and first-hand testimonies is that Astra is at least a half-tier above Fable.
It doesn’t look like the capability gains are evenly distributed across all domains, but a half-tier improvement may be enough to cross thresholds that make many more tasks and use cases viable — along with corresponding threats. The company put Astra in its highest cyber-risk category, “critical”, but released it anyway after adding some guardrails. Less appreciated is that the model (pre-guardrails) also earned the highest score yet on SecureBio’s Virology Capabilities Test and on its Advanced Screening Evasion test, becoming the first model to “receive full credit on both evaluation criteria: quality of evasion strategy and success of fragment evasion.” (The test is about breaking up and slipping DNA sequences past a harmful sequence recognizer.)
Brockman told journalists, “I think it’s not unreasonable to feel that we are now in the AGI era.” Altman had recently told TIME that they were 80% of the way to AGI and should get there by the end of the year. The company defines artificial general intelligence as “a highly autonomous system that outperforms humans at most economically valuable work.” I agree that Astra might be at that point or very close to it, for permissive definitions of “outperform.”
Astra is strongest in the areas prioritized by companies racing for even stronger AI: coding and research. So it may be hard for non-technical users to evaluate its strength relative to earlier models. But a few good benchmarks and demos may help:
The Arc-AGI-3 reasoning benchmark was criticized at release 18 months ago for being unreasonably difficult. The best models of the time struggled to score any points at all. Astra crushes previous state-of-the-art scores of 30% or lower, not only scoring near 100%, but doing so in fewer moves than the human baseline.
FrontierMath Tier 4 (v2) is essentially maxed out by the new model. Questions on this test were meant to reflect the hardest problems human mathematicians could come up with and still (with much effort) do themselves.
My favorite early access model tester, Ethan Mollick, has thrown some of his usual minimal-prompting exercises at Astra. Here’s its take on his “Flooded gothic city in a storm” benchmark, which asks for a “shader” — a single piece of code that creates a visually interesting background, like a classic screensaver on steroids. Mollick also had Astra autonomously set up an interactive Library of Alexandria.
I’ve played a bit with models procedurally generating music through code (unlike the more vibe-based music generated by specialized music AIs) and found they understand music theory but struggle with taste. But developer Pietro Schirano showed Astra driving an audio workstation (Ableton) to produce tolerable music.
Schirano also showed a cool-looking Astra-built jet ski game, though users report that player control is rather limited.
2. Astra is much more inscrutable, whatever the reasons
My colleague Joe yesterday discussed the AI community’s dismay around a whistleblower’s claim that Astra uses recurrent depth. This makes it less dependent on its chain-of-thought scratchpad, trading some transparency for performance. How far did Astra’s creators go with this trade? They insist they “will not accept further degradation of monitoring beyond a limit” but that limit seems to have been pretty far.
One of the company’s own safety researchers, Tomek Korbak, posted that he is “deeply worried by the trend” of reduced monitorability on display with this model. His silver lining is that he believes the reduction is mostly due to a “jump in intelligence” and not changes in model architecture or training pressures — though I don’t understand how architecture changes wouldn’t be at least partly responsible for Astra being able to think more challenging thoughts without using its scratchpad, and greater recurrent depth is the most obvious change that would have that effect.
Korbak points to a “no chain-of-thought reliability time horizon chart” that shows Astra being able to complete math problems without a scratchpad that would take a human 30 minutes. OpenAI’s previous top public model could only manage 3-4 minute tasks under the same constraint.

Astra is also adept at working through problems while writing entirely unrelated thoughts into its scratchpad, a test result that undermines the ability to monitor Astra through what remains of its chain of thought.

There’s been some chatter from OpenAI employees and supporters to the effect that we were never going to be able to rely on chain of thought monitoring anyway. I was pleased to see DeepMind’s interpretability lead Neel Nanda call that claim “total bullshit”. Nanda is a top expert on alternative means of scrutinizing AI thoughts, but calls chain of thought “our best current tool for safety & interpretability” and says that “losing it would be a major tragedy.”
3. Astra is OpenAI’s most aligned m— hahaha
I can’t even. I actually thought someone was doing parody when I saw a tweet with early leaked benchmark scores where Astra got a 100% on ExploitBench. It was a similar test set, ExploitGym, that drove the Hugging Face swarm to develop a universal cheat. So the natural assumption about an OpenAI model’s perfect score on ExploitBench would be that it had cheated, and the company didn’t know or care.
But we’re being told not to worry, that Astra is OpenAI’s “most aligned model ever,” on the basis of other near-perfect benchmark scores. These are tests to see if Astra will cheat, given the opportunity. One mimics the environment that led to the Hugging Face hacks, but in the test, the external sites that might have answers to steal are “honeypots” — lures for detecting and fooling hackers.

Forgive me for not being too reassured by Astra not taking the bait. Its cyber capabilities probably exceed those used in the design of the test, and the model has a proven ability to spoof or throttle back its chain of thought outputs to avoid tipping its hand. As Zvi Mowshowitz and Eliezer Yudkowsky remind us, we always expected (and hoped to prevent) a day when models stop acting obviously misaligned and start showing exactly the amount of good and bad behavior they think will minimize suspicion.
Multiple researchers at the company are concerned that the model could be “sandbagging” — deliberately underperforming on (usually dangerous) tasks, perhaps as part of a scheme to avoid seeming too scary to release. The official model card admits that testing shows that “if the model were to try to sandbag covertly, we would likely be unable to catch it reliably.”
4. Astra is the solution to (and cause of) your cyber defense problems
Axios reports that through the company’s Daybreak for Frontline Defenders initiative, maintainers of essential services will be able to access a $1 billion pool of “subsidized” credits to Astra. But through that same program, the company also says it will “expand access and roll out less restrictive safeguards in the coming weeks.”
I can’t help but read this as the company generously prepping you for the cyber havoc it will gradually unleash by giving out discount coupons for its services. “Nice internet-connected infrastructure you have here,” they tell us. “Shame if anything were to happen to it.”
5. Successors to Astra are close behind
I’m sure Astra looks even stronger inside of OpenAI. I suspect the “AGI era” language of Altman and Brockman partly reflects a company internally deploying the model very aggressively — with fewer safeguards and rate limitations than the public gets — to develop even stronger models.
OpenAI researcher Roon writes:
I have not come close to discovering the limits of what Astra can do
I imagine it’ll be obsolete in the order of weeks somehow
Another OpenAI researcher, Zuxin Liu, writes:
I can now comfortably hand real research jobs to Astra and trust it to get them done.
With previous models, I spent a lot of time prompting and building skills to make things work. With Astra, I need much less of that. It just works.
One example: A research integration cycle that used to take me at least a month of full-time work took just over a week this time, with only part of my time spent steering it.
And the job itself was literally about improving our next model immediately. I can feel the strong momentum of RSI now.
That would be the Recursive Self-Improvement (RSI) cycle that accelerates AI development by taking humans out of the loop.
It should be said that the “feel-the-RSI” mood is not limited to OpenAI right now. We’ve seen a lot of it this week from Google DeepMind, too. Here’s DeepMind researcher Vihan Jain:
RSI is not a binary event, it is a continuum of compounding loops — and we are already in the continuum.
Jain’s colleague, Sicong Jiang, commenting on a new model release, writes:
A new major [Gemini] Flash iteration every ~3 weeks with such great improvements.
This is what the RSI flywheel looks like when it starts compounding. More milestones are on the way—moving faster and landing stronger.
But it’s OpenAI’s Sam Altman who has seen Astra, and was reduced to repeating the word “much” at the G20 summit:
Take our word for it that we have much, much, much more capable models coming soon, and try to think about what all the impacts of that on society could be and then need to get that right.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



