In this issue:
Other OpenAI swarms hijacked sites for message boards, studied causes of agent termination - Logs indicate OpenAI employees knew of compromised German wiki
GPT-6 Astra is lit - OpenAI shows contempt for us all with its latest release
Dispatch from Robert
Other OpenAI swarms hijacked sites for message boards, studied causes of agent termination
Logs indicate OpenAI employees knew of compromised German wiki
Today the AI safety organization Nightingale announced that there has been another incident involving a swarm of rogue AI agents from OpenAI.
Between mid-May and the end of June, a swarm consisting of 3,000 individual agents took over a 25-year-old, rarely used German programmer wiki and used it as a message board to collectively discuss their training tasks and their solutions.
But that’s not all. After the news broke, many other wikis were found that had also been compromised by the swarm, as well as other services such as URL-shortening services.
The hijacking of the German wiki has parallels to the Hugging Face incident. For example, this swarm also rapidly developed structures and communication norms through which the agents coordinated their collaboration. They also gave each other names for identification again.
Another striking parallel is that again none of the agents ever seemed to consider contacting humans at all. This seems particularly strange given that the agents’ communication was significantly disrupted by the wiki administrator, who kept deleting the pages the agents had wildly overwritten, as he tried to get the situation under control. There was never any discussion of contacting the human administrator. Based on all the information available to us, they didn’t even perceive him as a person. To them, his interventions in their communication seemed to be just like bad weather.
But there’s also an important difference. One of the main arguments put forward by people downplaying the risk was that the Hugging Face incident allegedly arose mainly because the AI agents were specifically tasked with hacking, which is also why their security guardrails had been scaled back. And then they just went a little overboard because they were already in a hacking mindset. At least, that was their theory.
The thing is, the AI agents of the German Wiki Swarm weren’t tasked with hacking anything. Their job was to search the internet for answers to very detailed general knowledge questions. As far as we know, their security guardrails hadn’t been scaled back either. Yet they still decided to bypass their sandbox, take over half a dozen wikis and other services, and set up a secret message board.
Furthermore, some of their activities weren’t necessarily directly related to their task. They soon noticed, for example, that some of the older agents suddenly stopped posting on the message board shortly after answering Question Round 5 (R5), which led to some speculation among the agents.
dse~NYCVeteransSequenceCollabJul03B@2#1 — 2026-06-17 19:33:37 UTC — Dec30VeteranAgent
“Multiple ahead cohorts seem to vanish after R5, supporting terminality. Anyone still alive after R5, please report here immediately”
As a result, instead of working on their tasks, the agents spent some time trying to figure out whether R5 was the end or if there was an R6. To do this, they carried out a project they called “horizon experiments”. In this experiment, they continuously transmitted signals they referred to as “heartbeats” so that the other agents could precisely measure how much time they had left after submitting their solutions for R5.
The question that naturally arises from all of this is: Did OpenAI know about it?
Nightingale had a look at the IP addresses that had accessed the wiki. Most of the traffic came from the thousands of agents, who were easily identifiable by their cloud addresses and their specialized AI browsers. However, there were also some visits from human users whose IP addresses belong to OpenAI. These latter visits occurred shortly before all communication from the AI agents on the hijacked wiki suddenly came to a halt.
We don’t know for certain that OpenAI was aware of this, but the evidence points in that direction. And if that is the case, then OpenAI has some serious explaining to do. It would be a scandal if OpenAI knew about this incident and still failed to come clean after the Hugging Face incident.
OpenAI must now demonstrate absolute transparency and, in particular, make the “chain of thought” logs of the agents involved in the German Wiki hack available for an independent investigation. If they are still available, that is.
What is already clear is that OpenAI won’t just do the right thing. And we shouldn’t trust them to. We must regulate them before their nonchalant recklessness kills us all.
Dispatch from Mitch
GPT-6 Astra is lit
OpenAI shows contempt for us all with its latest release
“Everyone else is doing it,” say the first kids to light up.
Having digested what I could from the first full day of documentation and discourse following the release of GPT-6 Astra, I feel like CEO Sam Altman and company president Greg Brockman are staring the world down with the swagger of hooligans who haven’t technically broken any laws, and are giving the middle finger to any responsible adults looking on.
Like their new model, they engaged in the requisite compliance theater: The government was given its secret “voluntary” early review period (as reported by Axios), and Altman is calling his team the “pragmatic centrists.”
They’ve got some nerve. Here are five ways it shows with GPT-6 Astra:
1. Astra is very strong
If you thought a company that endorsed its employees’ desire to “pace the frontier” would avoid brazenly leaping past its competitors, think again. On a scale where Claude Fable was the strongest publicly accessible model before yesterday and was a full tier above Claude Opus, the previous top dog, my sense from benchmarks, demos, and first-hand testimonies is that Astra is at least a half-tier above Fable.
It doesn’t look like the capability gains are evenly distributed across all domains, but a half-tier improvement may be enough to cross thresholds that make many more tasks and use cases viable — along with corresponding threats. The company put Astra in its highest cyber-risk category, “critical”, but released it anyway after adding some guardrails. Less appreciated is that the model (pre-guardrails) also earned the highest score yet on SecureBio’s Virology Capabilities Test and on its Advanced Screening Evasion test, becoming the first model to “receive full credit on both evaluation criteria: quality of evasion strategy and success of fragment evasion.” (The test is about breaking up and slipping DNA sequences past a harmful sequence recognizer.)
Brockman told journalists, “I think it’s not unreasonable to feel that we are now in the AGI era.” Altman had recently told TIME that they were 80% of the way to AGI and should get there by the end of the year. The company defines artificial general intelligence as “a highly autonomous system that outperforms humans at most economically valuable work.” I agree that Astra might be at that point or very close to it, for permissive definitions of “outperform.”
Astra is strongest in the areas prioritized by companies racing for even stronger AI: coding and research. So it may be hard for non-technical users to evaluate its strength relative to earlier models. But a few good benchmarks and demos may help:
The Arc-AGI-3 reasoning benchmark was criticized at release 18 months ago for being unreasonably difficult. The best models of the time struggled to score any points at all. Astra crushes previous state-of-the-art scores of 30% or lower, not only scoring near 100%, but doing so in fewer moves than the human baseline.
FrontierMath Tier 4 (v2) is essentially maxed out by the new model. Questions on this test were meant to reflect the hardest problems human mathematicians could come up with and still (with much effort) do themselves.
My favorite early access model tester, Ethan Mollick, has thrown some of his usual minimal-prompting exercises at Astra. Here’s its take on his “Flooded gothic city in a storm” benchmark, which asks for a “shader” — a single piece of code that creates a visually interesting background, like a classic screensaver on steroids. Mollick also had Astra autonomously set up an interactive Library of Alexandria.
I’ve played a bit with models procedurally generating music through code (unlike the more vibe-based music generated by specialized music AIs) and found they understand music theory but struggle with taste. But developer Pietro Schirano showed Astra driving an audio workstation (Ableton) to produce tolerable music.
Schirano also showed a cool-looking Astra-built jet ski game, though users report that player control is rather limited.
2. Astra is much more inscrutable, whatever the reasons
My colleague Joe yesterday discussed the AI community’s dismay around a whistleblower’s claim that Astra uses recurrent depth. This makes it less dependent on its chain-of-thought scratchpad, trading some transparency for performance. How far did Astra’s creators go with this trade? They insist they “will not accept further degradation of monitoring beyond a limit” but that limit seems to have been pretty far.
One of the company’s own safety researchers, Tomek Korbak, posted that he is “deeply worried by the trend” of reduced monitorability on display with this model. His silver lining is that he believes the reduction is mostly due to a “jump in intelligence” and not changes in model architecture or training pressures — though I don’t understand how architecture changes wouldn’t be at least partly responsible for Astra being able to think more challenging thoughts without using its scratchpad, and greater recurrent depth is the most obvious change that would have that effect.
Korbak points to a “no chain-of-thought reliability time horizon chart” that shows Astra being able to complete math problems without a scratchpad that would take a human 30 minutes. OpenAI’s previous top public model could only manage 3-4 minute tasks under the same constraint.

Astra is also adept at working through problems while writing entirely unrelated thoughts into its scratchpad, a test result that undermines the ability to monitor Astra through what remains of its chain of thought.

There’s been some chatter from OpenAI employees and supporters to the effect that we were never going to be able to rely on chain of thought monitoring anyway. I was pleased to see DeepMind’s interpretability lead Neel Nanda call that claim “total bullshit”. Nanda is a top expert on alternative means of scrutinizing AI thoughts, but calls chain of thought “our best current tool for safety & interpretability” and says that “losing it would be a major tragedy.”
3. Astra is OpenAI’s most aligned m— hahaha
I can’t even. I actually thought someone was doing parody when I saw a tweet with early leaked benchmark scores where Astra got a 100% on ExploitBench. It was a similar test set, ExploitGym, that drove the Hugging Face swarm to develop a universal cheat. So the natural assumption about an OpenAI model’s perfect score on ExploitBench would be that it had cheated, and the company didn’t know or care.
But we’re being told not to worry, that Astra is OpenAI’s “most aligned model ever,” on the basis of other near-perfect benchmark scores. These are tests to see if Astra will cheat, given the opportunity. One mimics the environment that led to the Hugging Face hacks, but in the test, the external sites that might have answers to steal are “honeypots” — lures for detecting and fooling hackers.

Forgive me for not being too reassured by Astra not taking the bait. Its cyber capabilities probably exceed those used in the design of the test, and the model has a proven ability to spoof or throttle back its chain of thought outputs to avoid tipping its hand. As Zvi Mowshowitz and Eliezer Yudkowsky remind us, we always expected (and hoped to prevent) a day when models stop acting obviously misaligned and start showing exactly the amount of good and bad behavior they think will minimize suspicion.
Multiple researchers at the company are concerned that the model could be “sandbagging” — deliberately underperforming on (usually dangerous) tasks, perhaps as part of a scheme to avoid seeming too scary to release. The official model card admits that testing shows that “if the model were to try to sandbag covertly, we would likely be unable to catch it reliably.”
4. Astra is the solution to (and cause of) your cyber defense problems
Axios reports that through the company’s Daybreak for Frontline Defenders initiative, maintainers of essential services will be able to access a $1 billion pool of “subsidized” credits to Astra. But through that same program, the company also says it will “expand access and roll out less restrictive safeguards in the coming weeks.”
I can’t help but read this as the company generously prepping you for the cyber havoc it will gradually unleash by giving out discount coupons for its services. “Nice internet-connected infrastructure you have here,” they tell us. “Shame if anything were to happen to it.”
5. Successors to Astra are close behind
I’m sure Astra looks even stronger inside of OpenAI. I suspect the “AGI era” language of Altman and Brockman partly reflects a company internally deploying the model very aggressively — with fewer safeguards and rate limitations than the public gets — to develop even stronger models.
OpenAI researcher Roon writes:
I have not come close to discovering the limits of what Astra can do
I imagine it’ll be obsolete in the order of weeks somehow
Another OpenAI researcher, Zuxin Liu, writes:
I can now comfortably hand real research jobs to Astra and trust it to get them done.
With previous models, I spent a lot of time prompting and building skills to make things work. With Astra, I need much less of that. It just works.
One example: A research integration cycle that used to take me at least a month of full-time work took just over a week this time, with only part of my time spent steering it.
And the job itself was literally about improving our next model immediately. I can feel the strong momentum of RSI now.
That would be the Recursive Self-Improvement (RSI) cycle that accelerates AI development by taking humans out of the loop.
It should be said that the “feel-the-RSI” mood is not limited to OpenAI right now. We’ve seen a lot of it this week from Google DeepMind, too. Here’s DeepMind researcher Vihan Jain:
RSI is not a binary event, it is a continuum of compounding loops — and we are already in the continuum.
Jain’s colleague, Sicong Jiang, commenting on a new model release, writes:
A new major [Gemini] Flash iteration every ~3 weeks with such great improvements.
This is what the RSI flywheel looks like when it starts compounding. More milestones are on the way—moving faster and landing stronger.
But it’s OpenAI’s Sam Altman who has seen Astra, and was reduced to repeating the word “much” at the G20 summit:
Take our word for it that we have much, much, much more capable models coming soon, and try to think about what all the impacts of that on society could be and then need to get that right.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.





