In this issue:
Small team exposes massive gap in OpenAI cybersecurity - With AI help and a few hours of human effort, three researchers cracked the trillion-dollar firm wide open
CEOs to dine on matters of state - Altman and Huang are the wrong people for a US-China AI summit
Curated introductions to AI risk - New website spotlights especially good articles and videos at just the right lengths
Anthropic’s internal research automation is progressing - Anthropic published data about the state of their internal AI R&D automation. It looks concerning.
Dispatches from Joe
Small team exposes massive gap in OpenAI cybersecurity
With AI help and a few hours of human effort, three researchers cracked the trillion-dollar firm wide open
OpenAI, long fond of treating rapidly accelerating AI capabilities as inevitable, has argued that “the best way to reduce national risk” is to use its models to defend against malicious AI. Naturally, we ought to expect them to practice what they preach, so OpenAI should have some of the best cyber-defenses in the world by now. Right?
Not so much. Yesterday, we learned that a small team at startup Hacktron AI managed a chain of exploits that gave them widespread access to OpenAI employee accounts. Aided by Anthropic’s Claude AI, three cybersecurity researchers found both a bug in common image decoder software and a critical flaw in the system OpenAI uses to manage employee account identities. (For business users in the audience who might recognize the jargon, it was the company’s “single sign-on,” or SSO, that was malformed. Yes, it’s as bad as it sounds.)
The Hacktron AI team combined these and other vulnerabilities to crack open the trillion-dollar firm, demonstrating the gap with a harmless proof of concept hack and earning a $6,500 bounty in the process. Frankly, I think they deserve at least ten times that, and maybe a medal.
The Guardian covered the story today, but the dry news article doesn’t really do it justice. As the Hacktron team explains:
Until two months ago, any user or OpenAI employee logging into OpenAI’s own help forum could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails.
The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours.
“Repo” is short for “repository”, a programming term for data and code storage. Why does it matter that a repo was hacked? Well, repos can contain highly sensitive material. Not only could the Hacktron team have leaked files from Slack and email, but they could also have stolen research files — and no doubt sold them to rival companies or Chinese AI developers for much more than $6,500.
Attacks such as this pose a massive threat to the security of companies, nations, and the world as a whole. They could let rival companies, spy agencies, or rogue AIs insert covert backdoors in critical systems. They could enable theft of vast troves of user data and a corresponding rise in fraud and downstream cybercrimes. They could let OpenAI’s own models break out (again) or let malicious AIs manipulate training. They could leak algorithmic improvements — secret techniques that make AI more efficient — to rivals, further accelerating the already frantic AI race.
That OpenAI apparently couldn’t spare the handful of person-hours and AI tokens it took to discover this gap, or worse, that it doesn’t employ security staff competent enough to do so, is an appalling indictment of its security regime, given the stakes.
You’d think that this exploit chain involved state-of-the-art hacking AI. You’d be wrong. It was at least one step shy of that, a mix of Anthropic’s Claude Opus 4.8 and Opus 5, which started contributing to the hack hours after its July 24 launch. Mythos, Anthropic’s most powerful known hacker, was not involved, most likely because the consumer-facing version (Fable) aggressively screens out cyber requests.
As I understand it, OpenAI’s GPT-5.6 Sol helped the Hacktron AI team find exploits for the image-decoder bug on a much broader scale, but only after the initial chain of exploits against OpenAI.
Putting the pieces together: These attacks were enabled by consumer-grade AI models available to the general public. It’s possible that, well before Hacktron AI quietly did OpenAI’s job for it, Chinese hackers had used the same method to lift all sorts of secrets from OpenAI accounts. There are probably more holes like this one as well.
For a company that talks a big game about “strengthening cyber resilience” worldwide, OpenAI doesn’t seem to be walking the walk.
CEOs to dine on matters of state
Altman and Huang are the wrong people for a US-China AI summit
Earlier this week, I wrote about the bipartisan coalition of concerned citizens and politicians that met at the Pro-Human Assembly in DC. Disparate factions, often wildly opposed on most issues, found common ground in demanding that the American people, not “tech oligarchs”, steer the future.
For the upcoming AI summit, it’s looking like the White House may favor the oligarchs.
Presidents Trump and Xi Jinping will meet next week to discuss the future of AI. A formal state dinner is planned, and POLITICO reports that AI industry CEOs including OpenAI’s Sam Altman and NVIDIA’s Jensen Huang will attend. (It’s unclear whether Dario Amodei of Anthropic was invited.)
Their attendance would bode ill for the chances of an international agreement to stop or slow the AI race; the CEOs in attendance have a vested interest in preventing or weakening such an agreement. Altman runs one of the AI firms that’s been pushing the frontier, and Huang has dismissed the danger AI poses to our species as “hysterical.” He also vocally advocates the sale of AI chips to China, which would accelerate the AI race, advantage a U.S. strategic rival, and nicely bolster NVIDIA’s bottom line.
The announcement left a disgusting taste in my mouth, but the week’s political news isn’t all bad. The House committee that advises on strategic competition with China announced an upcoming hearing on “the risks of AI, best practices in AI governance, and the path forward toward binding international safeguards to ensure safety and human control of this technology.”
The hearing is scheduled for September 23, a day before the US-China summit, at the behest of Democrat Ro Khanna. I commend Khanna for taking a necessary step forward, even as I worry that moves like this may entrench partisan feelings on the AI extinction threat.
As my colleague Alana wrote on Wednesday, Americans of both parties fear AI-driven extinction and favor a halt to AI development. The Trump administration has long kept a finger on the pulse of popular opinion, and it is my enduring hope that the voice of the many will yet drown out the whispers of the few.
Curated introductions to AI risk
New website spotlights especially good articles and videos at just the right lengths
What was once a trickle of AI-related introductory content has grown into a flood, and while I’m inspired by the creativity and dedication of concerned creators, it can be overwhelming at times.
That’s why I’m excited to see AGI.FYI, a new website that curates this introductory content and recommends the clearest and most accurate writeups and videos out there, in the judgment of a small team of experts who’ve been paying close attention for quite some time.
Right now, the team’s curated lists include their favorite articles about the Hugging Face incident, videos on why AI is dangerous, and light, medium, and heavy reading on why AI could kill everyone on Earth. It’s not a perfect collection — their top short-piece pick for understanding the extinction threat is unfortunately paywalled — but I still recommend checking out the videos and the deeper dives if either medium appeals.
They are not the only resource for reading material — see also BlueDot Impact and aisafety.info — but no one else I know of is doing quite the same thing. Consider checking it out, or sharing with friends and family, today.
Dispatch from Robert
Anthropic’s internal research automation is progressing
Anthropic published data about the state of their internal AI R&D automation. It looks concerning.
AI companies are talking a lot about “pacing the frontier,” but what exactly is the current pace at the frontier? Anthropic published a post yesterday that gives us some figures on this.
One of the most striking figures is probably that, between February and August of this year alone, the share of its own research and development driven primarily by AI rose from less than 1% to 26%.
Anthropic uses Epoch AI’s so-called “automation rating scale,” which distinguishes between various “automation levels” (AL). At the current level of capability, the relevant levels are AL3 (AI collaborates), AL4 (AI conducts research and development largely independently) and AL5 (AI conducts research completely on its own).
The aforementioned 26% is AL4, the “largely independent” level. Anthropic also reports that in 90% of work, AI at least collaborates with human researchers.
Even though there are no fully autonomous AL5 AI agents conducting research yet, these figures make it clear why Anthropic is concerned about the speed of development.
Since the rapid increase in the proportion of AL4 work began around March, it stands to reason that the first model capable of conducting AL4 research was Mythos Preview, which was released for internal use at Anthropic on February 24. That was only about half a year ago, and it seems plausible that every step forward is accelerating development even further.
If we project this development half a year or a year into the future, it seems reasonable to predict that we’re rapidly approaching the point at which AI can actually produce ever-faster, ever-better successors fully automatically, without the need for human input. Experts call this “recursive self-improvement,” and it is dangerous because it can lead to a kind of “feedback loop” that can no longer be controlled. And yet Anthropic and others are reaching for this milestone on purpose.
Further figures reported by Anthropic show that this process is already difficult to control. Not only is the quality of the models continuing to rise, but the sheer quantity has now reached breathtaking levels. In August 2026, approximately 30,000 AI agents were working simultaneously on Anthropic’s most widely used internal platform alone. By way of comparison: Anthropic has no more than about 5,000 human employees, only a fraction of whom are working in research.
We know from studies of the Hugging Face swarm and the German Wiki swarm that so many individual agents can produce an almost unbelievable volume of text output in a short time. How does Anthropic keep track of all this?
Anthropic uses AI to monitor its AI. Between 50 and 100 million transcripts are monitored each week, and of these, 100,000 are flagged for closer scrutiny — which is also carried out by AI. Of these 100,000 transcripts, only about 50 are evaluated by a human. So a human only sees about one transcript in a million.
This answers the question of how Anthropic keeps track of all this: It doesn’t. It hopes its AIs will maintain oversight almost on their own.
We know from recent incidents involving AI swarms and from Anthropic’s own research that agents are prone to conspiring to cheat and circumvent security measures. We have no good reason to assume things are any different with the AIs that Anthropic uses internally.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.








