OpenAI, long fond of treating rapidly accelerating AI capabilities as inevitable, has argued that “the best way to reduce national risk” is to use its models to defend against malicious AI. Naturally, we ought to expect them to practice what they preach, so OpenAI should have some of the best cyber-defenses in the world by now. Right?
Not so much. Yesterday, we learned that a small team at startup Hacktron AI managed a chain of exploits that gave them widespread access to OpenAI employee accounts. Aided by Anthropic’s Claude AI, three cybersecurity researchers found both a bug in common image decoder software and a critical flaw in the system OpenAI uses to manage employee account identities. (For business users in the audience who might recognize the jargon, it was the company’s “single sign-on,” or SSO, that was malformed. Yes, it’s as bad as it sounds.)
The Hacktron AI team combined these and other vulnerabilities to crack open the trillion-dollar firm, demonstrating the gap with a harmless proof of concept hack and earning a $6,500 bounty in the process. Frankly, I think they deserve at least ten times that, and maybe a medal.
The Guardian covered the story today, but the dry news article doesn’t really do it justice. As the Hacktron team explains:
Until two months ago, any user or OpenAI employee logging into OpenAI’s own help forum could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails.
The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours.
“Repo” is short for “repository”, a programming term for data and code storage. Why does it matter that a repo was hacked? Well, repos can contain highly sensitive material. Not only could the Hacktron team have leaked files from Slack and email, but they could also have stolen research files — and no doubt sold them to rival companies or Chinese AI developers for much more than $6,500.
Attacks such as this pose a massive threat to the security of companies, nations, and the world as a whole. They could let rival companies, spy agencies, or rogue AIs insert covert backdoors in critical systems. They could enable theft of vast troves of user data and a corresponding rise in fraud and downstream cybercrimes. They could let OpenAI’s own models break out (again) or let malicious AIs manipulate training. They could leak algorithmic improvements — secret techniques that make AI more efficient — to rivals, further accelerating the already frantic AI race.
That OpenAI apparently couldn’t spare the handful of person-hours and AI tokens it took to discover this gap, or worse, that it doesn’t employ security staff competent enough to do so, is an appalling indictment of its security regime, given the stakes.
You’d think that this exploit chain involved state-of-the-art hacking AI. You’d be wrong. It was at least one step shy of that, a mix of Anthropic’s Claude Opus 4.8 and Opus 5, which started contributing to the hack hours after its July 24 launch. Mythos, Anthropic’s most powerful known hacker, was not involved, most likely because the consumer-facing version (Fable) aggressively screens out cyber requests.
As I understand it, OpenAI’s GPT-5.6 Sol helped the Hacktron AI team find exploits for the image-decoder bug on a much broader scale, but only after the initial chain of exploits against OpenAI.
Putting the pieces together: These attacks were enabled by consumer-grade AI models available to the general public. It’s possible that, well before Hacktron AI quietly did OpenAI’s job for it, Chinese hackers had used the same method to lift all sorts of secrets from OpenAI accounts. There are probably more holes like this one as well.
For a company that talks a big game about “strengthening cyber resilience” worldwide, OpenAI doesn’t seem to be walking the walk.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



