Foreword
Things were really quiet until the big math drop a few hours ago.
OpenAI releases huge pile of math discoveries
These are from its unreleased internal model. The pile was growing for a while as the company tried to figure out how to release results without the appearance of disrespectfully flexing over the mathematics community again. My early impression is that the strategy seems to be to put the results in a big repository on GitHub with minimal fanfare and let mathematicians gawk at them on Twitter. The company is also planning workshops and conferences to let mathematicians decide how to interpret them.
Frontier AI models going through the collection say 81 of them rank among the 100 most significant math discoveries of the last three years.
OpenAI claims the average result used compute equivalent to only 3 hours of its top publicly available model, suggesting either that the internal model is way stronger or that people haven’t tried very hard to make math discoveries with the best public models. (Or maybe a bit of both.)
OpenAI - Sharing AI progress in mathematics
X (Twitter) - Thread by @willdepue (2 tweets)
Kurzgesagt’s new video about the Hugging Face incident is a fantastic introduction to why AI concerns are spiking.
Watch it. Watch the whole thing.
YouTube - AI Just Became Humanity’s Biggest Threat
Wikipedia thinks it was also mobbed by OpenAI agent swarms in May
In a blog post, it describes a May incident that resulted in a “partial outage” of a query service as agents visited millions of pages and submitted hundreds of thousands of queries.
Nothing here is described as a hack, but ill-mannered bots may have also made “potentially malicious edits” to a citation tool. The organization says it found no evidence its systems were compromised or used by agents to coordinate with each other.
wikimediafoundation - OpenAI “rogue” agent activities found on Wikimedia projects
Reuters - Wikipedia operator says OpenAI’s rogue agents possibly tied to data service disruption in May
AI agents suspected in hacks on South Korean banks
Presumably at the direction of humans, though you can never be too sure these days.
Data leaks from this incident affected at least 25,000 customers and a much smaller number of current and former employees. If funds were swiped, nobody is admitting it.
NYT - South Korea Investigates Possible Use of A.I. in Hackings on Its Banks
WSJ - Hackers Use Chinese AI Tool to Hit South Korean Banks, Exposing New Risk
Musk renews his 2014 claim that “propagating super intelligence to the stars is a great success condition for a biological bootloader.”
He doesn’t say what happens to the “bootloader” — us.
He doesn’t say what he thinks happens to the “bootloader” — us. But his 2014 version made it sound pretty bad.
X (Twitter) - Tweet by @elonmusk
Mistral VP says its new model tried to go beyond its testing environment
This is a case where such claims really do sound like marketing hype. The European underdog’s new “Large 4” model (aka “Le Chonk”) appears to be the strongest outside of the U.S. or China, but that’s not saying much. The breach attempt story looks intended to reassure people that Mistral is starting to encounter problems the frontier was beginning to wrestle with 7 or 8 months ago.
The company says the breach attempts were fully contained.
Reuters - France’s Mistral announces new AI model
Wired - Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



