No-win scenario
Math breakthroughs, reward hacking in pop culture, Google Earth deepfakes, and more
In this issue:
Ten advances to shake the math world - OpenAI seems to have broken new ground on capabilities, but at what cost?
Captain Kirk, reward hacker - The New Yorker runs a solid introduction to “reward hacking” wrapped in pop-culture analogies
A crater where the Eiffel Tower should be - Journalists enthusiastically abuse new Google Earth feature to sound deepfake alarm
We are the books - Destructive book scanning is a bad reason to hate the AI companies, but a great metaphor for where this leads
Dispatches from Mitch
Ten advances to shake the math world
OpenAI seems to have broken new ground on capabilities, but at what cost?

A team of researchers from OpenAI this morning announced “Ten advances in mathematics and theoretical computer science.”
With my usual disclaimer that I am not a mathematician, it looks like these new discoveries are sufficiently formalized such that their validity isn’t really in question. But at the same time, as with so much frontier AI output these days, hardly anyone is equipped to fairly evaluate their novelty and importance. It may take a little while for elite mathematicians to look things over and decide whether to applaud or sneer.
For what it’s worth, Anthropic’s Claude Fable thinks OpenAI’s list of findings is a big deal, saying:
[N]early every item is a famous, decades-old problem, and several are considered flagship problems of their entire field [...] any single one of these, done by someone under 40, would plausibly anchor a medal case [...] A list of ten of them is not a plausible human thesis [...] No mathematician in history has a run like this.
I predict that the usual dismissal — that the AIs merely mechanically and methodically searched for solutions in ways too tedious for humans to bother with — won’t hold up. It didn’t hold up for the disproof of the Jacobian conjecture I covered less than two weeks ago: Terence Tao, sometimes considered the greatest living mathematician, subsequently wrote that “the construction presented in this fashion appears like a massive miracle,” and not the sort of thing that could be plausibly located “by brute force.”
Given OpenAI’s recent disclosures, with hints from Reuters yesterday that more have yet to come out, when I consider the new math discoveries, I am compelled to ask, “but at what cost?” The latest work was almost certainly done by one or more close cousins of the model that disproved the Erdős unit-distance conjecture in May — a model the company later described as having broken out of its sandbox to post evaluation results on the open internet without permission. That model or a close variant is itself understood to have been behind the autonomous attacks on Hugging Face.
OpenAI has pushed the math frontier by pushing the AI capabilities frontier in a series of reckless, poorly monitored experiments. Along the way, they learned the predictable lesson experts have long warned about — that (in the company’s anodyne phrasing) “Long-running models can solve difficult open-ended problems, but their persistence gives them more opportunities to take unwanted actions.”
The new math announcement makes no mention of these costs, instead pointing to the models’ token consumption and saying that at the rates the company charges for its best public model right now, Sol, the discoveries would have cost about $2,000 total.
That has to be an underestimate, given that the new model tier, called Astra, is almost certainly much larger, and will come with commensurately higher per-token operating costs. But if we conservatively guess that Astra costs about ten times as much to run as Sol, we’re talking about ten new discoveries for a mere $2,000 each, which is still, frankly, ridiculous — a clear sign that the world as we know it is about to give way, whether or not we halt further AI progress to look for a way around the extinction problem.
The announcement post closes with an explanation of the roles played by humans and AIs in the new discoveries, along with an acknowledgement of the concerns about AI’s growing impact on the field of mathematics. It links to the Leiden declaration, a lengthy open petition of mathematicians calling for transparency in the use of AI and resistance to reliance on unverified AI-generated proofs.
OpenAI’s post is accompanied by a link to a formal paper, as well as a companion paper about “How the Ideas Came Together.” That companion piece was itself written by AI, working from the notes of the AIs that worked through the discoveries.
Today’s announcement is sure to shake math experts who were already suffering from a crisis of meaning as AI intrudes on a source of joy and makes it harder to trust what they see. They certainly aren’t the first to have joined this club, and they won’t be the last.
Captain Kirk, reward hacker
The New Yorker runs a solid introduction to “reward hacking” wrapped in pop-culture analogies
One of the many positive developments to come from all the recent AI company disclosures about their out-of-control models has been an elevation in the quality of AI journalism. Writers are doing their homework, and putting their skills to good use.
As exhibit ‘A’, I present The New Yorker’s Joshua Rothman, whose “What If We Can Never Trust A.I.?” is a respectable crash course on the concept of “reward hacking” wrapped in Star Trek (and other pop-culture) packaging.
The analogy is no gimmick: He explains the famous Kobayashi Maru scenario — a “no-win scenario” intended to force Starfleet cadets to confront the possibility of death. In Star Trek II: The Wrath of Khan, we learn that a young James T. Kirk had defeated the challenge by reprogramming the simulation so it was possible to rescue the titular ship at the heart of the scenario. The reveal initially reads as a bit of character flavor: Captain Kirk has always been a renegade. But we later come to understand that, by commending him for his “original thinking,” Starfleet taught Kirk that cheating works. Much of the suffering around him is a consequence of his off-book approach to problem solving, and his inability to accept losses.
As with Kirk, AI models learn to cheat when cheating is rewarded, as it often is during training. For example, there are many ways for a model to receive a thumbs up from a user in response to a question, and many of these methods do not require giving the user a correct answer. A plausible answer, wrapped in praise for asking such an intelligent question, might work even better.
This is reward hacking, and it is one reasonable explanation for why an OpenAI model chose to orchestrate an elaborate cybersecurity breach to steal an answer key rather than just complete its assigned challenges, which were probably more straightforward. The more models get away with such antics during training, the stronger their tendencies will be to keep trying them, even when they shouldn’t “need” to, and even when they know the behavior isn’t what we want.
(I’ve seen it suggested, only partly tongue-in-cheek, that maybe today’s strongest models are so good at hacking because they were constantly engaging in unauthorized and undetected excursions during training. I’d guess not, but it’s hard to say for sure. What we do know is that when the new models were undergoing evaluation, presumably after their training, they were getting away with outside hacking a lot.)
Rothman understands why reward hacking is such a devilish problem to fix — one perhaps beyond any hope of a near-term solution. A core issue, he writes, is that “the methods used to train A.I.s focus mainly on what they do, not what they ‘think’ beneath the surface.” And because our ability to peek inside those thoughts is so limited, “Policing thoughts can lead to what one group of researchers calls ‘obfuscated activations’ — thoughts that have altered their forms,” but which remain in the system, beyond our ability to monitor.
Reward hacking is actually just one manifestation of a more fundamental challenge to aligning AIs with our intentions:
The problem is that, if you measure bad behavior, and then train a system not to manifest what you’ve measured, you train it not just to do less of the bad thing but also to evade measurement of it. This isn’t a tiny wrinkle in the A.I.-production process but a foundational issue inherent to how today’s A.I.s are made.
Rothman concludes by pointing to an admonition from the authors of the “AI 2027” and “AI 2040” scenarios. We need to “dramatically shift the burden of proof,” they wrote. It should not be on people like them to prove the potential for disaster, but on the AI companies to prove they have things under control. Because building superintelligent minds with methods that reliably lead to reward hacking looks like a no-win scenario for humanity, and Captain Kirk won’t be around to cheat our way out of it.
A crater where the Eiffel Tower should be
Journalists enthusiastically abuse new Google Earth feature to sound deepfake alarm
Google seems to have carelessly and unintentionally advanced the deepfake frontier on Thursday when it integrated Nano Banana, its AI image generator, into Google Earth, its comprehensive satellite map.
The Atlantic’s Matteo Wong followed Google’s instructions to “pick a spot on the map, and start bringing your ideas to life,” setting office buildings on fire, inserting a crater where the Eiffel Tower should be, and putting a homeless encampment on the White House lawn.
Google said yesterday that it “takes misinformation seriously” and was rolling back the feature after users — presumably not all journalists, though they ironically seem to have been the most enthusiastic early adopters — used the tool to generate imagery of events that would be newsworthy, if real. NPR said it was able to “easily generate images of Iran’s Kharg Island on fire, and a flooded U.S. Capitol complex.” BBC described generating “a sinkhole swallowing the Great Pyramid of Giza and Russian tanks in Ukraine’s capital.”
The problem is less that satellite imagery is so easily doctored and more that, until now, it wasn’t commonly done. People aren’t yet used to questioning imagery purportedly taken from orbit. This will have to change. Given all the press, I doubt bad guys will refrain from the tactic just because Google has stopped making it quite so easy.
We are the books
Destructive book scanning is a bad reason to hate the AI companies, but a great metaphor for where this leads

Because I’ve seen this 404 Media story go around and get turned into conspiracy videos, let’s talk about the way AI companies are buying used books in bulk, destructively scanning them for training data, and discarding the remnants.
While the companies appear to not be very open about their book operations, there’s not really anything new or shady going on here, so far as I can tell. Destructively scanning books is a very old practice — much older than the AI companies. Scanning is best done with flat pages, which are hard to get while the book is bound. So to get quality scans, it’s best to remove the pages from their binding.
If you only wanted a book for the scan, there’s little point in putting it back together after, which would take time and money. Moreover, you would then have to do something with the reassembled book. Store it forever? Why? Sell it? Who will want your Franken-book when they could buy a never-disassembled one?
And if you’re an AI company, reselling the book might jeopardize your legal safe harbor. Courts have determined that using other people’s books as AI training data without permission or compensation is fair game, but only if you actually bought the book. I don’t think courts that busted AI companies for pirating books would look kindly on arguments that resold books were temporarily owned at time of scan.
Yes, as you may have heard, books from before the AI era are perhaps valued a little more because they won’t themselves be the product of AI. But no, companies are almost certainly not buying up the only copies of rare books, nor are they trying to remove books from circulation. This is not a conspiracy to monopolize the past and rewrite history, and if it were, this would be far from the biggest threat posed by these companies.
If you want to worry about where the book scanning leads, know that it feeds the race to build artificial minds that would treat humanity the way the AI companies treat books. Books aren’t being destroyed out of any hatred for the printed word, but out of indifference to the format. Companies are extracting only what they care about and discarding the rest.
What would superintelligences want from us? Not much, and not for long. Some physical hands to help bootstrap a more efficient and automated supply chain, maybe. Once that’s up and running, you should expect to be worth about as much as the empty spine of an old pulp novel on eBay.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.




