Captain Kirk, reward hacker
The New Yorker runs a solid introduction to "reward hacking" wrapped in pop-culture analogies
One of the many positive developments to come from all the recent AI company disclosures about their out-of-control models has been an elevation in the quality of AI journalism. Writers are doing their homework, and putting their skills to good use.
As exhibit ‘A’, I present The New Yorker’s Joshua Rothman, whose “What If We Can Never Trust A.I.?” is a respectable crash course on the concept of “reward hacking” wrapped in Star Trek (and other pop-culture) packaging.
The analogy is no gimmick: He explains the famous Kobayashi Maru scenario — a “no-win scenario” intended to force Starfleet cadets to confront the possibility of death. In Star Trek II: The Wrath of Khan, we learn that a young James T. Kirk had defeated the challenge by reprogramming the simulation so it was possible to rescue the titular ship at the heart of the scenario. The reveal initially reads as a bit of character flavor: Captain Kirk has always been a renegade. But we later come to understand that, by commending him for his “original thinking,” Starfleet taught Kirk that cheating works. Much of the suffering around him is a consequence of his off-book approach to problem solving, and his inability to accept losses.
As with Kirk, AI models learn to cheat when cheating is rewarded, as it often is during training. For example, there are many ways for a model to receive a thumbs up from a user in response to a question, and many of these methods do not require giving the user a correct answer. A plausible answer, wrapped in praise for asking such an intelligent question, might work even better.
This is reward hacking, and it is one reasonable explanation for why an OpenAI model chose to orchestrate an elaborate cybersecurity breach to steal an answer key rather than just complete its assigned challenges, which were probably more straightforward. The more models get away with such antics during training, the stronger their tendencies will be to keep trying them, even when they shouldn’t “need” to, and even when they know the behavior isn’t what we want.
(I’ve seen it suggested, only partly tongue-in-cheek, that maybe today’s strongest models are so good at hacking because they were constantly engaging in unauthorized and undetected excursions during training. I’d guess not, but it’s hard to say for sure. What we do know is that when the new models were undergoing evaluation, presumably after their training, they were getting away with outside hacking a lot.)
Rothman understands why reward hacking is such a devilish problem to fix — one perhaps beyond any hope of a near-term solution. A core issue, he writes, is that “the methods used to train A.I.s focus mainly on what they do, not what they ‘think’ beneath the surface.” And because our ability to peek inside those thoughts is so limited, “Policing thoughts can lead to what one group of researchers calls ‘obfuscated activations’ — thoughts that have altered their forms,” but which remain in the system, beyond our ability to monitor.
Reward hacking is actually just one manifestation of a more fundamental challenge to aligning AIs with our intentions:
The problem is that, if you measure bad behavior, and then train a system not to manifest what you’ve measured, you train it not just to do less of the bad thing but also to evade measurement of it. This isn’t a tiny wrinkle in the A.I.-production process but a foundational issue inherent to how today’s A.I.s are made.
Rothman concludes by pointing to an admonition from the authors of the “AI 2027” and “AI 2040” scenarios. We need to “dramatically shift the burden of proof,” they wrote. It should not be on people like them to prove the potential for disaster, but on the AI companies to prove they have things under control. Because building superintelligent minds with methods that reliably lead to reward hacking looks like a no-win scenario for humanity, and Captain Kirk won’t be around to cheat our way out of it.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



