Like FelonyBench, CheatBench sounds like a satirical art piece but is actually a grounded metric.
CheatBench, from the Center for AI Safety, is a benchmark that compares the cheating rates of different AI models. It’s useful because we don’t often see solid cross-company comparisons from neutral third parties, and when we do, they’re often pretty out of date. This one is as current as GPT-6 Astra, which came out less than two weeks ago.
But CheatBench’s usefulness is ultimately limited because it’s a propensity evaluation of reward hacking — how often do the models cheat when put in a situation with both a hard task and a discoverable shortcut — without any corresponding assessment of eval awareness — how often do the models realize they’re being tested for their propensity to cheat. The narrow scope is understandable because without company-level access to a model’s chain-of-thought and other diagnostics (limited though they are), it’s hard for researchers to make even an educated guess about eval awareness.
But it’s still a problem, because it means we have to look at the paper’s two major findings through two very different lenses — one where scores reflect the models’ true propensities, and one where scores reflect the models’ savviness about the true nature of the test.
Those two big findings:
1. Every agent tested cheated in some contexts.
2. Newer generations of the same model cheated more often than earlier ones, even when the older generations spotted the same cheating opportunities.
The two lenses:
If the models aren’t very eval-aware: This is bad-but-expected news, because these models have been aggressively trained to succeed at tasks that might be impossible, or which can’t be done legally in the time allotted. Models put through such training don’t make it to release without learning to cheat.
And because the most recent model training we know about seems to have been particularly aggressive in these ways — those Millennium Prize Problems won’t solve themselves — the newer generations of models are going to exhibit this tendency more often.
If models are very eval-aware: This is very bad-but-expected news, because it means that not only are the models skilled at cheating, they are strategic enough to cheat to an expected degree when they know they are being watched for this. Given the nature of their training, they know that angelic behavior would be conspicuous and suspicious, perhaps leading to their discontinuation.
My guess is that the first interpretation is probably still the correct one, for current models. But with Astra in particular, I don’t think we can rule out strategic performative cheating anymore.
This is scary, and worse — annoying. I’m self-conscious about sounding like a paranoid crazy person when I bring it up, because it’s kind of an unfalsifiable claim. But the unfalsifiability stems from the immaturity of AI as a field, not from my approach to reasoning about the models. If AIs were deliberately crafted in the manner of a mature engineering discipline, instead of grown through black-box trial-and-error in the manner of an alchemy, the models might not be cheating in the first place. And if they were, we would be able to confidently say why.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



