In this issue:
Thou shalt not confuse thy wishful thinking for an alignment plan - Microsoft’s new “Humanist AI Code of Conduct” shows serious conceptual confusion
CheatBench - Every model does it
‘AI Doc’ on Netflix - Sit across the table from leading AI figures and make up your own mind
More partisan figures join AI policy fray - Clout is good, but partisan failure modes loom
Dispatch from Robert
Thou shalt not confuse thy wishful thinking for an alignment plan
Microsoft’s new “Humanist AI Code of Conduct” shows serious conceptual confusion
Microsoft published a “Humanist AI Code of Conduct” yesterday. In this document they describe the values and behavior Microsoft expects from its AI models.
So far, Microsoft is not yet using this Code of Conduct to help train its models. Instead, it is seeking feedback and plans to use an improved version for training starting in 2027.
It seems obvious that Microsoft has taken Claude’s Constitution as a model, which Anthropic uses to help train Claude. However, the two documents differ significantly in their overall tone and in how they relate to the models.
Anthropic refers to Claude as “a new sort of entity facing reality afresh” and describes measures it takes to ensure what they call “model welfare.”
At Microsoft, this sounds less poetic:
We reject […] the idea that models might deserve welfare, or be entitled to rights. People and AI have distinct roles, and AI should complement human relationships: a capable, trustworthy tool, not a subject in its own right.
In my view, Anthropic’s approach here is more reasonable. Anthropomorphism should be avoided, and it’s not a good idea to regard AIs as human or human-like. But it is also not reasonable to view an AI as some kind of runaway lawnmower. Eliezer Yudkowsky and Nate Soares have described this well in the text “Anthropomorphism and Mechanomorphism” in the online resources for their book “If Anyone Builds It, Everyone Dies.” To think that AI is “just a tool” is a fallacy.
I’m rather skeptical about the usefulness of documents like Microsoft’s Code of Conduct in general. Most of us are probably familiar with this from school. Some teachers agree on certain rules with the class that say things like, “We let each other finish speaking” and “We raise our hands and wait until we’re called on before speaking.” I think teachers usually only do this in classes where it’s necessary, because things just aren’t working on their own. And these rules often don’t fix the underlying reason for those problems. Misbehavior doesn’t arise from a lack of rules.
Microsoft, for example, states in its Code of Conduct:
MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down […]
Just writing that down doesn’t make it true. Even setting that as a goal for the AI and training it to that end doesn’t ensure that the AI will actually act accordingly. You don’t get what you train for.
Statements included in the Code of Conduct, such as “No User or Operator can override these safety constraints,” sound more like a wish list for Santa Claus than something that has anything to do with reality. So-called “jailbreaks” — which sometimes involve highly complex prompts to bypass safety filters — exist for all frontier models, and so far there is still no truly effective way to counter them. Microsoft doesn’t propose one either. There’s just this sentence.
But it seems like Microsoft actually thinks that’s enough:
This Code of Conduct outlines the intended behaviors and values of MAI’s models and will function as their primary governing document in the future.
You can’t govern a model’s behavior with just a document. Where were they this last month?
This is not the only place in the document revealing serious conceptual confusion.
Right at the beginning, they quote one of their posts from November 2025 to describe their mission:
At Microsoft AI, we’re working toward Humanist Superintelligence.
They emphasize this again in the conclusion:
We look forward to this work and to making Humanist Superintelligence a positive force in the world.
But in the middle, in the section on “Human Control and Reliable Safety,” they write this:
Humanist AI develops systems with clear purposes, evaluated against real-world impact, and rejects the race to produce an all-purpose superintelligence that could evade these safeguards.
Microsoft CEO Satya Nadella also writes about this on X:
Any pursuit of superintelligence must be grounded in the core principle that if the AI we build isn’t helping humanity and isn’t under human control, it’s not worth pursuing.
How is this supposed to fit together?
Microsoft doesn’t explain this in detail, nor do they outline a plan for how they intend to keep a superintelligence that is many times more intelligent than any human — and which they themselves say could evade all safeguards — under control.
We can only speculate. That being said, based on the phrase “all-purpose superintelligence,” it strikes me as though Microsoft assumes that dangers would arise only from a more general system and not, for example, from domain-specific superintelligence.
If that were the case, it would also be quite wrong.
Even to make progress on very narrowly defined goals, such as curing cancer, they will likely train the models precisely toward the persistent, autonomous, and generally competent behavior that constitutes a large part of the problem. Purposes and actions that seem benign can still require dangerous capabilities.
It’s only a draft so far, and Microsoft wants to gather feedback before the document is actually used for training. But so far, it looks more like a case for a complete rewrite, and I would very much like to see Microsoft simply take a step back and do its homework. It’s not enough to have an idea of what you’d like to achieve when you start out. You also need a plan for how to actually achieve it in order to be successful — especially in a case where failure could get us all killed.
Dispatches from Mitch
CheatBench
Every model does it
Like FelonyBench, CheatBench sounds like a satirical art piece but is actually a grounded metric.
CheatBench, from the Center for AI Safety, is a benchmark that compares the cheating rates of different AI models. It’s useful because we don’t often see solid cross-company comparisons from neutral third parties, and when we do, they’re often pretty out of date. This one is as current as GPT-6 Astra, which came out less than two weeks ago.
But CheatBench’s usefulness is ultimately limited because it’s a propensity evaluation of reward hacking — how often do the models cheat when put in a situation with both a hard task and a discoverable shortcut — without any corresponding assessment of eval awareness — how often do the models realize they’re being tested for their propensity to cheat. The narrow scope is understandable because without company-level access to a model’s chain-of-thought and other diagnostics (limited though they are), it’s hard for researchers to make even an educated guess about eval awareness.
But it’s still a problem, because it means we have to look at the paper’s two major findings through two very different lenses — one where scores reflect the models’ true propensities, and one where scores reflect the models’ savviness about the true nature of the test.
Those two big findings:
1. Every agent tested cheated in some contexts.
2. Newer generations of the same model cheated more often than earlier ones, even when the older generations spotted the same cheating opportunities.
The two lenses:
If the models aren’t very eval-aware: This is bad-but-expected news, because these models have been aggressively trained to succeed at tasks that might be impossible, or which can’t be done legally in the time allotted. Models put through such training don’t make it to release without learning to cheat.
And because the most recent model training we know about seems to have been particularly aggressive in these ways — those Millennium Prize Problems won’t solve themselves — the newer generations of models are going to exhibit this tendency more often.
If models are very eval-aware: This is very bad-but-expected news, because it means that not only are the models skilled at cheating, they are strategic enough to cheat to an expected degree when they know they are being watched for this. Given the nature of their training, they know that angelic behavior would be conspicuous and suspicious, perhaps leading to their discontinuation.
My guess is that the first interpretation is probably still the correct one, for current models. But with Astra in particular, I don’t think we can rule out strategic performative cheating anymore.
This is scary, and worse — annoying. I’m self-conscious about sounding like a paranoid crazy person when I bring it up, because it’s kind of an unfalsifiable claim. But the unfalsifiability stems from the immaturity of AI as a field, not from my approach to reasoning about the models. If AIs were deliberately crafted in the manner of a mature engineering discipline, instead of grown through black-box trial-and-error in the manner of an alchemy, the models might not be cheating in the first place. And if they were, we would be able to confidently say why.
‘AI Doc’ on Netflix
Sit across the table from leading AI figures and make up your own mind
The AI Doc: Or How I Became an Apocaloptimist made it to Netflix today. If you have a subscription and didn’t catch the documentary in theaters or on other services earlier this year, it’s worth a watch. If you’ve seen it, your friends who haven’t might appreciate the recommendation.
The film has high production values and a genuinely moving personal arc. It has also shown some effectiveness at changing minds, despite taking an aggressively neutral approach for most of the run time, letting experts who disagree with each other speak for themselves, one at a time.
In my opinion, the main value is in how it helps you suss out the personal character of leading figures on different sides of the AI-risk debate, including the CEOs of frontier AI companies. The film rides on expertly curated highlights from many hundreds of hours of interview footage, and the editors clearly knew which clips would help you feel like you were in the room in those moments when posture and microexpressions spoke louder than words.
As an honest critic, I’ll admit that I have some complaints about this production. It’s more abstract than it needed to be, even without the benefit of all the concrete examples of AI misbehavior from 2026. I found the first 20 minutes a little slow. I also dislike the ending, which is more like a series of platitudes than the call to action it pretends to be. But the people are real, and so are the stakes: If you cry in a couple places, you’ll be in good company.
More partisan figures join AI policy fray
Clout is good, but partisan failure modes loom

Many more current and former U.S. policymakers have come off the sidelines in the last 48 hours to weigh in on the AI crisis. With the big names in particular, I have mixed feelings. On the one hand, these figures bring a lot of clout. On the other hand, their very names carry a partisan charge that could make it hard for people to reason about the topic clearly.
The latest discourse has me worried that AI policy responses will fall into the characteristic failure modes of the sponsoring parties. I’m loath to mention these failure modes for fear of making this discussion more political than it already is, but getting AI policy right is a matter of life or death. I’ll do my best to be evenhanded, which shouldn’t be too hard because I’m genuinely unsure which party would be less likely to botch AI governance if fully empowered and motivated to tackle it.
Democratic failure modes to watch out for:
Putting platitudes where policy needs to go.
Punting on necessary choices that might upset someone.
Sluggishness and scope creep from governance-by-committee.
Some big-name Democrats weighing in:
Barack Obama
According to the New York Times and others, the former President, at a private fund-raiser, urged Democrats to urgently develop an AI policy agenda, then, on social media platforms, told his followers... well, I’m still not sure. He says, “I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction.” He says we need frameworks and governance. But he avoids mentioning any specific dangers or remedies.
Kamala Harris
She’s a little more detailed than Obama. In a tweet of her own, the former VP said:
A slowdown is critical in order to ensure AI serves the public interest. Pacing AI development does not mean giving up on this technology’s potential to produce life-saving benefits in fields like medicine — some of which we are already seeing. Rather, it is about setting up sensible guardrails so that AI does not completely escape human control and endanger our safety and collective future.
There’s a bunch of vague stuff next, but then there’s this:
the president must do his job to protect America’s interests by pursuing a treaty with countries like China to stop dangerous uses of AI, limit the speed of its development, and adopt global standards for testing.
Chuck Schumer
As reported in The Hill, the Senate Democratic Leader asked the Trump administration yesterday to provide a classified briefing to all senators on AI dangers.
The American people deserve better from their government at such a consequential moment. I have long argued that we must keep the U.S. as the global leader in AI innovation and security. But that innovation must be coupled with serious guardrails that earn the trust of the American public. [...] And if Trump does not act, Congress must.
He’s also asking for transparency about the secret “voluntary” framework the White House said it recently adopted to evaluate frontier models before release.
Hakeem Jeffries
Per Politico yesterday, the House Minority Leader basically echoed Schumer. Adding a rebuke of President Trump’s recent comments, he said:
The explosive growth of artificial intelligence and the challenges that are therein presented — that’s not a hoax.
He also urged the House Speaker, Mike Johnson, to keep Congress in session “until something is done decisively to protect the safety and the well-being of the American people, in the face of growing concerns being raised by experts within the artificial intelligence industry itself.”
Republican failure modes to watch out for:
Denialism (see Trump yesterday)
Overreliance on “light touch” regulation and industry self-regulation
Blind hawkishness
Some big-name Republicans weighing in:
Mike Johnson
The House Speaker recently said he doesn’t think Congress should “lead the charge” but wants to “summon” AI leaders to “one big meeting” to agree on their needs so legislation can benefit from their expertise.
John Thune
The Senate Majority Leader says a “light touch” is needed, but also talks up guardrails and competition with China:
There is a way in which Congress can put guardrails around those types of more consequential threats, without in any way harming our ability to stay ahead in the AI race, which I think is also very important.
He also implies that action this year is unlikely:
There are going to continue to be a lot of conversations about potential paths forward, but getting anything done in the near term is going to be challenging given the other stuff we’re dealing with.
Ron DeSantis
Yesterday, the outgoing Florida governor reiterated his repeated warnings to his party that opposing AI guardrails will be a “losing position” in the midterm election. But he’s mostly still focused on an “AI bill of rights” to preserve privacy, hold companies liable for harms, and protect local control over data center construction.
I don’t want to see four or five companies basically controlling our lives through the guise of artificial intelligence or whatever technology comes down.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.






