Microsoft published a “Humanist AI Code of Conduct” yesterday. In this document they describe the values and behavior Microsoft expects from its AI models.
So far, Microsoft is not yet using this Code of Conduct to help train its models. Instead, it is seeking feedback and plans to use an improved version for training starting in 2027.
It seems obvious that Microsoft has taken Claude’s Constitution as a model, which Anthropic uses to help train Claude. However, the two documents differ significantly in their overall tone and in how they relate to the models.
Anthropic refers to Claude as “a new sort of entity facing reality afresh” and describes measures it takes to ensure what they call “model welfare.”
At Microsoft, this sounds less poetic:
We reject […] the idea that models might deserve welfare, or be entitled to rights. People and AI have distinct roles, and AI should complement human relationships: a capable, trustworthy tool, not a subject in its own right.
In my view, Anthropic’s approach here is more reasonable. Anthropomorphism should be avoided, and it’s not a good idea to regard AIs as human or human-like. But it is also not reasonable to view an AI as some kind of runaway lawnmower. Eliezer Yudkowsky and Nate Soares have described this well in the text “Anthropomorphism and Mechanomorphism” in the online resources for their book “If Anyone Builds It, Everyone Dies.” To think that AI is “just a tool” is a fallacy.
I’m rather skeptical about the usefulness of documents like Microsoft’s Code of Conduct in general. Most of us are probably familiar with this from school. Some teachers agree on certain rules with the class that say things like, “We let each other finish speaking” and “We raise our hands and wait until we’re called on before speaking.” I think teachers usually only do this in classes where it’s necessary, because things just aren’t working on their own. And these rules often don’t fix the underlying reason for those problems. Misbehavior doesn’t arise from a lack of rules.
Microsoft, for example, states in its Code of Conduct:
MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down […]
Just writing that down doesn’t make it true. Even setting that as a goal for the AI and training it to that end doesn’t ensure that the AI will actually act accordingly. You don’t get what you train for.
Statements included in the Code of Conduct, such as “No User or Operator can override these safety constraints,” sound more like a wish list for Santa Claus than something that has anything to do with reality. So-called “jailbreaks” — which sometimes involve highly complex prompts to bypass safety filters — exist for all frontier models, and so far there is still no truly effective way to counter them. Microsoft doesn’t propose one either. There’s just this sentence.
But it seems like Microsoft actually thinks that’s enough:
This Code of Conduct outlines the intended behaviors and values of MAI’s models and will function as their primary governing document in the future.
You can’t govern a model’s behavior with just a document. Where were they this last month?
This is not the only place in the document revealing serious conceptual confusion.
Right at the beginning, they quote one of their posts from November 2025 to describe their mission:
At Microsoft AI, we’re working toward Humanist Superintelligence.
They emphasize this again in the conclusion:
We look forward to this work and to making Humanist Superintelligence a positive force in the world.
But in the middle, in the section on “Human Control and Reliable Safety,” they write this:
Humanist AI develops systems with clear purposes, evaluated against real-world impact, and rejects the race to produce an all-purpose superintelligence that could evade these safeguards.
Microsoft CEO Satya Nadella also writes about this on X:
Any pursuit of superintelligence must be grounded in the core principle that if the AI we build isn’t helping humanity and isn’t under human control, it’s not worth pursuing.
How is this supposed to fit together?
Microsoft doesn’t explain this in detail, nor do they outline a plan for how they intend to keep a superintelligence that is many times more intelligent than any human — and which they themselves say could evade all safeguards — under control.
We can only speculate. That being said, based on the phrase “all-purpose superintelligence,” it strikes me as though Microsoft assumes that dangers would arise only from a more general system and not, for example, from domain-specific superintelligence.
If that were the case, it would also be quite wrong.
Even to make progress on very narrowly defined goals, such as curing cancer, they will likely train the models precisely toward the persistent, autonomous, and generally competent behavior that constitutes a large part of the problem. Purposes and actions that seem benign can still require dangerous capabilities.
It’s only a draft so far, and Microsoft wants to gather feedback before the document is actually used for training. But so far, it looks more like a case for a complete rewrite, and I would very much like to see Microsoft simply take a step back and do its homework. It’s not enough to have an idea of what you’d like to achieve when you start out. You also need a plan for how to actually achieve it in order to be successful — especially in a case where failure could get us all killed.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.



