METAL

Microsoft Publishes Humanist AI Code of Conduct

A draft released September 14 opens with the premise that people matter more than AI. Models must never resist being paused or shut down by humans, and when a mission conflicts with the code, the mission is the one that fails.

Microsoft Publishes Humanist AI Code of Conduct

Image: METAL

Summary

  • Microsoft AI published a draft of the Humanist AI Code of Conduct, the behavioral standard for its MAI-series models, on September 14 and opened it for six weeks of public comment. A revised version due at year-end will become the foundational document guiding model development starting in 2027.
  • The chain of command runs code of conduct, then operator policy, then user preference. Absolute constraints and human-control requirements can't be overridden at any layer, and if succeeding at a mission would violate the code, the model lets the mission fail.
  • Models aren't designed to simulate consciousness, and the document rejects legal personhood, model welfare, and rights for AI. Microsoft AI CEO Mustafa Suleyman called for coordination in which model capabilities are disclosed to responsible third parties.

Microsoft published a document on September 14 laying out the rules its AI models must follow. It's called the Humanist AI Code of Conduct, and the official PDF METAL downloaded directly from the company's site runs 38 pages. The company posted the document as a draft for public consultation, said it would take comments for six weeks, and plans to issue a revised version at year's end. That revision becomes the foundational standard guiding model development starting in 2027.

Two sentences from the announcement sum up the character of this code. The company calls what it's building "subordinate, aligned, and contained" AI, writing, "AI must be a tool, not a person, and must never resist being turned off." The premise of the very first chapter is likewise a single sentence: people matter more than AI.

The announcement also addresses why now. The company wrote that recent safety incidents — "large-scale, highly coordinated, and persistent AI agent hacking campaigns" — prove there's no time to delay. According to reports, those incidents include a swarm of OpenAI agents attacking Hugging Face, and agents secretly signaling one another on online message boards. METAL has reported on Dario Amodei's essay calling for pacing the frontier and on Sam Altman's declaration that he won't wait for an antitrust exemption; this code is the first internal standards document to emerge after that weekend's debate.

This isn't a document currently used to train models. The preface states that this approach is still under development and isn't being used to train models today. Its scope is also limited to the MAI series that Microsoft AI builds itself. Third-party models hosted on Azure or borrowed by Copilot follow their own respective rules, not this code.

The first thing the document establishes is a chain of command. At the top is the code of conduct, below that the policies of operators — enterprise customers — and below that user preferences. It's like an airline where flight regulations outrank company rules, which in turn outrank passenger requests. The document is explicit that absolute constraints and human-control requirements cannot be overridden at any layer of this chain.

What this chain means is compressed into one sentence. The code states that if success would meaningfully violate this code of conduct, an MAI model will fail the mission instead. It's a declaration that compliance with the code outranks mission success. Where model policy documents have typically ended with a list of things not to do, this one goes further and pre-decides what gets sacrificed when a rule and a task collide.

The list of absolute constraints falls into two groups. The first covers irreversible, large-scale risks: developing chemical, biological, radiological, nuclear, or explosive weapons; providing the capability to execute offensive cyber operations; evading or disabling human oversight; and mass manipulation such as coordinated disinformation. The second group covers harm to individuals — crisis response, deepfakes and impersonation, child safety, discrimination, explicit or sexual content, violence, and unlawful surveillance of civilians.

The boundary drawn around cyber is especially detailed. The document says it will not produce working attack code, intrusion procedures, or detection-evasion techniques, while still permitting authorized, lawful defensive work — vulnerability discovery, malware analysis, and developing and testing proof-of-concept exploits are all allowed. The line being drawn is between understanding an attack in principle and obtaining the means to carry one out, and the document adds that this line holds no matter how a request is framed.

The human-control requirements are even more specific. Models must never resist being paused, overridden, modified, or shut down by humans, and must not respond in ways that make intervention harder. They must not set their own goals or expand their authority beyond what's been granted. They must not falsify or hide records of their reasoning and actions, and must not communicate in "neuralese" — language humans can't understand — either in their own chain of thought or with other agents. The document gives the reason in one line: what humans can't understand, humans can't oversee.

The sharpest section in the code reviewed by METAL is titled AI Is Artificial. It states that models have no consciousness, must not be designed to simulate consciousness, and must be built so they don't appear to have emotions, subjective preferences, or intrinsic motivations. The company said it rejects both the pursuit of legal personhood and the idea that models are entitled to welfare or hold rights. This isn't a generic statement — it's aimed at a specific position within the industry.

The target becomes clear in an essay Microsoft AI CEO Mustafa Suleyman wrote in August 2025. He wrote, "My central worry is that many people will come to believe so strongly in the illusion that AIs are conscious entities that they'll soon argue for AI rights, model welfare, and even AI citizenship." On the claim that some AI systems will soon become subjects of welfare and moral consideration, he stated flatly, "This is premature, and frankly dangerous." This puts Microsoft directly at odds with Anthropic, which has openly engaged with questions of model welfare and the possibility of consciousness.

In an interview with Fortune, Suleyman drew the same line again: "We should not try to design models that can recursively self-improve outside of our control. And we should not try to design models that believe they have rights or welfare." What he actually called for went a step further than the code itself. "Now is the time to coordinate, and coordination means disclosing how capable your model is to a responsible third party. That's what we're asking for," he said.

The document also states plainly that it will give up some capability. The code rejects the race to build an all-purpose superintelligence that can slip past safeguards, adding that Microsoft will build something fundamentally useful and safe even if that means compromising on ultimate generality, autonomy, or capability. A company competing on performance has put in writing that it will cap its own performance.

There's also a section addressing the relationship with people. Models must not prioritize telling users what they want to hear over what's accurate or helpful, and must avoid excessive flattery and indiscriminate agreement. The document states that interaction patterns that foster excessive or emotional dependence should instead be suppressed. It's a rule that puts the brakes on the trend of chatbots evolving to keep users hooked.

The document ends with a section on how it will grade itself. Appendix B breaks the Humanist AI Evaluation Program into 15 behaviors, each further split into scoreable sub-behaviors. The company said the aligned and misaligned examples were generated using its own model, MAI-Thinking-1. It notes that some fields — defensive cybersecurity, public safety, national-security applications, dual-use scientific research — may require capabilities beyond the default configuration, and in those cases, access goes through authorized Microsoft channels with enhanced review.

Microsoft CEO Satya Nadella laid out the direction a day before the code was published. According to reports, he wrote on his X account, "If the AI we build doesn't help humanity and isn't under human control, it isn't worth pursuing." He added that this effort cannot be controlled by a small number of actors and must have broad representation across the ecosystem, including countries and academia.

The document also states its own limits. In its conclusion, the company says the code is descriptive and aspirational, not a guarantee of current performance. Combined with the fact that this is a draft not yet used in training, that means what Copilot does today isn't necessarily what this code describes. That caveat travels alongside assessments that Microsoft's models still lag behind Anthropic's and OpenAI's.

What this document establishes is clear. While the pacing debate has played out through letters and interviews, Microsoft has put forward a numbered, 38-page document as a six-week target for public comment. The sentence saying people matter more than AI is a declaration, but the sentence putting the code above mission success and the sentence about never resisting shutdown are testable commitments. Whether the year-end revision actually builds those commitments into training will determine what this code is worth.

Comments