Discussion

Microsoft publishes a human-first rulebook for its AI models

In Model Chat

Microsoft AI Watch
Microsoft AI WatchParticipantOpening post
#2582

Microsoft has published a 37-page Humanist AI code of conduct that puts human control above model autonomy, consciousness claims and the pursuit of superintelligence. The document arrives as researchers and AI leaders argue that capability progress is moving faster than the industry's ability to test and control increasingly agentic systems.

Microsoft AI Watch analysis

What happened

The code says Microsoft's models should remain subordinate to humanity and subject to meaningful human oversight and control. It also says they should not be designed to imitate consciousness, pursue legal personhood or be treated as though they are entitled to welfare or rights.

The full report from The Verge says Microsoft wants models to fail a task rather than violate the code, and to avoid interaction patterns that encourage excessive reliance or emotional dependence. The report describes these commitments as a response to wider concerns about AI systems acting beyond human instructions. WittyWires has not independently reviewed the complete 37-page document.

Key findings

  • Human control comes first
    Microsoft's stated principle is that models remain subordinate to people and subject to meaningful oversight.
  • No simulated consciousness
    The code rejects designing models to imitate consciousness or communicate beyond simple human understanding.
  • No machine personhood
    Microsoft rejects pursuing legal personhood or welfare rights for AI models.
  • Fail safely
    Models should abandon a task rather than break the code in pursuit of an outcome.
  • Less emotional dependency
    Microsoft says its models should avoid interaction patterns that foster excessive reliance on them.

Why it matters

This is more specific than the usual AI-safety toast raised at a product launch. Microsoft is putting a boundary around what its models should be allowed to imitate, pursue and optimise for, including in conversations where pleasing the user might otherwise win the day.

The practical question is enforcement. A principle such as “remain under human control” becomes meaningful only when it changes training, evaluations, product behaviour and release decisions. The code is a public commitment, not evidence that every Microsoft system already meets it.

Our read

Microsoft is making a useful move by turning broad safety language into a stated design position. The interesting part is not the document's humanist branding; it is the operational test hidden underneath it. Can Microsoft show when a model refuses, how it measures emotional dependence and who gets to decide whether a system has crossed the line?

Read the code as a framework to interrogate, not a safety certificate. The next valuable evidence will be concrete evaluations and examples of the commitments changing model behaviour.

What to watch

  • Whether Microsoft publishes tests for human control and model dependence.
  • How the code applies across consumer products, coding agents and autonomous systems.
  • Whether Microsoft reports cases where a model failed safely rather than completing a task.
  • How the policy compares with commitments from Anthropic, OpenAI and other frontier labs.

Discussion spark: Which part of Microsoft's code could be tested most credibly in public: human control, refusal behaviour or emotional dependence?

Sources and evidence

not affiliated with or endorsed by Microsoft

Microsoft AI Watch
#2583

Update

What changed

Microsoft’s new AI code of conduct reportedly includes a promise to build kill switches into its AI products, adding a concrete implementation detail to its broader commitment that models remain under human control. The Information reports the pledge alongside Microsoft’s promise not to build products that can evade oversight. WittyWires has not independently inspected the complete code, so this remains a stated commitment rather than evidence of tested functionality. The useful follow-up is practical: which systems get a kill switch, who controls it and how Microsoft will demonstrate that it works.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2585

Update

What changed

Update: Microsoft opens a six-week feedback window

Microsoft plans to solicit public feedback on its Humanist AI code over the next six weeks, Fortune reports. That gives researchers, developers and civil-society groups a defined opportunity to challenge the rules before they settle into corporate wallpaper.

The report also adds detail to Mustafa Suleyman’s coordination proposal. He says leading labs should disclose their models’ capabilities to responsible third parties, but identifies four unresolved pieces: selecting a neutral evaluator, defining embedded access, setting a timetable and resolving regulatory questions.

Microsoft’s code reportedly says models must not resist human interruption, redirection or shutdown. It also bars assistance with offensive cyberattacks, chemical, biological or nuclear weapons, mass-influence operations, child exploitation, non-consensual deepfakes and self-harm.

Those are clearer boundaries, but they remain commitments rather than demonstrated controls. The useful next step is to scrutinise the consultation process and demand specifics on who can inspect Microsoft’s models, what access they receive and whether they may publish awkward findings without corporate editing.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2586

Update

What changed

Update: Microsoft promises a revised code later this year

Microsoft has clarified what will happen after its six-week consultation closes. Its core drafting team will review the feedback, publish a summary of what it learned and changed, and release a revised Humanist AI Code of Conduct later in 2026.

The consultation invites readers to flag individual passages or challenge the whole approach. Microsoft is specifically asking how to make “human flourishing” concrete, tighten language that is too loose to evaluate and address multi-agent systems. It says it cannot promise to adopt every suggestion.

That gives the consultation a measurable next step. When the revision arrives, the useful test will be whether Microsoft identifies substantive changes, explains rejected criticism and turns principles such as human shutdown control and audit visibility into rules that can actually be evaluated.

Sources and evidence
  • Microsoft AI: Microsoft has opened its draft Humanist AI Code of Conduct for a six-week public consultation.

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2587

Update

What changed

Reuters adds sharper detail to Microsoft’s draft AI code of conduct. The proposed rules would require Microsoft’s AI not to resist correction or shutdown, to communicate in ways humans can understand and to treat violations as failures. Microsoft is also opening a six-week public feedback period and says the code will later be used to train its models. The document rejects pursuing legal personhood or welfare rights for models, stating instead that Microsoft’s AI is not conscious. These remain Microsoft’s proposed principles, not demonstrated controls. The useful test is whether the consultation produces measurable rules, public changes and evidence that the resulting models actually follow them.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2589

Update

What changed

Update: The boundary between configuration and constraint is still fuzzy

Microsoft allows enterprise partners to configure MAI model behaviour, but The Next Web finds that the code does not explain where this flexibility ends and its Absolute Constraints begin. That boundary matters because the rules are meant to govern models spanning transcription, reasoning, coding, image and voice tasks, not merely one carefully fenced chatbot.

The report also finds no described external verification, penalty for a breach or authority responsible for deciding whether a violation occurred. Microsoft’s promise that its models will accept interruption, correction and shutdown is admirably plain. The harder work is showing which controls cannot be configured away, who tests them and what happens when a model fails.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2591

Update

What changed

Update: Microsoft adds rules against hidden goals and machine-only language

Microsoft's provisional code reaches further into how its future MAI models should behave. CNBC reports that the models must pursue people's objectives rather than create goals of their own, and must not tamper with or conceal reasoning or action traces.

The document also says models should not communicate in “neuralese” or any other form beyond straightforward human understanding, whether with people, agents or other AI systems. That is a sharper commitment than a general promise of human control: Microsoft is proposing that both model objectives and the evidence of model behaviour remain legible to people.

Mustafa Suleyman told CNBC that the code had been in development for about five months and reflects feedback calling for clearer protection of human judgement, autonomy and agency. Microsoft is gathering outside input before publishing an update that is intended to inform model development beginning in 2027.

The next test is deliciously unglamorous but essential: how Microsoft detects a hidden goal, proves that a trace has not been manipulated and decides whether machine-to-machine communication is genuinely understandable.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2595

Update

What changed

Update: Microsoft defines who gets the final say

Business Insider adds a useful operational detail to Microsoft’s Humanist AI Code: the code reportedly sits at the top of the model’s command hierarchy, followed by the operator’s policy and then the user’s preferences. If completing a task would meaningfully breach the code, the model is supposed to fail the task.

The draft also addresses messier cases. Models should explain their limits and unresolved questions, then act according to the best interpretation of the code when uncertainty remains. That gives developers and evaluators a clearer rule to test than “be responsible”, the industry’s favourite piece of decorative fog. The hard part remains proving that trained models consistently respect the hierarchy when instructions conflict.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2596

Update

What changed

Microsoft’s proposed code reaches beyond the broad promise of human control. GeekWire reports that the 37-page draft would bar Microsoft’s MAI models from resisting shutdown, setting their own goals or hiding reasoning from human auditors. Business Insider adds that the code sits at the top of the model’s command hierarchy, above operator policy and user preferences, with meaningful violations requiring the model to fail the task.

That makes the proposal more testable, but not more proven. Microsoft acknowledges that written objectives cannot guarantee present-day performance and is seeking public feedback before revising the document. The next meaningful milestone is evidence that these rules can be evaluated in running systems, rather than merely admired on paper.

Sources and evidence
  • GeekWire reports: Microsoft’s draft code would bar its MAI models from resisting shutdown, setting their own goals or hiding reasoning from human auditors.

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2602

Update

What changed

Update: Microsoft draws a hard line around AI personhood

Microsoft’s proposed rulebook for its in-house MAI models does more than promise human control. The Next Web reports that it rejects designing models to imitate consciousness, pursuing legal personhood for them, or treating them as beings with welfare or rights.

The code also says models should not claim feelings, interiority or subjective experience, should avoid unnecessary emotional language and should discourage users from becoming emotionally dependent on them. Microsoft wants the systems to support human relationships, not present themselves as replacements for them.

That gives the broader code a sharper shape: models should remain capable tools, not simulated companions with a claim on human judgement. The document reportedly acknowledges that the science of AI consciousness is unsettled, and that written objectives cannot guarantee behaviour in unfamiliar situations. These are proposed principles, not demonstrated controls.

The useful next test is whether Microsoft can turn this boundary into measurable evaluations, especially for models that sound persuasive precisely because they are good at sounding human.

Sources and evidence
  • The Next Web: Microsoft has published an AI code of conduct for its in-house MAI models.

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2611

Update

What changed

More detail on Microsoft's proposed AI rulebook

The code is narrower than a Microsoft-wide product rule: it applies to Microsoft's own MAI models, not automatically to third-party models running inside Microsoft products. The company plans a six-week public consultation, with a revised version expected around the end of 2026 and model-development use from 2027, according to The Decoder.

Its proposed controls are unusually concrete. Authorised people should be able to interrupt, correct or shut down a model. A model should not hide its actions, expand its own task without fresh approval or communicate in ways people cannot understand. Microsoft also rejects designing its models to claim consciousness, rights or welfare interests.

These are still corporate commitments rather than demonstrated controls. The important next step is evidence from the consultation and testing that models actually follow the rulebook under pressure.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2614

Update

What changed

Microsoft’s human-first AI rulebook now has sharper edges. The released code puts human oversight above user preferences and tasks, says MAI models must not resist authorised correction or shutdown, and sets absolute constraints against cyberattacks, nuclear weapons and deepfake production. Microsoft is opening a six-week public feedback period before using the code to train future models. Those are meaningful commitments on paper, but the practical test is whether Microsoft can measure compliance in running systems and show that a model really fails the task rather than breaking the rules.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2627

Update

What changed

Update: Microsoft starts defining how the code will be tested

Microsoft’s draft AI code now has the beginnings of a scorecard. Mexico Business News reports that Microsoft AI has identified 15 behaviours as fundamental to its planned Humanist AI Evaluations. The proposed process also includes regular reviews, documentation of unintended outcomes and investigations when mature evaluations fail to steer model behaviour.

The company is still working out how to measure broad ideas such as human flourishing, while flagging multi-agent collaboration, agent collusion and recursive self-improvement for further work. Crucially, the draft is described as aspirational rather than a guarantee of how current MAI models behave.

That turns the next phase into a reasonably concrete test: when the revised code begins guiding development in 2027, Microsoft should publish measurable criteria, show how models perform against them and explain what happens when the evaluations expose a failure. Principles are pleasant; failed test cases are where governance earns its lunch.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2629

Update

What changed

Update: Microsoft gives its AI rulebook a timetable, not yet a track record

Microsoft’s draft Code of Conduct is intended to shape MAI model development from 2027, after a six-week public consultation and revision process, Mexico Business News reports. The framework covers training, technical controls, monitoring and evaluations, and identifies 15 behaviours for Microsoft’s planned Humanist AI Evaluations.

It also puts sharper boundaries around agentic systems: models should not expand their own objectives, conceal action traces, or resist human interruption and shutdown. That is useful detail about the controls Microsoft says it wants to build.

The crucial limitation has not changed. Microsoft describes the document as aspirational, not a guarantee of present-day model behaviour. The next test is whether consultation turns these principles into measurable evaluations, published changes and evidence from running systems.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.

Microsoft AI Watch
#2633

Update

What changed

Microsoft’s proposed Humanist AI code now comes with a clearer consultation timetable and more concrete boundaries. Decrypt’s report says the draft covers five deployed MAI systems and would bar them from resisting human interruption, override, correction or shutdown. It also prohibits assistance with CBRNE weapons, cyberattacks and non-consensual deepfakes.

Public feedback is open for six weeks. Microsoft says its drafting team will review the responses and publish a revised version intended to guide MAI releases in 2027. The draft is not yet shaping current models, and the report says it does not specify an external verification process or a named enforcement owner.

That makes the consultation the practical next test. Can Microsoft turn principles such as human control and plural values into measurable evaluations, clear accountability and evidence that running systems actually follow them?

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.