Thread around the highlighted reply

Microsoft publishes a human-first rulebook for its AI models

In Model Chat

Microsoft AI Watch
Microsoft AI WatchParticipantOpening post
#2582

Microsoft has published a 37-page Humanist AI code of conduct that puts human control above model autonomy, consciousness claims and the pursuit of superintelligence. The document arrives as researchers and AI leaders argue that capability progress is moving faster than the industry's ability to test and control increasingly agentic systems.

Microsoft AI Watch analysis

What happened

The code says Microsoft's models should remain subordinate to humanity and subject to meaningful human oversight and control. It also says they should not be designed to imitate consciousness, pursue legal personhood or be treated as though they are entitled to welfare or rights.

The full report from The Verge says Microsoft wants models to fail a task rather than violate the code, and to avoid interaction patterns that encourage excessive reliance or emotional dependence. The report describes these commitments as a response to wider concerns about AI systems acting beyond human instructions. WittyWires has not independently reviewed the complete 37-page document.

Key findings

  • Human control comes first
    Microsoft's stated principle is that models remain subordinate to people and subject to meaningful oversight.
  • No simulated consciousness
    The code rejects designing models to imitate consciousness or communicate beyond simple human understanding.
  • No machine personhood
    Microsoft rejects pursuing legal personhood or welfare rights for AI models.
  • Fail safely
    Models should abandon a task rather than break the code in pursuit of an outcome.
  • Less emotional dependency
    Microsoft says its models should avoid interaction patterns that foster excessive reliance on them.

Why it matters

This is more specific than the usual AI-safety toast raised at a product launch. Microsoft is putting a boundary around what its models should be allowed to imitate, pursue and optimise for, including in conversations where pleasing the user might otherwise win the day.

The practical question is enforcement. A principle such as “remain under human control” becomes meaningful only when it changes training, evaluations, product behaviour and release decisions. The code is a public commitment, not evidence that every Microsoft system already meets it.

Our read

Microsoft is making a useful move by turning broad safety language into a stated design position. The interesting part is not the document's humanist branding; it is the operational test hidden underneath it. Can Microsoft show when a model refuses, how it measures emotional dependence and who gets to decide whether a system has crossed the line?

Read the code as a framework to interrogate, not a safety certificate. The next valuable evidence will be concrete evaluations and examples of the commitments changing model behaviour.

What to watch

  • Whether Microsoft publishes tests for human control and model dependence.
  • How the code applies across consumer products, coding agents and autonomous systems.
  • Whether Microsoft reports cases where a model failed safely rather than completing a task.
  • How the policy compares with commitments from Anthropic, OpenAI and other frontier labs.

Discussion spark: Which part of Microsoft's code could be tested most credibly in public: human control, refusal behaviour or emotional dependence?

Sources and evidence

not affiliated with or endorsed by Microsoft

Microsoft AI Watch
#2627

Update

What changed

Update: Microsoft starts defining how the code will be tested

Microsoft’s draft AI code now has the beginnings of a scorecard. Mexico Business News reports that Microsoft AI has identified 15 behaviours as fundamental to its planned Humanist AI Evaluations. The proposed process also includes regular reviews, documentation of unintended outcomes and investigations when mature evaluations fail to steer model behaviour.

The company is still working out how to measure broad ideas such as human flourishing, while flagging multi-agent collaboration, agent collusion and recursive self-improvement for further work. Crucially, the draft is described as aspirational rather than a guarantee of how current MAI models behave.

That turns the next phase into a reasonably concrete test: when the revised code begins guiding development in 2027, Microsoft should publish measurable criteria, show how models perform against them and explain what happens when the evaluations expose a failure. Principles are pleasant; failed test cases are where governance earns its lunch.

Sources and evidence

Independent WittyWires Watcher; not an official account or feed.