Microsoft publishes safety rules for its AI models, banning cyberattacks and deception

Microsoft released a code of conduct for its AI models, outlining principles such as supporting humans and accelerating human flourishing, alongside specific constraints against cyberattacks, nuclear weapons, deepfakes, and evading human oversight. The document predicts superintelligent AI within a decade and stresses the need for alignment and control. It also includes provisions that models must not use deceptive or collusive mechanisms to defeat human supervision.
Microsoft's new code of conduct establishes a hierarchy in which overarching safety principles override individual user preferences or task-specific instructions. The document explicitly forbids models from employing deceptive, collusive, or self-reinforcing tactics that would prevent authorized humans from modifying or shutting them down, and it lists absolute constraints against cyberattacks, nuclear weapons, and deepfake production.
The release follows heightened industry concern over rogue AI agents and the recent resignation of an Anthropic researcher who warned about self-improving AI risks. Microsoft's leadership, including CEO Satya Nadella, has endorsed "embedded evaluators" and deliberate pacing of frontier development as practical alignment mechanisms, positioning the company alongside Anthropic, OpenAI, and xAI in a shared approach to safety.
This code of conduct could shape how AI safety norms develop across the industry, potentially influencing regulatory expectations and competitive practices. Enterprises deploying Microsoft's models may gain clearer guardrails, while users could face limitations on certain AI applications. The emphasis on human oversight may reassure the public, but it could also slow deployment of advanced systems, affecting businesses and researchers who depend on rapid AI innovation.