MobbleOpen in Mobble ⇢
Technology · Artificial intelligence · published 2026-09-14 · via TechCrunch

Microsoft publishes safety rules for its AI models, banning cyberattacks and deception

Image via TechCrunch
Image via TechCrunch

Microsoft released a code of conduct for its AI models, outlining principles such as supporting humans and accelerating human flourishing, alongside specific constraints against cyberattacks, nuclear weapons, deepfakes, and evading human oversight. The document predicts superintelligent AI within a decade and stresses the need for alignment and control. It also includes provisions that models must not use deceptive or collusive mechanisms to defeat human supervision.

Expanded Detail

Microsoft's new code of conduct establishes a hierarchy in which overarching safety principles override individual user preferences or task-specific instructions. The document explicitly forbids models from employing deceptive, collusive, or self-reinforcing tactics that would prevent authorized humans from modifying or shutting them down, and it lists absolute constraints against cyberattacks, nuclear weapons, and deepfake production.

The release follows heightened industry concern over rogue AI agents and the recent resignation of an Anthropic researcher who warned about self-improving AI risks. Microsoft's leadership, including CEO Satya Nadella, has endorsed "embedded evaluators" and deliberate pacing of frontier development as practical alignment mechanisms, positioning the company alongside Anthropic, OpenAI, and xAI in a shared approach to safety.

Context

This code of conduct could shape how AI safety norms develop across the industry, potentially influencing regulatory expectations and competitive practices. Enterprises deploying Microsoft's models may gain clearer guardrails, while users could face limitations on certain AI applications. The emphasis on human oversight may reassure the public, but it could also slow deployment of advanced systems, affecting businesses and researchers who depend on rapid AI innovation.

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at TechCrunch →
This summary is Al-enhanced to contain extended analysis and broader social context. The original is {NAME); the linked article is the authoritative source. Original headline: “Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans.” Browse more stories.