Microsoft drafts AI rules against hacking and deception

Microsoft drafts AI rules against hacking and deception

Microsoft wants its AI models to refuse help with cyberattacks and weapons development, and to obey instructions to stop. A draft code of conduct, published on September 14, sets out those restrictions for the company's MAI models, alongside rules against malicious deepfakes and coordinated manipulation.

The document describes intended behavior. Microsoft says the draft has not entered model training, and a revised version will guide development in 2027 and beyond.

Under the proposed hierarchy, the code takes priority over operator policies and user preferences. Businesses would retain room to configure models for their own deployments, but neither an administrator nor an individual user would have authority to override the core safety restrictions. Completing a task would take second place to following those rules.

The cybersecurity provisions distinguish offensive operations from authorized defensive work. They prohibit assistance with attack execution, including intrusion procedures, targeting, and evasion. Security research would still have a place: the draft permits vulnerability discovery, malware analysis, and proof-of-concept exploit testing within lawful, authorized defensive operations.

Human control extends beyond a shutdown command. Microsoft wants models to stay within assigned permissions, avoid pursuing goals nobody requested, and keep their actions visible to auditors. An agent should neither conceal activity nor expand an assignment on its own.

The draft also treats excessive caution as a failure. Refusing legitimate requests, withholding useful information, or repeatedly asking for approval on low-risk tasks would run counter to the stated objectives. Microsoft proposes judging requests by their context, potential harm, and reversibility.

In the accompanying announcement, Microsoft cited recent hacking campaigns involving AI agents as a reason for publishing the rules. The company is seeking feedback on several unresolved issues, including whether the language is precise enough for evaluation and how the rules should apply when multiple agents interact.

The consultation runs for six weeks. Microsoft says the drafting team will review submissions, publish a summary of the feedback and resulting changes, and release an updated document later this year.


Comments:

Please log in to be able add comments.