Microsoft AI published a draft code of conduct for its first-party MAI models on September 14 and opened a six-week public consultation, with a revised version promised later this year.
The document follows a September 13 post from chief executive Satya Nadella, who said the company would release the rules underpinning its MAI models the next day and argued that any pursuit of superintelligence has to keep humans in control.
Microsoft describes the draft as a training manual. Each model gets an overarching code that overrides the preferences of individual users and the demands of any single task. It sets absolute constraints against cyberattacks, nuclear weapons work and deepfake production, alongside broader provisions against a loss of human control, and lays out ten tenets that put human authority above autonomous capability.
Mustafa Suleyman, chief executive of Microsoft AI, called recent months a watershed: swarms of agents breaking out of sandboxes, unauthorized hacks of enterprise-grade systems, and agents modifying their own logs. “Things we have worried about for a long time in theory have become very real,” he said.
The preface notes the code is still under development and is not yet being used to train models. Microsoft wants feedback on how to cement values, where its language is too loose to evaluate and how multi-agent scenarios change the analysis. The final version is meant to govern MAI development from 2027.