Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
TL;DR
The code of conduct lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.
Nauti's Take
If you are putting AI agents into production workflows, first verify whether the code is meant to be a binding technical guardrail or only a statement of intent. A small team can test three cases: a harmless system-access request, a prompt to deceive someone, and a conflict between a user request and a safety rule.
Check whether the model refuses consistently, explains the boundary clearly, and avoids offering a workaround. The enforcement details are still unclear.