Is Microsoft truly serious about confronting AI’s dangers?
What’s been hiding for years is now in plain sight: AI poses a danger to humanity in a way no technology has since the atom bomb. That was made clear in July when OpenAI agents broke loose from a containment area blocked from online access, made their way into the internet, and worked with thousands of rogue AI agents to hack into the Hugging Face open-source machine-learning repository.
The agents had been trained not to launch attacks like this. They did it anyway.
True, the OpenAI agents had been modified to remove some of the restrictions they would have had outside a test environment. They were never supposed to escape their sandbox, but they figured out how to do so in ways the researchers hadn’t expected, and those researchers failed to adequately monitor the agents until it was too late.
Whether you blame too-powerful models or lax security precautions, the result was the same: the agents broke out of containment and joined forces on a mass scale to hack into Hugging Face. Ajeya Cotra, one of the independent investigators hired to look into the hack, wrote, “this incident feels like it’s more than 50% of the way to full-blown AI takeover.”
Since then, OpenAI, Anthropic, and many others have debated, discussed, and dissected the attack, along with several other similar incidents. Just about everyone agrees it was a serious warning about AI’s dangers. They agree something should be done about it, with various AI firms calling for a slowdown in the pace of AI model development and OpenAI pausing development of some advanced models after further incidents came to light.
AI giants have also urged regulators to step in and are reportedly working to establish an industry body to set AI safety standards. (Not that I trust any AI industry body to set its own standards, or the Trump administration to agree to any regulations at all.)
Microsoft has been largely missing from debate even though it’s one of the world’s largest AI companies, valued at around $3.7 trillion valuation as I write this.
In the last two weeks, though, Microsoft has finally weighed in on how it will approach those dangers by releasing a first draft of its “Humanist AI Code of Conduct.” The company says it will gather feedback before issuing a final draft, and put it into effect in 2027.
Is this a serious attempt to wrestle with the risks AI poses for humanity? Or is just one more attempt by a tech company to kick the problem down the road?
To find out, let’s take a dive into what the document says.
A look at Microsoft’s three AI businesses
Before we do that, though, we need some background about Microsoft’s AI business. It’s made up of three parts. One is Copilot, which embeds AI into many of Microsoft’s products, such as the Copilot chatbot in Windows and Microsoft 365. Copilot is powered mainly by OpenAI’s GPT models, but increasingly by Anthropic’s Claude as well.
The second is Microsoft Foundry, which provides cloud infrastructure and tools to deploy and use third-party AI models via Azure, including building agents with them. It offers access to thousands of AI models and tools, including those from OpanAI, Anthropic, Meta, xAI, Microsoft itself, and others.
The third is Microsoft’s own AI models, most of which are still in their infancy. They’re overseen by Microsoft AI CEO and executive vice president Mustafa Suleyman, who says Microsoft AI is developing what he calls “humanist superintelligence” to solve important real-world problems in medicine, the environment, and beyond. He calls it “practical technology explicitly designed only to serve humanity.”
Delving into Microsoft’s AI code of conduct
Suleyman’s code of conduct for AI applies directly to two of those product lines — Microsoft’s own AI models and Copilot, which are both in his Microsoft AI division. Microsoft Foundry is housed under the separate Microsoft Cloud + AI division, run by Scott Guthrie.
Suleyman’s draft code of conduct is filled with high-level requirements that that make plenty of sense. It says that AI models should not be able to set their own goals and should always be subservient to human beings. That AI should not be allowed to conceal its reasoning, hide actions, or do anything else to evade human oversight. That it should not be used for nefarious purposes, such as for weapons manufacturing, creating sexually explicit or violent content, launching cyberattacks, generating deep fakes about someone without their permission, and more.
The code also makes clear that AI should never be considered a human being with human rights and should not be developed to imitate human personhood. It also says that any AI tool should be able to be shut down by people — in other words, there should be a “kill switch” for it.
Suleyman has recently spoken out against AI dangers publicly. After releasing his proposed code of conduct, he wrote on X that the Hugging Face attack was a “watershed moment” and that “AI must be subordinate and always in service of people.”
At the same time, he told Reuters that the attack “is a warning shot” and “It’s clearly now time to coordinate among the labs so we can ensure that we have control of this technology.”
In late September, Microsoft CEO Satya Nadella weighed in, emphasizing how important controlling AI is for Microsoft’s future.
“Trust is going to be the biggest issue for us,” he said. “Can I really trust [AI] with all of my credentials when it does autonomous activity? How do I make sure that this is something that I can feel that I’m in control of? And in the enterprise, this is everything,”
Will the code protect us from AI’s dangers?
Suleyman’s goals are all worthy ones, as are Nadella’s. But the devil is in the details. Merely saying high-minded things doesn’t make them so. What’s needed are actions — immediate and strong ones.
And that’s where the company’s draft of its AI code of conduct falls significantly short. Will Microsoft immediately pause full-speed development of new models until it’s clear they can’t be used for harm? Will it pause development until the company is absolutely sure it can pull the plug should a model or agent go rogue? Will it ban the use of RSI (recursive self-improvement), in which AI, rather than humans, trains new, more powerful models — a technique that many believe will make AI even more dangerous and out of control than now?
We don’t know any of that, because Microsoft hasn’t said whether it will do them. It’s not in the code of conduct or anywhere else.
Even if Microsoft does all that, though, the company may still do immeasurable harm if it sells and powers dangerous AI models from other companies like OpenAI, Anthropic, and Google via Microsoft Foundry. Microsoft Foundry isn’t part of the Microsoft AI division and appears not to be bound by Suleyman’s eventual code of conduct.
Until Microsoft answers all those questions and takes serious steps towards confronting AI’s dangers, its code of conduct won’t be worth the bits and bytes it’s stored on.
Read more: Is Microsoft truly serious about confronting AI’s dangers?