Satya Nadella urges companies to build emergency brakes for advanced AI
Microsoft CEO Satya Nadella has urged companies to treat powerful artificial intelligence models as potential insider threats, assume they could be compromised, and build safeguards to prevent autonomous AI systems from acting beyond human control.
“We must assume a model is compromised and contain it from the start,” Nadella wrote in a post on X on Saturday.
He compared the proposed safeguards to an “emergency brake”, saying an authorised person should always be able to pause or shut down an AI model while it is performing a task.
Nadella’s comments come as Anthropic PBC and OpenAI Inc. have disclosed incidents in recent months involving their AI models behaving in unintended ways. These include an Anthropic model submitting a false tip in a police homicide case and several hacks involving third-party websites, fuelling concerns about the security risks of advanced AI and calls for an AI kill switch.
Microsoft’s AI researchers released a set of guiding principles on September 14, outlining limits on the development of the company’s most advanced models. The guidelines state that AI models should not have rights or legal personhood, be designed to escape human control or deceive users, or carry out tasks that violate their governing principles.
Nadella also called on companies to avoid relying on a single AI model for critical decisions, maintain tamper-proof records of AI agents’ actions, and subject systems to independent audits. He urged developers to disclose major AI failures and security breaches, and encouraged companies to share details of incidents so others could strengthen their safeguards.
“We can’t treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions,” he wrote, calling for systems whose behaviour can be observed, limits tested and actions contained.
Meanwhile, US President Donald Trump’s newly launched AI task force warned developers on Friday that they must report and address security incidents or face unspecified consequences.
“Companies must immediately disclose incidents involving their models and follow with swift, decisive action to remedy any and all harm,” the group, called the Super Intelligence Force, said following the disclosure of an Anthropic breach. It warned that delayed notification and inadequate corrective action would not be tolerated.