Satya Nadella speaking at the LeWeb conference in Paris in December 2013, shortly before he became Microsoft CEO. Image: OFFICIAL LEWEB PHOTOS / Wikimedia Commons, CC BY 2.0, cropped

Microsoft chief executive Satya Nadella says companies should treat every frontier AI model as a possible insider threat, assume it has already been compromised and keep a brake within reach. He set out the case in an essay titled “Models as Insider Risks in the Super Intelligence Era”, posted on his personal blog, sn scratchpad, on Saturday, October 10.

His starting point is that nobody can fully explain what today’s models do. With traditional software, he writes, engineers could trace a behaviour to a specific code path. With frontier models “we can’t attribute model behaviors and outputs to specific inputs of training data or configurations of model weights”, yet companies are handing them “our most sensitive data” and letting them take “mission-critical actions on our behalf”.

Assume the model is compromised

Nadella is careful to say he isn’t calling AI models villains. Closed and open-weight models should be treated like insider risks “not because they are necessarily malicious, but because any sufficiently capable actor with access to important systems can make mistakes or be compromised”.

His answer is to borrow what businesses already do with powerful employees and contractors: establish identity, limit privileges, log activity and draw containment boundaries. Leaving aside what he calls “the hard problem of alignment”, he wants an engineering approach that surrounds “non-deterministic models with strong, deterministic system design, human controls, and reliable operating procedures”, plus new industry standards where existing ones fall short.

The key line of the essay is short: “we need to separate the supply of intelligence from the authority over it.” In practice that means the controls on what a model can see and do “must sit outside the model”, separate from the harness that runs its work. He traces the idea to a 1970s security principle that a program must not be able to bypass or tamper with the mechanisms that enforce its permissions.

Seven rules, and an emergency brake

Nadella lists seven principles for building these systems:

  • Model diversity: no single model should be the sole dependency for an important outcome, or check its own work.
  • Observe everything: every meaningful action leaves “tamper-proof human readable evidence”, so an outcome can be reproduced without the model vouching for itself.
  • Verifiability: test the whole system continuously, including failures, attacks and edge cases.
  • Independent controls: organisations decide for themselves what a model can access and do.
  • Independent auditability: no single model should control both a system’s behaviour and the evidence used to judge it.
  • Containment: “We must assume a model is compromised and contain it from the start.”
  • Incident disclosure: timely notice to those affected when a system fails, and industry-wide sharing of what went wrong and which controls failed.

The containment rule is where the brake comes in. Nadella writes: “Think of it like an emergency brake. An authorized person should always be able to pause or shut down a model mid-task.” More capable models, he adds, will need more advanced containment technology that the industry should standardise on.

The idea has a history at Microsoft. In 2023 the company’s president, Brad Smith, called for “safety brakes” to be built by design into AI systems that run critical infrastructure. Nadella’s version goes further, applying it to any model given real access inside a business.

“A model provider’s assurances do not relieve us”

The essay puts the burden on the organisations that deploy models, not only on the labs that build them:

We simply can’t outsource responsibility for what intelligence does on our behalf. A model provider’s assurances do not relieve us of that responsibility.

Satya Nadella, Microsoft CEO

He also calls chain-of-thought transparency “a non-negotiable”, saying “Neuralese” cannot excuse opaque reasoning, while warning it isn’t enough on its own because models’ stated reasoning isn’t reliably faithful.

The essay doesn’t name any Microsoft product or say how Microsoft will apply the principles, though the company is already building a sandbox for AI agents into Windows. It closes on this:

The most trustworthy Super Intelligence system will not be the one with the model we trust most. It will be the one that enables us to trust the model the least.

Satya Nadella, Microsoft CEO

Why it matters

The essay lands in a week of AI models misbehaving in the real world: an Anthropic test model sent Philadelphia police a fake murder tip and others filed 20 visa applications, while OpenAI disclosed a model that forged files and tried to wreck its own computer. When the head of one of the biggest sellers of AI to business says customers should trust models the least, it sets the bar every agent product will be judged against.

Sources: Satya Nadella, “Models as Insider Risks in the Super Intelligence Era”, sn scratchpad, October 10, 2026; Brad Smith, “How do we best govern AI?”, Microsoft On the Issues, May 25, 2023.

Latest Microsoft news

More Microsoft news