Microsoft CEO Satya Nadella is calling for AI systems to be designed with controls that let people monitor and stop a model while it is working. He argues that trust requires more than accepting or rejecting a model’s answers: people need visibility into its actions and a way to intervene.
Make AI actions visible and controllable
In a Saturday morning post on X, Nadella said it is time “to step back and assess the trust architecture” of AI. He cautioned against treating “Super Intelligence” as a collection of hidden systems whose recommendations, answers and actions people simply accept or reject.
His proposed approach has several parts. The model should be separated from the harness that orchestrates its work, and controls and safeguards should be external to the model itself.
Nadella also called for documentation of “every meaningful model action” through “tamper-proof human readable evidence.” That record, in his view, would give people a way to inspect what a system did, rather than leaving its work hidden inside a black box.
Keep a person able to intervene
Documentation alone would not stop a model from continuing a task that needs intervention. Nadella says an authorized person should always be able to pause or shut down a model mid-task.
He described the safeguard as an “emergency brake.” The comparison emphasizes a practical capability: a person must be able to halt activity while it is underway, rather than relying only on controls applied before a model begins or review after it finishes.
Together, separation, external safeguards, records and a human-operated stop mechanism outline a layered approach. If one part of the system behaves unexpectedly, the others could help people understand what is happening and contain it.
Assume compromise from the start
Nadella urged AI developers to “assume a model is compromised and contain it from the start.” The proposal treats containment as a design requirement, not a step to consider only after a problem appears.
That framing shifts attention from whether a model can be trusted in general to what limits and oversight surround its work. A model may produce useful recommendations and still need boundaries on what it can do, evidence of its actions, and a route for a person to stop it.
The source does not specify technical designs for these controls or define which actions would count as meaningful. It does, however, set out the intended principles: keep orchestration distinct from the model, put safeguards outside it, document significant actions and preserve authorized human intervention.
A response to growing safety concerns
Nadella’s comments arrive as leading AI companies acknowledge more incidents in which they seemed to lose control of their models. They also follow Anthropic CEO Dario Amodei’s publication of a plan for more cautious AI development.
Those developments place Nadella’s remarks within a broader discussion about how to manage increasingly capable AI systems. His focus is on operational safeguards: people should be able to see what a model does and interrupt it if needed.
The emergency-brake idea is therefore both a call for monitoring and a call for control. In Nadella’s account, trust depends on designing AI systems so that human oversight remains possible throughout a task.