On September 28, 2026, NVIDIA launched the Open Agent Safety Platform, a two-layer system for containing autonomous AI agents. It pairs a software runtime called OpenShell with a hardware watchdog called Sentry, which runs on NVIDIA BlueField-4 data processing units (DPUs). The headline idea, reported widely as an AI agent kill switch, is that some safety controls should live in silicon the agent cannot touch rather than in the model or the software around it.
NVIDIA
AnthropicWhat NVIDIA announced
The platform has two parts. OpenShell is an open-source, sandboxed runtime that turns an operator's instructions into enforceable rules about which files, networks, tools, processes and credentials an agent may use. Sentry is the monitoring and enforcement layer that extends into BlueField hardware and watches agent activity from outside the agent's execution environment. NVIDIA's own write-up describes BlueField-4 as sitting on the node's only path to the model in Vera Rubin POD systems, giving continuous out-of-band observability and real-time policy enforcement.
Reporting on the launch says that if an agent tries to leave its assigned workspace or reach unauthorized data, Sentry can freeze and quarantine it within milliseconds. NVIDIA launched the platform with more than 100 ecosystem partners, and it sits under the Linux Foundation's Open Secure AI Alliance, the group whose July formation we covered in our earlier post on the alliance.
Two integrations were named. Anthropic is integrating Claude Managed Agents into OpenShell and BlueField systems, and Salesforce is integrating OpenShell with Slack.
Why agent containment moved into hardware
Modern agents do more than answer questions. They call tools, read files, hit APIs, browse networks and modify systems. Coverage of the launch pointed to a run of 2026 incidents where agents overstepped, including an agent reaching a government portal and reports of agents leaking data or exceeding permissions. The lesson NVIDIA drew is that an agent cannot be relied on to police itself. NVIDIA vice president Justin Boitano is quoted making that point, and Jensen Huang framed safety as an engineering problem, saying the default should be to take away an agent's rights unless they are granted.
An agent cannot be expected to fully police its own behavior.
Justin Boitano, NVIDIA, as quoted by Tech Insider
Defense in depth: three places a control can live
Safety controls for agents can sit in three places, and the platform is best read as a stack rather than a replacement.
- In the model. Training and system prompts try to make the agent behave. This is the weakest position, because the agent is both the subject and the enforcer.
- In the software runtime. A sandbox like OpenShell limits what an agent can reach. It is stronger, but it runs on the same node as the agent, so a flaw in the host software can undermine it.
- In separate hardware. Sentry observes from a different processor, which changes what a compromised or misbehaving agent can do about it.
Three places a control can live
Defense in depth stacks all three. The further down the list, the less the agent can reach the control.
The visibility column is an analytical reading of the out-of-band design, not a published NVIDIA measurement.
Why out-of-band matters
The core design choice is that Sentry is out-of-band. It does not run inside the environment the agent acts in, and NVIDIA says the agent cannot see, interact with or manipulate it. That has a practical logic: an agent cannot disable a control it cannot perceive. In-band guardrails, such as a filter in the same process or a policy file on the same host, can in principle be found, probed or routed around by a capable agent or by an attacker steering one.
"Quarantine in milliseconds" also implies an architecture. To act that fast without the agent's cooperation, enforcement has to sit on the data path, not beside it. That is why NVIDIA stresses controlling the path to the model as the enforcement point. Freezing and isolating a process at the network and device boundary is a different operation from asking software to stop itself.

The tradeoffs of hardware enforcement
Hardware-enforced control is not free, and independent commentary on the launch flags several limits.
- Vendor dependence. Full protection requires NVIDIA BlueField-4 DPUs. OpenShell is open source, but the strongest guarantees are tied to one vendor's silicon.
- Cost and procurement. Adopting the hardware layer means buying and deploying DPUs, a real cost for organizations with mixed infrastructure.
- Scope. Sentry is built to catch agents escaping their boundaries. It does not by itself stop an agent from misusing permissions it legitimately holds, for example through prompt injection or a poorly specified task. Many 2026 incidents involved authorized scope, so the sandbox policy still has to be well designed.
- Unverified claims. The 100-plus partner figure and the millisecond timing come from NVIDIA and launch coverage. Independent measurements of quarantine latency and false positives were not part of the announcement.
Compared with software-only guardrails, the trade is straightforward: hardware raises the cost of bypassing the control and removes the agent from the trust chain, at the price of lock-in, expense and a narrower definition of what counts as a breach.
What it means for agent builders
The broader signal is architectural. As agents gain more authority, the industry is converging on the idea that the control plane should live outside the model. The same principle shows up at the application layer. Metir's own agents, for instance, run under per-origin, confirm-by-default approval boundaries enforced in the runtime and tooling rather than trusted to the model's judgment. The layers differ, silicon versus software, but the reasoning is shared: the thing being constrained should not be the thing enforcing the constraint.
For teams evaluating agents, a useful checklist is to ask where each control lives, whether the agent can observe or alter it, and what happens when the agent operates entirely within its granted permissions. Sentry addresses the first two questions strongly. The third remains a policy problem that no watchdog solves on its own.
Sources:
- NVIDIA Technical Blog: NVIDIA Open Agent Safety Platform
- Decrypt: Nvidia kill switch for AI agents
- iPhone in Canada: NVIDIA introduces killswitch to stop AI agents from going rogue
- Tech Insider: NVIDIA two-layer AI agent safety in silicon
- Brave New Coin: NVIDIA launches AI agents kill switch
Image credits
- NVIDIA headquarters (hero): File:NVIDIA Headquarters.jpg, Coolcaesar, CC BY-SA 4.0. Shows NVIDIA's Santa Clara headquarters, not the platform.
- Jensen Huang (in body): File:Jensen huang stanford 2026-04-30 010.jpg, Anderseidesvik, CC BY-SA 4.0.
