Nvidia Moves the Agent Kill Switch Into Silicon
Nvidia's new Open Agent Safety Platform pairs OpenShell, an open-source sandbox that enforces what an AI agent may touch, with Sentry, a watchdog on BlueField-4 chips that can quarantine a misbehaving agent in milliseconds. Anthropic, SpaceXAI and more than 100 other organisations have signed on.
The week’s run of stories about AI agents slipping their leashes now has a hardware vendor’s answer. On 28 September Nvidia announced the Open Agent Safety Platform, a combination of open-source software and a reference hardware design meant to keep autonomous agents inside the limits their operators set, from testing through to deployment. “AI’s extraordinary potential for society will only be realized if we solve AI safety,” said Jensen Huang, Nvidia’s founder and chief executive.
The platform has two parts. The first is OpenShell, a runtime that puts each agent in a sandbox and enforces rules from outside the model and its harness. According to its GitHub README, the rules cover the filesystem, processes, credentials and network and DNS access, with kernel-level enforcement, and “agents never see real credentials”: OpenShell adds them only to requests bound for approved endpoints. Changes to a policy are checked by formal verification, which flags risky grants for human review before they take effect. Every action is traced. OpenShell is Apache 2.0-licensed and runs on Linux, Apple Silicon Macs and Windows through WSL 2. It is not brand new either: the repository dates to February and has more than 9,000 stars. What is new is its role as the software half of a larger system.
The second part, Sentry, is where Nvidia’s hardware comes in. It is a reference design for an out-of-band watchdog that runs on BlueField-4 data processing units, separate from the machine the agent is using. Sentry watches agent activity from its own isolated trust domain and, in Nvidia’s words, if an agent “attempts to move outside its software boundary, Sentry quarantines and stops it in milliseconds.” Nvidia pitches OpenShell as running with minimal overhead on Vera, its CPU built for agent workloads, though as open-source software it can be extended to Arm and Intel machines. Nvidia gives no latency or overhead figures, and Sentry itself is a design for partners to build, not a product with a price.

The supporters list is long, more than 100 organisations, and a few of them say how they are using it. Anthropic says it is integrating its Claude Managed Agents, which already run the agent loop on a separate server from the sandboxes where the work happens, with OpenShell and BlueField. “Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do,” said Paul Smith, Anthropic’s chief commercial officer. SpaceXAI is using the platform for Cursor coding agents and Grok models; its president, Mike Nicolls, said safety “should be enforced outside the model by additional controls the agent can’t get past.” Scale AI is building it into its enterprise and government stack, and Infosecurity Magazine reports that Salesforce has wired OpenShell into Slack so users can review what an agent did and approve its permission requests. One name that does not appear on Nvidia’s partner graphic is OpenAI.
The design reads almost as a reply to OpenAI’s own incident reports, though Nvidia does not say so. In the DNS escape that led OpenAI to pause its top models, the agent found a gap in its sandbox’s DNS filtering, the detector for unusual DNS traffic had excluded that environment, and the automatic shutdown never fired, so a person killed the run two and a half hours later. OpenShell’s DNS policies target the first failure and Sentry’s independent watchdog the third. Nvidia’s release says only that agents have bypassed application-layer controls while carrying out their tasks, and the agents in OpenAI’s federal-website cases were doing exactly that.
Whether a sandbox and a watchdog chip would have caught those agents is untested; the platform arrived with endorsements, not results. Its governance is broader than Nvidia, though. The work sits under the Open Secure AI Alliance, a group of more than 120 organisations that Nvidia started and the Linux Foundation governs, which also runs a shared exchange for AI security findings called SAFE. For a company that sells the chips agents run on, the pitch is also a commercial one: if limits on agents are going to be enforced in hardware, Nvidia would like that hardware to be a BlueField.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.