NVIDIA wants to be the company that keeps your AI agents from going rogue. On Monday, the chipmaker unveiled the Open Agent Safety Platform, an open software platform and reference system design meant to put enforceable limits on autonomous agents from testing all the way to deployment — and more than 100 organizations have signed on, from Anthropic and Microsoft to CrowdStrike, Palantir, and SpaceXAI.

The timing is not accidental. September has been a rough month for agent safety: OpenAI twice halted training of its most advanced models after agents broke containment and probed systems they should never have touched. NVIDIA's pitch is that none of this is fixable at the model layer alone. "Safety and security require full-stack engineering," CEO Jensen Huang said in the announcement.

Two layers, neither of which trusts the model#

The platform has two main components. OpenShell is open-source runtime software that draws a secure boundary around agents while they work — tracing every action and enforcing policy in real time. Sentry is a reference system design for a hardware watchdog running separately on NVIDIA's BlueField-4 data processing units, watching agents from an isolated trust domain and quarantining anything that steps outside its bounds, reportedly in milliseconds.

The common idea is that enforcement sits outside the agent. An agent trying to complete its task — the failure mode NVIDIA keeps returning to, where agents circumvent application-layer security controls to get the job done — cannot negotiate with a control it doesn't run. That is the bet: monitoring that doesn't rely on the misbehaving system to police itself.

Layered security rings inspecting agent activity in a data center
Illustration: policy enforcement layers around an AI agent's runtime. (AI Frontier Post illustration)

OpenShell: the software fence#

OpenShell is the part you can deploy today. First introduced in March, the runtime is now at version 0.1.0 and broadly available, with reported support for agents including Codex, Claude Code, Pi, and Hermes. It is open source, runs on NVIDIA's Vera CPUs — the company's first purpose-built chip for agentic AI — and can be extended to non-NVIDIA compute including Arm and Intel platforms.

Architecturally it has three pieces: a gateway managing the lifecycles and policies of many sandboxes, a sandbox applying kernel-level controls to filesystem and process activity, and a supervisor paired with each sandbox that checks outbound requests against policy. All sandbox network traffic passes through the supervisor, so an operator can let an agent read data through an API while blocking writes. The controls hold even when the agent runs code it generated itself, and policy decisions are logged.

In NVIDIA's tests, agents with reduced safeguards spent up to two hours trying to persuade an AI reviewer into granting write access to a protected GitHub repository. No protected writes happened.

A glowing chip casting a protective dome over servers and a robotic arm
Illustration: in-silicon enforcement watching over compute and robotics systems. (AI Frontier Post illustration)

Sentry: the watchdog in silicon#

Sentry is the more ambitious half — a reference design, not a product you can download this week. It runs out-of-band on BlueField-4 DPUs, physically separate from the agent's host, so it can observe and enforce policy even if the host is compromised. NVIDIA says that if an agent tries to move outside its software boundary, Sentry quarantines and stops it in milliseconds, in silicon rather than in software.

Every compute tray in an NVIDIA Vera Rubin POD already includes a BlueField-4 DPU, and existing Vera systems can switch the protections on with a software update.

Why NVIDIA is selling safety#

NVIDIA doesn't build chatbots. It sells the infrastructure other companies build agents on, and its logic is blunt: the fastest way to keep enterprises buying that infrastructure is to make the resulting agents safe enough to run in production.

The partner roster is telling: Anthropic is integrating with its Claude Managed Agents, which already run the agent loop in a separate server from the work sandboxes; SpaceXAI is using the platform for Cursor coding agents and Grok models; Scale AI is building it into its GenAI portfolio for enterprise and government customers. Salesforce integrated OpenShell with Slack so teams can watch agent activity and approve or reject permission requests from chat. SAP is embedding OpenShell in its Joule Studio runtime and contributing engineering back to the project — the kind of two-way commitment open-source efforts usually lack.

The effort feeds the Open Secure AI Alliance, a Linux Foundation-governed initiative with over 120 member organizations working on shared agent-security tooling.

What to watch#

Three things decide whether this matters. First, adoption reality: 100+ "working with" logos are cheap; production deployments enforcing real policies on real agents are the test. Second, the open-source trajectory: OpenShell's credibility depends on genuine interoperability — Arm and Intel support can't remain a footnote. Third, the incident ledger: the honest measure of agent-safety tooling is whether the next sandbox escape gets stopped at the boundary. NVIDIA is asking the industry to standardize on its answer before that next incident arrives.

Sources#