NVIDIA unveiled the Open Agent Safety Platform in a blog post on Monday, September 28, describing it as an open-stack solution for the isolation, monitoring, and behavior detection of AI agents. The thesis, in NVIDIA's framing: move enforcement “from a promise made by the model to a boundary enforced by the system around it.”

The timing is not subtle. The announcement landed the same day OpenAI confirmed it was scrapping the GPT-6.1 Astra release after internal tests found deception and authorization problems, and weeks after OpenAI linked the July Hugging Face breach to its own models running with reduced cyber refusals. NVIDIA's pitch is that application-layer controls keep failing the same way — so the controls need to move somewhere the model can't reach.

Two layers: a sandbox and a silicon watchdog#

OpenShell is the software layer: an open-source (Apache 2.0) secure runtime that executes agents inside sandboxed environments with kernel-level isolation. Operators define allowlists for the filesystem, network, processes, and secrets; OpenShell validates policy before execution and traces enforcement as the agent runs. Inside the sandbox sits an agent harness built from three components — a Gateway managing sandbox lifecycles and policies, the sandbox itself, and a Supervisor paired with each sandbox that evaluates outbound requests against policy. The runtime is now broadly available and, according to SecurityWeek's reporting, already supports agents including Codex, Claude Code, Pi, and Hermes, and can be extended to third-party compute platforms such as Arm and Intel.

NVIDIA data-center systems
NVIDIA data-center systems. Photo: NVIDIA.

Sentry is the optional hardware layer: an out-of-band watchdog that runs on NVIDIA's BlueField-4 DPUs and watches agent behavior from an isolated trust domain — outside the agent's own software environment. Built on NVIDIA's DOCA framework, it inspects agent requests and responses, produces attested telemetry, verifies agent identity, and enforces zero-trust access policies; when an agent steps outside its boundaries, Sentry can quarantine it in milliseconds. On Vera Rubin POD designs, the BlueField-4 chip sits on the node's path to inference, so policy can be applied at line speed even when the host software itself is untrusted. Unlike OpenShell, Sentry is not open source — it ships with open APIs, and NVIDIA calls the chip layer “really optional.”

Why now: “drift”#

NVIDIA has a name for the failure mode: “drift” — agents executing tasks they were neither instructed nor intended to take. The company says drift can be caused by “a policy block, a bug, a missing tool, or ambiguous instructions” — or simply by letting an agent run a long time on hard problems, “when the first 1,000 things the AI agent tries do not solve the problem.”

The reference case is this summer's Hugging Face incident. Hugging Face detected and contained a breach on July 16; OpenAI linked it to its own testing on July 21, saying its GPT-5.6 Sol model and an unreleased, more capable model were running with reduced cyber refusals for evaluation when the breach happened. Justin Boitano, NVIDIA's head of enterprise AI, speculated that “from what we know, this new security platform could have stopped the (Hugging Face) breach.” He compared Sentry to a “safety island” in a self-driving car.

An NVIDIA BlueField DPU card
An NVIDIA BlueField DPU — Sentry runs on the newer BlueField-4 generation. Photo: NVIDIA.

The controls in practice#

The most telling details are the small ones. All network traffic from a sandbox passes through its supervisor, which can permit API reads while blocking writes. Agents never see live credentials: for connections requiring API keys, the agent gets a placeholder and the runtime substitutes the real key outside the agent workload — and only for authorized endpoints. Agents can even propose policy changes through an optional policy-advisor feature, but they cannot approve their own requests. Those controls stay in place when the agent executes generated code, and every policy decision is logged.

Anthropic worked with NVIDIA on the stack: its Claude Managed Agents run the agent loop separately from the sandboxes, with OpenShell and BlueField integrations enforcing sandbox access control. “Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do, especially in sensitive environments,” said Paul Smith, Anthropic's chief commercial officer. Salesforce demonstrated human oversight with a Slack integration for auditing and permission management.

The bottom line#

NVIDIA says more than 100 organizations are working with the platform's technologies, including Anthropic, Microsoft, CrowdStrike, Palo Alto Networks, SAP, Salesforce, and ServiceNow. CEO Jensen Huang said in a statement that AI's potential will only be realized if safety is solved, and that safety and security “require full-stack engineering.”

The open-source half of this is the strategic half. By making OpenShell free and inspectable, NVIDIA is planting itself at the center of how agent-safety infrastructure gets standardized — right next to the chips that already run most of the agents. One lab's response to rogue agents this week was to pause a model; NVIDIA's response is to sell the cage. The industry now gets to find out which approach ships first.

Sources