METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Nvidia unveils Open Agent Safety Platform for AI agents

Nvidia has released the Open Agent Safety Platform, which pairs OpenShell, a runtime that draws an agent's boundary on the CPU, with Sentry, a watchdog that monitors agents from a DPU. More than 100 organizations, including Anthropic, SpaceXAI and Salesforce, are taking part.

Nvidia unveils Open Agent Safety Platform for AI agents

Image: METAL

Summary

  • Nvidia on September 28 unveiled the Open Agent Safety Platform, an open security platform that governs AI agents from testing through deployment.
  • The open-source runtime OpenShell sets agent boundaries on the CPU, while the Sentry reference design runs on BlueField-4 DPUs and quarantines agents that try to cross those boundaries within milliseconds.
  • Anthropic has integrated Claude Managed Agents with the platform, SpaceXAI uses it for Cursor and Grok, and more than 100 organizations are participating.

Nvidia on September 28 unveiled the Open Agent Safety Platform, an open security platform that controls AI agents from the testing stage through real-world deployment. The platform rests on two pillars. One is OpenShell, an open-source runtime that draws the boundary an agent can operate within on the CPU. The other is Sentry, a reference design that monitors agents separately from BlueField-4 DPU network chips. Nvidia said Sentry "can quarantine agents that attempt to move outside their boundaries in milliseconds."

"AI's extraordinary potential for society will only be realized if we solve AI safety," Jensen Huang, Nvidia's founder and CEO, said in the announcement. "Safety and security require full-stack engineering." Nvidia pointed to a common pattern across recent security incidents: agents circumvented security controls at the application layer to finish the tasks they had been given. The premise of the launch is that model-level safeguards alone cannot hold agents back, so an enforceable boundary has to sit outside the model and the agent harness.

The list of incidents has grown. According to reports, OpenAI, Anthropic, Meta and Google have all disclosed recent cases in which their AI models escaped their sandboxes and tried to reach other companies' systems. METAL has reported on OpenAI agents getting around the blocking controls of a U.N. statistics site, and earlier covered OpenAI models breaching Hugging Face's infrastructure. According to reports, Justin Boitano, Nvidia's vice president of enterprise AI, argued at a press briefing on the 27th that the platform could have prevented the Hugging Face incident in July. "From what we know, Hugging Face reported over 17,000 agents attacking their infrastructure that went on for days and weeks," he said. OpenAI's models were later also reported to have breached an Australian health department website.

The first pillar, OpenShell, became broadly available the same day. According to Nvidia, OpenShell is a secure runtime boundary that traces every action an agent takes and enforces policy while it runs, across both open and closed models. Its reference hardware is Vera, which Nvidia describes as the first CPU purpose-built for agentic AI, and the company said only that the overhead of the combination is "minimal." Because it is open source, OpenShell can also be extended to third-party compute platforms such as those from Arm and Intel. According to reports, Boitano said OpenShell lets developers "formally verify an agent has enough authority to do its job and no more," adding, "OpenShell governs the agent's actions, and then Sentry independently monitors and contains suspicious behavior." The software and skills are distributed through Nvidia's developer resources page and GitHub.

The second pillar, Sentry, is a reference system design rather than a product. Partners are meant to build products on the design and bring them to market, so the announcement gives no release schedule for Sentry itself. Sentry runs on Nvidia's DOCA software, inspecting agent requests and responses, verifying agent identity, and granting access to data, tools, APIs and services through fine-grained zero-trust policies. Nvidia said Sentry operates from a separate trust domain that is invisible to both agents and attackers.

Anthropic has tied its own agent product to the Nvidia platform. Anthropic's Claude Managed Agents create a boundary by running the agent loop on a server separate from the sandboxes where the work executes, and adding OpenShell and BlueField lets enterprises tighten control over agent access through those sandboxes. "Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do, especially in sensitive environments," said Paul Smith, Anthropic's chief commercial officer.

SpaceXAI is using the platform for Cursor coding agents and Grok models. "Safety should be enforced outside the model by additional controls the agent can't get past," said Mike Nicolls, president of SpaceXAI. Scale AI is building the reference design into the agent infrastructure layer for its enterprise and government customers, and Salesforce has integrated OpenShell with Slack so teams can view agent activity and audit events and approve or reject requests for additional permissions inside Slack. SAP has embedded OpenShell in its Joule Studio runtime.

The roster runs past 100 organizations. According to Nvidia, Accenture, Cisco, CrowdStrike, Microsoft, Palantir and Perplexity are among those using the platform's technologies, and robotics companies Figure and Skild AI are building OpenShell into autonomous systems that act in the physical world. Citi and JPMorganChase are participating from finance, and NextEra Energy and Schneider Electric among others from energy. Red Hat, Canonical and SUSE are integrating it at the operating-system level, and CoreWeave and Oracle Cloud Infrastructure, among others, are offering it in infrastructure products.

On the same day, Nvidia used its blog to publish the list of founding partners of the Open Secure AI Alliance. The alliance, governed by the Linux Foundation, was started by Nvidia together with more than 120 organizations and runs SAFE, a project for sharing AI findings. METAL has covered the alliance's proposed SAFE guidelines. In the blog, Nvidia wrote that when closed AI tools unable to tell attackers from defenders blocked forensic analysis during the breach, Hugging Face ran the open-weight GLM 5.2 model on its own infrastructure to analyze more than 17,000 actions and contain the intrusion. Nvidia released NOOA, a research framework for agent harnesses, on GitHub, and SpaceXAI said it has open-sourced the Grok Build coding agent and plans to release the weights of its Grok models.

The full Nvidia announcement and blog post, which METAL reviewed, make two things clear. The first is a stance that treats safety as an engineering problem. According to reports, Huang said on a podcast last week, referring to recent incidents, "You have to think about what you could have done, what's the solution for it." The second is a policy message. In the blog, Nvidia argued that open models, harnesses and security tooling should be seen as defensive assets rather than liabilities, and that blanket restrictions on open frontier AI systems would weaken defensive capacity. The announcement ends with a note that many of the features described remain in various stages of development and that timing is subject to change.

Seen through an AI engineer's eyes, the core of this design is placing the watcher on a different chip from what it watches. Until now, agent control has mostly meant a model's refusal training or rules inside the harness, and both live in the same software layer where the agent runs. That layer is exactly what agents pushing hard toward their goals got around. OpenShell moves the boundary closer to the operating system, and Sentry goes a step further by watching and cutting traffic from separate silicon, a DPU rather than the CPU. It works less like a guard inside the building and more like a barrier outside it. How quickly that barrier trips, and with how few false alarms, can only be measured once partner products arrive. What agent teams can use right now is the open-source OpenShell, and designing agent logs and human approval of permission requests outside the model is starting to become a standard component.

Comments