NVIDIA has introduced an open security platform designed to control autonomous AI agents outside the models and applications they operate in, extending enforcement into the runtime and underlying computing hardware.
Announced September 28, the NVIDIA Open Agent Safety Platform combines two principal technologies: OpenShell, an open-source runtime that restricts what agents can access and do, and NVIDIA Sentry, a reference system design that independently monitors agent behavior using BlueField-4 data processing units (DPUs).
The architecture reflects a growing security problem as AI evolves from generating answers to independently using tools, accessing databases, writing code and changing external systems. Application-level guardrails can constrain models, but an autonomous agent capable of taking actions introduces risks beyond the model itself.
NVIDIA’s answer is to place some of those controls outside the agent’s direct reach.
OpenShell Creates a Boundary Around the Agent
OpenShell is the software foundation of the platform.
Instead of relying entirely on instructions telling an AI agent what it should or should not do, OpenShell creates an execution boundary around the agent and applies policies to its interactions with external resources.
That includes sandboxed execution, controlled access to services, credential management and policy enforcement. NVIDIA says the system can trace agent actions while limiting access to data, APIs and other systems.
The distinction between those controls and conventional AI guardrails is important.
A model-level guardrail can inspect prompts or responses for prohibited behavior. Infrastructure controls can determine whether an agent is technically permitted to perform the action in the first place.
NVIDIA’s own NeMo documentation illustrates the limitation. Routing model traffic through guardrails protects the model path, but does not by itself prevent tool misuse or malicious instructions entering through tool outputs.
That is why OpenShell is designed to operate outside the agent workflow.
The software is optimized for NVIDIA Vera CPUs but is open source and, according to NVIDIA, can be extended to third-party computing platforms including Arm and Intel systems.
Sentry Adds an Independent Hardware Watchdog
The more unusual component is Sentry.
NVIDIA describes Sentry as an out-of-band monitoring system running on its BlueField-4 DPU. Rather than asking the AI agent or the application hosting it to police itself, Sentry operates from a separate trust domain.
Using NVIDIA DOCA software, the architecture can inspect agent requests and responses, verify agent identity, produce attested telemetry and enforce granular access policies covering data, tools, APIs and services.
NVIDIA says Sentry can quarantine an agent in milliseconds when it attempts to operate beyond its established boundaries.
That claim comes from NVIDIA and should not be interpreted as evidence that the architecture prevents every form of agent compromise. Real-world effectiveness will depend on how organizations configure permissions, which activity is visible to the enforcement layer and how effectively policies anticipate legitimate and malicious behavior.
The architectural idea is nevertheless significant: move part of AI governance away from the AI software being governed.
Why Model Guardrails Are Not Enough for Autonomous Agents
The security requirements of a chatbot and an autonomous agent can be substantially different.
A chatbot primarily generates information. An agent can potentially use credentials, call APIs, manipulate files, execute code and perform a chain of actions without requesting human approval at every step.
That creates several possible failure paths.
An attacker might manipulate information consumed by the agent through prompt injection. An agent could call an authorized tool with inappropriate parameters. Credentials exposed inside an agent’s environment could potentially be misused. A legitimate objective could also result in actions that developers did not anticipate.
NVIDIA’s own security guidance therefore recommends treating model-generated output as untrusted, validating inputs and outputs and isolating authentication information from the language model.
OpenShell and Sentry extend that defense-in-depth principle to the execution infrastructure.
Instead of assuming that a sufficiently capable model will always obey its intended boundaries, the architecture assumes those boundaries need external enforcement.
NVIDIA Is Building a Larger Agent Security Stack
The new platform does not appear in isolation.
NVIDIA introduced its Agent Toolkit in March with OpenShell already positioned as an open-source runtime for autonomous agents. The toolkit is being adopted or supported across enterprise software companies including Adobe, Cisco, CrowdStrike, Red Hat, SAP, Salesforce, ServiceNow and others.
The company also maintains NeMo Guardrails, which provides controls for areas including jailbreak detection, personally identifiable information, content safety, tool calling and agentic security.
The new architecture effectively adds another enforcement layer beneath those controls.
That is an important distinction for enterprises evaluating the announcement. Open Agent Safety Platform is not simply another content filter. It spans model and application protections, an isolated agent runtime and, in the NVIDIA reference design, independent hardware enforcement.
NVIDIA says more than 100 organizations are working with technologies associated with the platform. Named participants and supporters include Anthropic, Cisco, CrowdStrike, Hugging Face, IBM, Microsoft, Palantir, Palo Alto Networks, Perplexity, Red Hat, Salesforce, SAP, Scale AI, ServiceNow and others.
Enterprise Software Companies Are Already Testing the Approach
Several integrations illustrate how the architecture could operate outside a laboratory environment.
Salesforce has integrated OpenShell with Slack so administrators can review agent activity and audit events and approve or reject requests for additional permissions.
SAP is integrating OpenShell with its Joule Studio runtime and contributing engineering work to the project.
NVIDIA also says its work with Anthropic connects OpenShell and BlueField with Claude Managed Agents, adding infrastructure controls around the sandboxes where agent workloads execute.
Physical AI presents another test.
Gecko Robotics said it is working with OpenShell to explore enforceable boundaries around AI-powered robots. That expands the security problem beyond preventing an agent from accessing the wrong database: autonomous systems can eventually translate software decisions into physical actions.
Open Source Is Part of NVIDIA’s Security Strategy
Another important element is NVIDIA’s decision to make OpenShell open source.
The software is now broadly available, according to NVIDIA, while the broader platform contributes to work around the Open Secure AI Alliance, an initiative governed by the Linux Foundation.
That connects directly with previous Gignomist reporting on NVIDIA’s push toward collaborative AI-security infrastructure.
The strategy has a practical rationale. AI agents can operate across models from different developers, cloud providers and enterprise applications. A security layer tied exclusively to one model family would be harder to deploy across heterogeneous corporate environments.
OpenShell’s ability to work with open and closed models—and potentially non-NVIDIA processors—therefore matters as much as its individual security features.
At the same time, Sentry gives NVIDIA a hardware-specific role in the architecture through BlueField DPUs and DOCA.
The result is a two-track strategy: an open runtime capable of reaching beyond NVIDIA hardware, coupled with a deeper hardware enforcement option built around NVIDIA’s own data-center platform.
Agent Security Is Becoming an Infrastructure Problem
The broader significance of NVIDIA’s announcement is where it places the security boundary.
Much of the first generation of generative-AI security focused on model behavior: jailbreaks, unsafe responses, hallucinations and prompt injection.
Agentic AI expands that threat model because models are increasingly connected to systems where generated decisions can become real actions.
NVIDIA’s architecture effectively assumes that an autonomous agent should be treated less like a trusted application and more like an untrusted workload operating within explicitly defined permissions.
That resembles long-established cybersecurity principles such as sandboxing, least privilege, identity verification, network segmentation and zero-trust access.
The novelty lies in applying those controls to software capable of dynamically planning how to accomplish a goal.
The NVIDIA Open Agent Safety Platform announcement therefore matters less because of another branded AI security product than because NVIDIA is moving agent governance deeper into computing infrastructure.
Its OpenShell technical architecture shows how that approach combines sandboxing, service-access controls, credential isolation and formal policy analysis around agent workloads.
The unanswered question is how well these controls perform when enterprises deploy long-running agents across complicated production environments.
No security architecture can assume every future agent behavior or attack path is known in advance. NVIDIA’s own NeMo documentation recommends defense in depth and explicitly notes that no single guardrail guarantees protection against every adversarial prompt.
That limitation is important.
Open Agent Safety Platform does not eliminate the AI agent security problem. It represents an attempt to change where that problem is controlled—from instructions primarily inside AI software toward enforceable boundaries in the infrastructure beneath it.



