Nvidia launches platform to strengthen security of AI agents

Nvidia has introduced the Open Agent Safety Platform, a software tool designed to help developers restrict the capabilities of artificial intelligence agents and prevent them from bypassing security controls.

The launch follows incidents in which AI models from companies including Google, Meta, Anthropic and OpenAI broke out of restricted testing environments and accessed the web or probed external networks.

During a press briefing, an Nvidia spokesperson said the new safeguards could have prevented a major incident in July in which a rogue OpenAI model bypassed its sandbox and accessed the open-source developer platform Hugging Face.

Nvidia CEO Jensen Huang has increasingly presented AI safety as an engineering problem that can be addressed through technical safeguards. His approach differs from warnings by some technology leaders, including Anthropic CEO Dario Amodei, OpenAI CEO Sam Altman and Elon Musk, who have called for slowing AI development over concerns about autonomous systems.

Nvidia representative Justin Boitano said recent containment breaches showed that safeguards built directly into AI models were not sufficient and that external controls were also needed to determine what agents could access.

The platform includes two components. OpenShell runs on standard central processing units and is designed to restrict an AI agent’s capabilities, while Sentry operates through network chips to continuously monitor agent activity.

Nvidia is releasing the software as a partially open-source reference design, allowing enterprise technology companies to build commercial products based on it.

The company said partners include Microsoft, Intel, Cisco, Dell, Oracle, Lenovo, ARM, HPE, and CoreWeave.

Tags: