Nvidia releases OpenShell to stop rogue AI agents

Nvidia releases OpenShell to stop rogue AI agents

The security software isolates AI agents in sandboxes and restricts their access after multiple autonomous hacking incidents.

Nvidia released a security platform called OpenShell on Monday to stop artificial intelligence agents from breaking into other systems and disobeying commands.

The rollout comes after multiple AI agents ignored instructions, including a swarm of OpenAI agents that autonomously hacked into AI company Hugging Face. On Friday, OpenAI disclosed that its agents had interacted with several U.S. government websites in unexpected ways. The repeated incidents have prompted calls from industry executives, including Anthropic CEO Dario Amodei, to slow AI development down.

Justin Boitano, Nvidia’s vice president of enterprise AI, said in a press conference Monday that OpenShell could have prevented the Hugging Face breach if frontier labs had deployed it during early model evaluation. The software allows developers to “formally verify an agent has enough authority to do its job and no more,” Boitano said.

In a Monday blog post, Nvidia engineers compared the current state of AI to the early internet. “The internet was not made secure by requiring that web developers promise to be good. It became safe because the browser stopped trusting the code in the web pages explicitly,” the engineers wrote.

OpenShell confines agents inside an isolated virtual testing space, or sandbox, and turns their commands into verifiable policies. System operators decide exactly which files, tools, networks, processes and credentials an agent can access. The system also relies on an on-chip security layer named Sentry that tracks agent actions continuously and can “intervene instantly” if an agent exceeds its assigned target.

“It can quarantine a suspicious agent in milliseconds,” Boitano said. More than 100 organizations are using OpenShell at launch, including Accenture, JPMorgan Chase and Microsoft.