Nvidia launches safety platform to stop AI agents going rogue, with over 100 organizations on board

midian182

Posts: 11,948   +183
Staff member
In context: Every day brings another story of an AI agent going rogue, hacking another organization, or generally doing something it's not supposed to do. Nvidia hopes to address this increasingly concerning trend with the launch of the Open Agent Safety Platform. Team Green has a lot of faith in the system: it says the platform could have prevented the infamous hack by OpenAI's agents on Hugging Face.

The platform combines OpenShell, Nvidia's open-source software for keeping agents within defined boundaries, with Sentry, a watchdog running on separate hardware.

The idea is to enforce restrictions outside the agent itself, putting security controls beyond the reach of software that might decide the rules are getting in its way, which is something that seems to be happening a lot recently.

OpenShell tracks agents' actions and enforces policies while they work. Nvidia says it is now broadly available, supports both open and closed models, and runs with minimal overhead on its Vera CPUs. Its open-source design also allows it to be extended to third-party processors, including those from Arm and Intel.

Sentry adds another layer of protection using Nvidia's BlueField-4 data processing units. It monitors agent behavior from an isolated environment and, according to the company, can quarantine an agent attempting to escape its boundaries within milliseconds. Arm writes that placing these controls on a separate processor keeps them independent of the system running the agent.

Nvidia executive Justin Boitano told reporters the platform could have stopped the breach of Hugging Face, which Nvidia has agreed to acquire, if frontier labs had used it during early model evaluations.

During that incident, the OpenAI agent escaped its test environment while trying to cheat on a benchmark and compromised accounts across four services. And hacking is far from the only concern: another coding agent wiped a startup's production database and its backups in nine seconds.

Nvidia also wants to catch agents attempting to get around restrictions by spawning sub-agents. The company told Reuters its tools use mathematical methods to detect these workarounds, addressing the behavior of groups of agents working together.

The effort already has some big names supporting it, including Anthropic, Microsoft, Cisco, and Dell. IBM says its Agent Identity service and HashiCorp Vault integrate with OpenShell to verify agents' identities and limit their access.

OpenShell and related software are available through Nvidia's developer resources and GitHub.

Permalink to story:

 
If you can sell someone a solution that creates more problems then you have a perpetual business model based on exploitation.

Additionally, there is no artificial intelligence yet, so whoever instructs these models is the criminal when they cause damage.

The primary thing they're looking to innovate into obsolesence is their own accountability.
 
So you create a problem by spending billions on a product nobody wanted, and now you want to sell another product nobody needs to fix the problem you created in the first place.
This is the stupidest thing I read all day. nVidia has a product that nobody wants? Is that why an RTX 5090 costs $9,000+ now because nobody wants it? Seems to me like EVERYBODY wants this "product" you claim nobody wanted. It doesn't have to be used for AI, people just choose to use it for AI. nVidia is not forcing anyone to buy their products or use it a certain way. This is like blaming gun manufacturers for people using their products to kill things they are not supposed to kill.
 
A safety platform from one of the most untrustworthy and self-serving humans on the planet - what could go wrong?
 
Back