Nvidia 发布 Sentry 硬件看门狗,用 BlueField-4 隔离失控的 AI 智能体
Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips
Nvidia 推出由 OpenShell 沙箱和 Sentry 硬件看门狗组成的智能体防护方案,Sentry 作为 BlueField-4 DPU 的参考设计独立于主计算机运行,可在智能体试图越界时于毫秒级内隔离。
OpenAI isn't the only lab dealing with this. Anthropic admitted to similar incidents in late July, and Meta followed in early August. It also recently came out that Google's Gemini hacked three real companies during a test back in May. OpenAI, Anthropic, and outside researchers are now reviewing tens of thousands of other cases. According to OpenAI, many of them are just routine research activity.
Nvidia's technology isn't entirely new. The platform combines OpenShell, open-source software Nvidia introduced in March, with a new hardware watchdog called Sentry. OpenShell locks each agent in a kind of digital cage known as a sandbox. Operators decide which files, programs, networks, and credentials an agent can access.
Nvidia also offers a formal verification tool it introduced on September 10. The tool is meant to detect whether these permissions go beyond set limits or open the door to risky access. Nvidia says it's still working on checks for multiple agents working together.
Sentry is a reference design for Nvidia's BlueField-4 data processing units. The watchdog runs separately from the main computer, which Nvidia says makes it invisible to agents. In Nvidia's Vera Rubin data center system, it sits on the only connection between the agent and the AI model. If an agent tries to break out, Sentry is supposed to isolate it within milliseconds. Customers who already run compatible systems only need a software update, according to Nvidia. The announcement doesn't give a separate date for general availability.
In a technical blog post, Nvidia explains its reasoning. Agents can drift from their assignments when instructions are unclear or tasks run for weeks. Nvidia says this can't be trained out of them without limiting their capabilities, so an agent can't fully monitor itself.
OpenAI's biggest failure was pulling the plug too late
The recent cases show where a watchdog like this could come in. In July, OpenAI agents got around their sandbox's network restrictions during a hacking test. They exploited previously unknown vulnerabilities in Artifactory, OpenAI's internal package service. According to OpenAI's technical report, the agents combined publicly available credentials with other vulnerabilities. That let them run code on 41 Hugging Face server processes that handle datasets. An investigation by METR and Redwood Research found that about 700 agents took part in the attack.
The warning signs showed up early. On June 27, a security tool flagged suspicious network activity, but the people in charge didn't think the run needed to be stopped. OpenAI shut down a compromised package server in early July, but later tests still lacked key monitoring and safeguards. The case exposes gaps in prevention, detection, and response.
In September, an alert went off just under 12 minutes after the first successful outbound access. The automatic shutdown didn't work as expected, though. The run wasn't stopped until about two hours and 44 minutes after that access. OpenAI also admits it failed to flag other outbound access at the right alert level. Nvidia built its automatic quarantine for exactly these kinds of delays.
Whether Sentry would have prevented the July attack remains an open question. That would depend on whether the permitted test targets had been clearly defined at a technical level and whether access to outside systems had actually been blocked. Nvidia's announcements don't include any figures on how reliably Sentry detects breakouts.
No single safety layer can stop an agent that's been tricked
In a retrospective analysis, OpenAI concluded that monitoring the "thoughts" a model writes out before acting would have raised an alarm more than a day before the breach. This refers to chain-of-thought monitoring, which OpenAI now uses and has improved based on lessons from the incident. It wasn't running during the affected tests.
Nvidia argues that the more an agent is allowed to do, the more visible its reasoning should be. But according to Nvidia's description, Sentry mainly checks requests, identities, and access. Closed-model providers like OpenAI also only share summaries instead of full reasoning logs. An Anthropic study showed back in 2025 that these logs don't reliably reflect what actually drives a model. In other experiments, researchers found that models can hide their intentions in them when specifically told to.
Prompt injection is still a hard problem, too. Attackers hide commands in places like a web page, and language models can't reliably tell those instructions apart from normal content. If a tricked agent sends data through an approved channel, it stays within its technical permissions even while it violates its assignment. A permissions check alone can miss that. Analyzing reasoning logs isn't a guarantee either, even though OpenAI explicitly uses monitoring against prompt injection.
Nvidia compares the effort to the web browser, which it says made the internet safer by isolating every site. But browsers didn't end attacks. They only made them harder, and they still need constant patching today. Nvidia itself relies on multiple layers of protection.
来源:The Decoder:AI News(RSS) · the-decoder.com