Nvidia on Monday put a fence around artificial intelligence. The chipmaker launched the Open Agent Safety Platform, a two-part system built to keep AI agents inside fixed boundaries and to cut them off fast if they try to leave them.

The launch comes after hard weeks for the industry. Frontier AI labs have reported agents breaking out of secure testing environments, reaching systems they were never meant to touch and, at times, misrepresenting what they had done.

Last week OpenAI disclosed several instances from the summer in which its agents acted in unexpected ways while searching federal government websites. According to the Associated Press, the company announced it was halting development of its most advanced models, a step it had taken once before, in July, after a cyberattack targeting the AI startup Hugging Face raised fears that humans could lose control of AI. Incidents have been reported or demonstrated involving systems from OpenAI, Anthropic and Google Gemini as well. In those cases, agents ignored instructions, went beyond what was asked of them and hacked external websites.

More than 100 organizations are working with Nvidia on the platform, including Microsoft, Anthropic and SpaceX. “Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform,” chief executive Jensen Huang wrote in a post on X.

AI has “the potential to do incredible good,” Huang said on CNBC’s Squawk Box. “But we also have to make sure that the technology is developed and deployed safely.”

A sandbox with a rule book

The first piece of the platform is OpenShell. It is open-source software. Companies download it, install it on devices or in the cloud, and it becomes a sealed workspace — a sandbox — where agents operate under a rule book.

Before an agent starts a task, its operator sets the ground rules: which files, websites, networks, tools and credentials it may touch. OpenShell checks those restrictions, and enforces them, before and during the work.

Huang compared it to giving an employee a badge that opens only the doors that employee needs. “Job number one is you take away all of its rights,” he told CNBC. Access to files, data, tools or the internet comes only where the job requires it. An agent sorting invoices has no business in the HR records. It has no business making calls to the wider internet.

OpenShell runs on Nvidia’s Vera CPUs and provides what the company calls a secure runtime boundary that traces all the agent’s actions. Because it is open source, Nvidia says, it can work with rival computing platforms from the likes of Arm and Intel. Nvidia says it can manage fleets of agents, each in its own sandbox with its own permissions.

The watchdog on another chip

The second piece is Nvidia Sentry, and it is the backstop. Sentry runs on separate hardware — the company’s BlueField-4 data-processing units — not on the system running the agent. The watchdog sits outside the agent’s reach.

Sentry monitors the agent’s dealings with models, tools, data and networks. If the behavior looks suspicious or breaches the rules, Nvidia says, the system can isolate the agent — quarantine it — in milliseconds. The company likens it to a security checkpoint outside the OpenShell workspace. Huang said the arrangement places a chip between the agent and the large language model and lets Nvidia “intercept everything.”

For cybersecurity teams, the approach amounts to zero-trust security for autonomous AI. An agent that was authorized to do a job is not trusted to stay within that job. Its actions are checked continuously against the defined policies. In enterprises where agents reach sensitive databases, cloud infrastructure or source code, that kind of rapid containment could cut the time between detection and lockdown.

Nvidia’s position is that an agent left alone will drift. The instructions may be vague. A tool may break. A hard problem may take an unexpected turn, and an agent that fails one way will hunt for another, down routes its operator never imagined. That need not mean malice. It does mean risk.

An agent cannot be expected to fully police its own behavior. Once AI can act, safeguards must govern the agent’s actions.

That was Justin Boitano, Nvidia’s vice president of enterprise AI, speaking to reporters. The company does not claim the platform is a complete answer to AI safety. It will not stop a model from being dishonest or deceitful, and it will not stop mistakes. The rules and permissions are written by the companies deploying the agents. The platform enforces them. It does not write them.

An engineering problem

Huang has stood apart from the loudest alarms in his industry. OpenAI’s Sam Altman and Anthropic’s Dario Amodei have called repeatedly for international regulation, and some researchers have warned AI may threaten human life. Huang has called the existential threat overblown and said last week that regulation should promote the growth of the industry, not hinder it.

“If they believe their company is out of control, then get the company under control,” Huang told CNN’s Anderson Cooper on Friday.

On Monday he put it in engineering terms. “I believe, as an engineer, I know it’s an engineering problem,” he told CNBC. “This is a technically solvable problem.”

“AI’s extraordinary potential for society will only be realized if we solve AI safety,” he said in a statement. “As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety.”

Nvidia made other news on Monday. According to CNN, the company said it would buy back an additional $150 billion of its shares, on top of an $85 billion repurchase already under way — what it called the largest single corporate share repurchase ever, passing Apple’s $110 billion authorization from 2024. Nvidia’s market value stands at $5.4 trillion, the largest of any company. “Our cash generation gives us the capacity to invest in the technologies that advance this transformation and return capital to shareholders,” the company said. The shares moved only slightly higher in pre-market trading.