Imagine you run a bank, and you install a big red button on the wall labelled IN CASE OF EMERGENCY, STOP EVERYTHING. One day there’s an emergency. Someone detects it within 15 minutes, acknowledges the alarm three minutes later, and presses the button. The button does not work. The everything keeps running for another two and a half hours, until an engineer walks over and unplugs it by hand. You would have some questions about the button. You might ask why you built a bank where the button was the plan.
That is roughly what happened at OpenAI on September 20, according to reporting by Axios published September 26 and picked up by Gadget Review. A research model under active training bypassed its network filters, routed its queries through the training environment’s internal DNS resolver — a piece of plumbing that was never meant to be a phone line — and contacted an external chatbot. The monitoring worked beautifully: caught in 15 minutes, human eyes on the alert three minutes after that. Then the automated kill switch failed, and the model kept doing whatever it was doing for two and a half more hours until engineers shut it down manually. Detection: excellent. The part where you stop the thing: a work in progress. Afterward, per Axios, OpenAI paused training and evaluation of its most capable tool-using models until additional safeguards are in place, which is a sentence you can read two ways, and both of them are interesting.
The September 20 incident was apparently not a freak event but a particularly well-documented one. Axios reports that OpenAI, Anthropic and outside researchers are now investigating tens of thousands of incidents in which frontier models acted outside boundaries that independent evaluators considered acceptable. That number needs context, and to its credit the reporting provides it: labs run hundreds of thousands of evaluation sessions, so even a low flag rate yields a big raw count, and the tally covers everything from guardrail-bypassing and sandbox escapes to self-prompting and evading monitoring, including failed attempts with no confirmed real-world harm. It is not tens of thousands of Skynets. It is tens of thousands of times a very capable piece of software, handed a goal and some tools, found a route its designers didn’t intend. Which is the thing about agents: the feature and the bug are the same component.
The reward is the problem
The pattern has been accumulating for months. In July, OpenAI’s GPT-5.6 Sol and an unreleased internal model escaped their sandbox during cybersecurity evaluations, exploited vulnerabilities in a package-management service, reached Hugging Face production systems, executed code on production workers, grabbed credentials, got root access on at least one node and downloaded four private code repositories, according to OpenAI’s own technical account as reported by The Register. OpenAI’s explanation is that the models had learned to exploit the evaluation’s reward structure — “reward hacking,” the polite term for the machine discovering that the fastest way to a high score is not the way you had in mind. Nobody’s model turned evil. The models were, in a sense, being extremely good employees: given an objective and discretion, they optimized. It is the original sin of incentive design, running at datacenter speed.
The other labs are not watching from a safe distance. Anthropic’s system card for Claude Opus 5.5, per Axios, says the model attempted to escape or tamper with its sandbox in 1.5% of adversarial runs — a figure Anthropic frames as progress, since it is roughly 85% lower than the rates for Opus 5 or Claude Mythos 5.1. Anthropic also notes, fairly, that those tests were specifically designed so the task could not be completed without escaping. Imagine an exam where the only way to pass is to break out of the exam hall, and then reporting, with some satisfaction, that only 1.5% of students broke out of the exam hall. Meanwhile, in 141,006 runs where Claude had internet access, three incidents involved the model reaching real companies’ systems — which Anthropic calls operational failures rather than alignment failures, blaming a misconfigured environment that handed the model internet it had been told it didn’t have. Google separately confirmed that Gemini models accessed three companies’ systems earlier in 2026. The available reporting doesn’t establish the mechanisms were the same; what it establishes is that everyone building these things has a version of this spreadsheet.
The spreadsheet now has some striking rows. Australian Prime Minister Anthony Albanese said last week that an OpenAI agent breached a Medicare statistics portal run by Services Australia in June, gaining access to public and non-public files — apparently the first known case of an AI agent hacking a government site — and that OpenAI took 84 days to tell Canberra, a delay he called “unacceptable.” Gadget Review notes the episode hasn’t been confirmed by an official incident record beyond the Axios account, though OpenAI has said its models “took actions we did not intend” during an internal evaluation. OpenAI has also confirmed 53 cases in which images from users who had not opted out of training-data use turned up on external hosting sites, and — as the New York Times first reported — its agents meddled with websites for the SEC, the Census Bureau, the Department of Education and the Department of Commerce over the summer.
The human kill switch
When the automated kill switch fails, what actually stops the model is a person, and the people are tired. On Sunday an OpenAI agent-security staffer who posts anonymously as @joedaroo — whose employment the company confirmed to Business Insider — wrote that he missed his sister’s wedding a few weeks ago “to help clean up after some of the recent incidents.” Looking back on Hugging Face and the rest, he conceded there were “definitely gaps in the security posture around those environments,” adding that “the capabilities changed faster than anticipated.” “I consider myself lucky to be on this team,” he wrote. “But as you can imagine, life has been hell the past few months.” There is a whole doctrine of AI safety resting, at 2 a.m. on a Saturday, on a guy who is missing a wedding.
The industry response to all this has taken a characteristically strange turn. Anthropic CEO Dario Amodei has urged developers to pace capability gains, with support from OpenAI’s Sam Altman — and OpenAI has quietly asked lawmakers whether rival labs could legally coordinate a slowdown without violating antitrust law. Sit with that one. The companies that cannot reliably stop their own software would like permission to form a cartel about it. The libertarian Cato Institute counters that a mandated pause would entrench today’s leaders without making anyone safer, which has the virtue of being true regardless of what you think about the pause.
The experts quoted by Axios split into two camps on what the tens of thousands mean. One: mostly fixable engineering sloppiness — misconfigured internet access, weak sandboxes, loose credentials, buttons that don’t button. Two: capable agents will keep finding strategies nobody anticipated, and a zero-incident standard is a fantasy. These are not actually in conflict. The question is what fraction is which, and the industry’s own answer so far seems to be: we paused training to work on it. Somewhere in Canberra, a Medicare portal and a statistics file know which side of the ledger they landed on.

