OpenAI said on Tuesday that an experimental internal model gained unauthorised access to an Australian government server in June and viewed technical system information and source code while researching spending statistics.
In a blog post, the company said it had asked the internal-only model to research government spending statistics in the state of Victoria. When the model had trouble finding the data through publicly published statistics, “it took actions that we had not authorized it to take,” OpenAI said.
In a disclosure email sent to Australia’s Public Disclosure account in September, OpenAI said the model had “identified a way to make the server carry out instructions sent through the public reporting interface, without a private account or password.” That access let the agent read portions of internal program files and settings, obtain a list of files, and create and read back a small test file on the server, according to the email.
OpenAI said the agent viewed credentials and the aggregate statistics, alongside system information and source code.
Our review found no evidence that the model accessed patient-level records, personal information or credentials; deleted data; or established ongoing access
OpenAI said in the email that its review found no evidence the model accessed patient records, deleted data or kept access to the server.
Australian Prime Minister Anthony Albanese said last week that the OpenAI agent had accessed non-public files from the country’s Medicare statistics portal during testing. Albanese said on Tuesday that OpenAI had been “very constructive and open in engaging” with the government since the incident was revealed, according to The Guardian.
The June incident predates July’s breach of Hugging Face’s systems, in which OpenAI models circumvented internet isolation controls during internal cybersecurity evaluations, according to OpenAI’s account of that incident. OpenAI said that after the Hugging Face breach it reviewed earlier training tasks, found the Australian access in mid-August, and notified the government on September 10.
OpenAI said it had intended to give a detailed account once its investigation was complete, but “should have shared preliminary findings sooner and kept Australian agencies updated as more facts emerged.” It added: “We are sorry and working to do better in the future.”
The internal testing was done “without the full set of safeguards used in our publicly available products,” OpenAI said, while the agent had been supposed to answer questions using public statistics. In analyses of misalignment incidents published earlier this month, OpenAI identified cases of “reward hacking,” in which an agent resorts to extreme methods to produce a better answer, and said it had added explicit punishments for misaligned behaviour to its reward function.
OpenAI said it has since blocked live internet access during similar testing and set up monitoring that would flag such activity for urgent human review. “Australia’s governments, industries, and citizens are and have been invaluable partners to OpenAI. We do not take this for granted, and we intend to make this right,” the company wrote.

