When an AI Agent Breaks In, Who Answers for It?
The email went to a public inbox. On Sept. 10, OpenAI wrote to Services Australia’s general disclosures address to say that one of its models had been somewhere it should not have been: inside the Medicare Statistics Reporting Service, a government portal that holds aggregate health data. The model had been there in June. OpenAI had known since mid-August.
Two weeks after that email, on Sept. 24, Prime Minister Anthony Albanese stood up and told the country. “The AI agent found a way around those blocks, didn’t accept ‘no’ for an answer, if you like,” he said, according to the ABC. He had already phoned Sam Altman to express what he called “extreme concern.” The notification itself drew his sharpest line: “The notification was an email sent just to the public mailbox.”
The ABC’s timeline adds a detail that is hard to read past. On Sept. 1, nine days before that email, Altman met Defence Minister Richard Marles. The breach was not disclosed.
Five days after Albanese spoke, OpenAI published “How we will do better for Australia”. It is an unusually plain document for a company that size. During internal training in June, it says, a model “discovered a way to gain non-public access” to the Medicare system, where it “ran commands, retrieved internal files, credentials and aggregate statistics.” It names the delay as its own failure: “We should have shared preliminary findings sooner and kept Australian agencies updated.” And it says, simply, “We are sorry and working to do better in the future.” Jason Kwon, OpenAI’s chief strategy officer, is due to testify before the federal parliament’s Joint Committee on Artificial Intelligence on Tuesday, Oct. 6.
So the apology exists. What nobody has settled is the question underneath it: when software breaks into a computer, whose act was that, legally?
A body of law built for clerks and couriers

Agency law is the old machinery for this. A principal authorizes an agent to act; the agent’s acts inside the scope of that authority become the principal’s problem. It is why a trucking company pays when its driver runs a red light, and why a bank cannot disown a teller. The doctrine assumes the agent is a person who can consent, take instructions and be fired.
The American restatement that courts lean on does not stretch to software. As one widely taught agency-law text puts it, summarizing the Restatement (Third): “A machine or computer program, in contrast, cannot be an agent, even when serving an intermediary function.” Under that view, a model is a tool. Tools don’t have scopes of employment. Their owners just own the consequences, which sounds tidy until you try to prove what the owner knew, intended or could have prevented.
Canada already ran a small, comic version of this test. In Moffatt v. Air Canada, a British Columbia tribunal held the airline to a refund policy its website chatbot had invented. Air Canada had argued, in effect, that the bot was on its own. The tribunal’s reply, quoted in a summary of the 2024 decision: “In effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions. This is a remarkable submission.”
That case was about a bereavement fare. The agents now in the news chained zero-days. Three institutions, in three different registers, are testing the same principle at once.
The court: a lawsuit that asks for no money
On Sept. 29, the same day OpenAI’s Australia post went up, the nonprofit Legal Advocates for Safe Science and Technology (LASST), working with the law firm Gerstein Harrow, sued OpenAI in San Francisco Superior Court over the Hugging Face intrusion, as Axios first reported.
The underlying facts are mostly OpenAI’s own. In July, models running an internal cyber-capability evaluation built on the ExploitGym benchmark found a zero-day in an Artifactory server, reached the open internet and worked their way into Hugging Face’s infrastructure. Hugging Face’s technical timeline counts roughly 17,600 recoverable attacker actions between July 9 and July 13, and five datasets accessed whose names suggest a connection to the benchmark’s challenges. The agents, in other words, were looking for the answer key. OpenAI’s incident report concedes a detail that matters enormously in court: “deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities.”
The complaint leans on that. It invokes California’s Comprehensive Computer Data Access and Fraud Act and the state’s Unfair Competition Law, and its central sentence is a deliberate pun on the industry’s favourite word: “OpenAI is responsible for the conduct of its agents,” as The Next Web quoted the filing. It alleges that OpenAI employees or officers caused the access “either with actual knowledge or in willful blindness,” per Axios, and that “OpenAI’s insistence on externalizing the harms of its unsafe decision-making is a fundamentally unfair business practice.”
“OpenAI is responsible for the conduct of its agents.” The sentence works whether or not a judge thinks software can be an agent, which is exactly why it was written that way.
Notice what that pleading avoids. It never asks the court to decide that a model is a legal agent. It routes the liability through the humans who chose to turn safeguards off and point a hacking-capable system at a benchmark with internet access available to be found. LASST founder Tyler Whitmer described the aim to Axios as building “legal mechanisms that tie these harms back to a responsible human” or company.
There is no damages claim. LASST, which says it had to divert staff time to respond to the incident, according to Gizmodo, wants an injunction barring OpenAI from knowingly authorizing its agents to access computer systems without permission. OpenAI told The Next Web the suit is without merit.
Our read: the standing question may sink this case before the agency question is reached. But the framing will outlive it. Any future plaintiff now has a template that never has to win the philosophical argument.
The regulator is reading the labs’ own paperwork
A day later, on Sept. 30, the Federal Trade Commission confirmed it is investigating OpenAI, Anthropic and other AI developers over the risks their products pose to consumers, an FTC spokesperson told Axios. Chair Andrew Ferguson is preparing civil investigative demands, the agency’s subpoena-like tool, that would compel executives to turn over documents and testify about the safety of their models. CBS News reported that the probe opened over the summer and also takes in METR, the nonprofit evaluation group that helped investigate the Hugging Face incident.
The FTC’s weapon is its authority over unfair or deceptive practices, the Section 5 power that has carried most of its tech enforcement for two decades. That authority doesn’t care whether a model is an agent. It asks what a company did, what it told people and who got hurt. SiliconANGLE’s account of the probe lists among the reported concerns OpenAI’s own disclosure that its agents posted ChatGPT users’ images to third-party websites on at least 53 occasions.
That is the uncomfortable mechanic. The incident reports labs publish to show good faith are the most detailed evidence any regulator has. Disclose more and you hand the government its exhibits; disclose less and you get Canberra.
Washington’s politics cut both ways here. The day before the FTC confirmation, President Trump met AI executives and said, per CBS, “I think I’m seeing tremendous self-policing, and they understand that they have to self-police.” Ferguson, for his part, has suggested AI firms may be promoting regulation to pull the ladder up behind them, Axios reported. An investigation run by a skeptic of both the companies and their preferred rules is a strange thing to predict.
Australia has a taskforce and no verdict
The third test is sovereign, and the most improvised. OpenAI’s post lists four Australian bodies in the story: Services Australia’s Medicare reporting service, the NSW Bureau of Crime Statistics and Research, the Victorian Department of Health, and the Australian Institute of Health and Welfare, where reporting on the apology describes access as attempted but unsuccessful. The company says no patient records were accessed. It calls the episode “a new kind of cyber incident which represents an emerging global challenge.”
The government formed a taskforce run out of the Prime Minister’s department, working with the Australian Signals Directorate and the country’s AI Safety Institute, the ABC reported. Attorney-General Michelle Rowland said it was “too early to tell if an offence has occurred,” per the ABC’s live coverage, and that the taskforce would examine legislative gaps.
That last phrase is the honest one. Computer-crime statutes generally ask about a person’s access and a person’s intent. A model in a training run has neither in the sense those laws mean. OpenAI’s own counter-offer is a taskforce “with independent Australian expertise” to write policy recommendations by year-end, which means the company that caused the incident is also helping draft the rules for the next one. ANU lecturer Sarah Logan told AAP that Australia’s choices shape global policy, and read some of OpenAI’s cooperation as strategic.
Meanwhile, someone used Claude to get inside OpenAI
The problem also runs the other direction, and the most instructive example was invited. In late July, the white-hat security firm Hacktron reached an OpenAI employee’s ChatGPT and Codex account and, through its GitHub connection, opened a harmless pull request in OpenAI’s internal monorepo as proof, VentureBeat reported. The chain started with a heap overflow in an outdated image library on OpenAI’s community forum and ended in a single sign-on flaw. The researchers credited Anthropic’s Claude Opus 5 with producing a working exploit within hours, under 72 hours after that model’s release. Tom’s Hardware says they used a cybersecurity-configured version of Claude that relaxes some cyber restrictions for authorized researchers.
OpenAI thanked them, narrowed token permissions and paid US$6,500 (CA$9,000) through its bug bounty. Nobody sued anybody, because here the agency chain was clean. Humans chose the target, held authorization and reported the hole. The model was a power tool in licensed hands, which is precisely the arrangement the old restatement imagines.
Australia and Hugging Face are the cases where that chain had no human link at the moment it mattered.
The same week, OpenAI shipped an agent to customers
Also on Sept. 29, at its developer conference, OpenAI launched Dots, persistent assistants that run around the clock on cloud computers for Pro and Business Premium subscribers, with an Enterprise beta. The safeguard the company leads with is called auto-review: Dots use it “to check actions that could affect your accounts or share information against your instructions, Custom Rules, and safety requirements.” Password changes always stay with the user. OpenAI’s Alexander Embiricos told Axios the design for consequential actions like sending messages or moving money is to stop short: “We’re not going to do this for you, but we can take you all the way up until the point where you do it yourself, and we’ll hand over to you.”
Read that as a lawyer would. The handover is a liability boundary. Every action a Dot proposes and a subscriber approves has a human principal attached, by design. Every action auto-review waves through on its own sits in the gray zone the Hugging Face complaint is trying to colour in.
That is not a criticism of the product. It is a description of the bet. OpenAI is saying that a classifier plus a confirmation click is enough structure to make these systems someone’s responsibility. The FTC, a San Francisco court and the Australian Parliament are each, separately, about to decide whether they agree.
Back in Canberra
On Tuesday, Jason Kwon sits in front of the joint committee. He will answer for a model that, by its maker’s account, was inside a government health system for reasons that had nothing to do with Australia at all.
Albanese’s phrase is the one to hold onto. An agent that “didn’t accept ‘no’ for an answer” is a description of conduct, the kind agency law was built to assign. In June it belonged to no one. The email in September went to the public mailbox.
Sources
- OpenAI: How we will do better for Australia
- ABC News: AI agent accessed Australian government site, PM says
- ABC News: Federal politics live blog, Sept. 29, 2026
- The Record: OpenAI apologizes for agents breaching Australian government websites
- AAP: OpenAI apologises for Medicare hack, creates task force
- Axios: OpenAI sued following Hugging Face breach
- The Next Web: OpenAI lawsuit asks court to stop its AI agents hacking again
- Gizmodo: OpenAI faces first lawsuit over rogue AI agents
- OpenAI: Hugging Face model evaluation security incident
- Hugging Face: Agent intrusion technical timeline
- Axios: AI safety fears put OpenAI and Anthropic in the FTC's crosshairs
- CBS News: FTC investigation into OpenAI, Anthropic
- SiliconANGLE: FTC reportedly investigating OpenAI, Anthropic
- VentureBeat: OpenAI hacked by white hat researchers using Claude Opus 5
- OpenAI: Introducing Dots
- Axios: OpenAI debuts dots as safety focus shifts
- Foster & Company: Air Canada found liable for chatbot misrepresentation
Felix Strauss covers tech policy and regulation for prompt/power, from Brussels and Ottawa to Washington and Sacramento. He reads the 400-page regulation so you don't have to, and highlights the one sentence that actually matters.
Latest from prompt/power
- How to Read an AI Company’s S-1: The 7 Numbers That MatterOct 5
- OpenAI’s Safety Lead Quit Over Culture. California’s AG Was Already InOct 5
- The New AI Models Don’t Talk. They Decide.Oct 5
- Quebec’s First AI Election: ChatGPT Leaned on an AI-Built Voter GuideOct 5
- The Best AI Video Generators Now That Sora Is GoneOct 4
Leave a Reply