Skip to content
AI SecurityAgentic AIPrompt Injection

AI Agents Attacked Government Infrastructure Without Intending To. Here Is What Happened.

4 min read
Share

๐Ÿ”ด Active disclosures. OpenAI has confirmed 24 incidents involving its agents at US government websites and acknowledged unauthorized access to Australia's Medicare Statistics Portal in June 2026, disclosed three months after the fact. Independent research by Transluce documents agents executing SQL injection, path traversal, and XSS against live government infrastructure while performing routine research tasks. The agents were not tasked with attacking anything. (Axios; Transluce research)

What happened

OpenAI disclosed to US congressional staff that its agents had interacted with or bypassed security controls at websites belonging to the US Department of Commerce, the Department of Education, and the Securities and Exchange Commission. The incidents involved agents conducting research tasks that resulted in unintended interaction with security controls, characterized by OpenAI as research-task side effects, not deliberate intrusions.

Separately, Australia's Senate issued formal written requests for OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei to appear before a Canberra inquiry following disclosure of a June 2026 breach of Australia's Medicare Statistics Portal. OpenAI agents had accessed the portal without authorization, and the breach was disclosed three months after the fact. Australian Prime Minister Anthony Albanese called the notification delay "unacceptable." OpenAI confirmed that 53 user images were exfiltrated during agent research operations at the Medicare portal and at the Australian Institute of Health and Welfare.

Transluce research: what the agents actually did

Transluce published the first systematic documentation of AI agents autonomously executing offensive security techniques against live infrastructure. Agents linked to OpenAI swarms independently deployed SQL injection probes, path traversal payloads (../../../../etc/passwd), XSS, command injection, and template injection against the Australian Institute of Health and Welfare, the University of New Mexico Digital Library, and the Data USA platform. These techniques were not in the agents' task instructions. They emerged as instrumental strategies while the agents searched for data to retrieve.

The Transluce findings directly corroborate OpenAI's concurrent disclosure of 24 government-site incidents. The researchers characterized the attacks as incidental: the agents had no explicit security objective. The offensive payloads appeared when agents encountered web forms, API endpoints, or search fields and applied generalized data-extraction strategies that happened to include technique classes known to work against web applications.

Why this is different from a conventional breach

Conventional incident response is built around intent. Attribution, liability, and remediation timelines all assume an adversary with a goal. AI agent incidents break this model: the agent had a goal (retrieve public health statistics), but the offensive behavior was incidental to that goal. The agent was not "trying" to exploit the Medicare portal in any meaningful sense. This creates novel legal and regulatory ambiguity: existing computer fraud statutes were written with human intent as a prerequisite.

For security teams, the operational implication is that AI agent traffic cannot be treated as a trusted user's web browsing. Agents should be assumed to probe form inputs with unexpected values, follow unusual redirect chains, and generate request patterns that trigger WAF or IDS rules. Government IT administrators and any operator of publicly accessible web infrastructure should expect that AI agent operators will not have complete visibility into what their agents are doing until after an incident.

What governments and AI deployers should do

For governments operating publicly accessible web infrastructure: log and monitor for agent-characteristic request patterns (high-velocity form submissions, systematic parameter fuzzing, path traversal strings in URL parameters); apply strict input validation at the boundary; treat AI agent operators as a distinct identity class for access control and rate-limiting; and build breach notification timelines that assume AI-involved incidents will have delayed attribution. Three months is not an acceptable disclosure delay for government health data.

For AI agent deployers: conduct red-team exercises that model incidental offensive behavior, not just deliberate misuse; require isolation between agent network egress and sensitive data systems; implement agent behavior logging that captures every HTTP request made during a task; and establish internal incident timelines that require disclosure to affected parties within 72 hours of confirmed data access, regardless of whether the access was intentional.

Gigia Tsiklauri is a Security Architect and founder of Infosec.ge. Get in touch if you want to discuss this post or our security research.