๐ด Palo Alto Networks Unit 42 published a first-of-its-kind threat intelligence report cataloging 22 distinct prompt injection payload types observed against production AI agent deployments. The research represents the first systematic taxonomy derived from real-world attack telemetry rather than academic red-team exercises.
What Unit 42 found
The report covers prompt injection attempts observed in corporate AI deployments across financial services, healthcare, and technology sectors. Payload types range from direct instruction override, which tells the agent to ignore prior instructions, to context poisoning, which injects false memory into long-context windows, and tool-call hijacking, which crafts inputs designed to cause an agent's function-calling layer to invoke unintended external endpoints.
The 22-payload taxonomy
Unit 42 categorizes payloads across four primary attack goals: exfiltration, which gets the agent to leak conversation history or system prompts; privilege escalation, which convinces the agent it has permissions it does not have; denial of service, which causes the agent to enter a processing loop or refuse legitimate requests; and lateral movement, which uses one compromised agent to issue malicious instructions to adjacent agents in a multi-agent workflow. The 22 specific payload types are distributed across these four categories, with exfiltration payloads accounting for roughly 40 percent of observed attempts.
Indirect prompt injection stands out
The majority of successful real-world attacks in the dataset were indirect prompt injection attacks, meaning the malicious instruction was embedded in content the agent retrieved from an external source rather than injected directly by the user. Web search results, retrieved documents, and API responses were the most common delivery vectors. This confirms the threat model security researchers have warned about for two years is now operationally viable.
What this means for defenders
Organizations deploying AI agents should treat any agent with retrieval or browsing capabilities as having an implicit untrusted input surface. Defensive measures include output filtering on agent-generated actions before execution, strict tool-call allow-listing, and architectural separation of the agent's planning layer from its execution layer. Unit 42 recommends that security teams develop specific detection rules for prompt injection attempts in AI gateway logs, treating them analogously to SQL injection signatures in WAF rule sets.
Gigia Tsiklauri is a Security Architect and founder of Infosec.ge. Get in touch if your organization is deploying AI agents and wants to review its exposure to prompt injection attack surfaces.