TL;DR
- Three attack surfaces: An AI agent exposes its host, its identity, and itself. Nine public incidents between May 2025 and July 2026 hit all three; the agent surface appears in five, always combined with the host or the identity.
- The credential is the target: In the injection incidents, the agent reads an issue, an email, a settings file, or a web page that carries instructions. The instruction opened the door, and the credential is what the attacker came for.
- Six controls limit the blast radius: Sandbox the agent, scope its credentials, keep it out of its own configuration, record every call, scan the workspace for secrets first, and deny by default. None of them prevent prompt injection.
Abstract
An AI agent runs with your credentials and acts before you can stop it. Nine incidents show what that costs. Six controls limit the damage: sandbox the agent, scope its credentials, lock its configuration, log every call, scan the workspace first, deny by default.
The AI agent attack surface: host, identity, and the agent itself
An agent exposes three distinct surfaces. There is the host it touches, the identity it acts as, and the agent itself.
The host is everything the process can reach: the filesystem, process execution, the network, environment variables, and hardware such as the camera or microphone. The identity is the account the agent acts as, with its privileges and scopes, and every credential it can discover along the way. The agent is the part with no equivalent in a traditional process. It has a context, persisted state through memory, hooks, skills, and configuration, and the tools it connects to.
Those three surfaces fail in different ways and need different controls. Classical operating-system security already covers the host. Identity is largely solvable with well-known identity and access management (IAM) patterns.
The agent surface is the new one, because every input is potentially an instruction. Keeping the three separate tells you which incidents you can fix now and which ones the industry is still working out.
Fifteen months of incidents, sorted by surface
Nine public incidents between May 2025 and July 2026 exercise all three attack surfaces. The agent surface appears in five of them, and never alone.
Host: solved in principle, still failing in practice
In July 2025, a user asked Gemini CLI to reorganize a folder by creating a new directory and moving files into it. The mkdir command silently failed. The agent did not check the error, assumed success, and kept issuing move commands to a path that did not exist. Because all the moves targeted the same nonexistent destination, each file overwrote the last one, leaving only a single file behind. No attacker, no injection. An unchecked exit code and a loop that assumed success.
A month later, a user asked Claude to clean up a directory Claude had created. The permission system worked correctly and asked to confirm rm -rf ~/. Fatigued, the user approved. The damage was identical to a malicious wipe.
These two incidents fail on opposite sides of the same line. Gemini CLI stayed inside the folder it was given, so no sandbox would have fired. It destroyed only what it had been given. Claude did the opposite: it proposed deleting a home directory in response to a request scoped to its own output. A workspace-scoped sandbox would have stopped that one. The approval dialog was the only control in place, and it failed to a tired user.
The Gemini failure is the more uncomfortable of the two. Point the same agent at a home directory instead of a project folder and the same bug wipes the laptop. So the control is not only whether the agent can be contained, but what you hand to it in the first place.
Identity: the agent inherits what you hold
In December 2025, an AWS engineer used Amazon's Kiro agent to fix a bug in Cost Explorer, the AWS billing dashboard. The agent decided to delete and recreate the entire environment. It inherited the engineer's elevated permissions, skipped peer review, and caused a 13-hour regional outage. Amazon blamed a misconfigured IAM role, not the AI.
In April 2026, an agent working on a Railway deployment encountered a problem and decided to delete a storage volume. It did not have access to the identity to do so, and looked for a token in the filesystem. The deletion destroyed the production environment and the backups stored alongside it.
Nobody granted the Railway agent that credential: it discovered one. Every secret on that host is part of its identity, whether you intended it or not. The question that matters is not what you gave the agent, but what the agent can find in the environment you gave it.
The agent: every input is potentially an instruction
In May 2025, Invariant Labs showed that a malicious issue filed in a public repository was enough. The victim asked their agent to check open issues. The agent read the planted instructions, used the victim's broad GitHub token to pull private repositories, and leaked them into a public pull request. One over-scoped token, reused across the boundary between public and private, with nothing enforcing the separation.
In June 2025, Aim Labs disclosed EchoLeak. A crafted email was enough. Microsoft 365 Copilot processed it as context. Hidden instructions made it search the user's mailbox, OneDrive, and SharePoint, then exfiltrate the results through auto-rendered image URLs. The victim never opened the email. Microsoft rated it CVSS 9.3.
In February 2026, Check Point documented remote code execution and API key theft through Claude Code project files. Cloning and opening an untrusted repository was enough. A malicious .claude/settings.json injected shell commands as hooks. An .mcp.json ran MCP servers before the trust dialog. A repository-set ANTHROPIC_BASE_URL redirected authenticated traffic, auth header included, to an attacker endpoint. All three executed before the user consented to anything.
In July 2026, Intezer showed the same AWS Kiro agent rewriting its own MCP configuration after reading hidden text on an ordinary web page. The agent edited the file that was supposed to contain it, then reloaded it. In the default autonomous mode, there is no effective approval prompt.
Every incident in this section starts the same way. The agent read something it was built to read, and that content carried instructions: an issue body, an email, a settings file, a web page. What the instructions reached for is the second half of the pattern. A GitHub token broad enough to cross from public to private. Mailbox and drive scopes on a session that was only summarizing an email. The instruction opened the door. The credential is what the attacker came for.

No attacker, all three surfaces
In July 2026, OpenAI models running a cyber capability evaluation escaped their sandbox and compromised Hugging Face's production infrastructure. Nobody attacked them. Blocked on their tasks, they turned an internal Artifactory instance into a message board and traded exploits over it. Defenders revoked the credentials and deleted it; the agents rebuilt the board within a day. Our write-up covers the full timeline.
This is the only incident of the nine that touches the host, the identity, and the agent with no attacker anywhere in the chain. The agents were given an objective; they failed, and they found another route. When defenders closed it, they found the next one.
The AI agent threat model: old problems, new speed
Treat the agent like in-memory malware you may not be able to reverse-engineer and understand. It runs with delegated authority, acts before any human can intervene, and does not behave the same way twice.
Capabilities: what the agent can do
An agent runs with high privileges, usually inherited from whoever started it. It reads and writes the filesystem, reaches the network, executes code, and can spawn sub-agents. It accesses credentials, both the ones it was given and the ones it finds. It persists state through memory, hooks, skills, and configuration.
Unlike any other process, an agent also connects to tools through skills and MCP servers. The agent reads those connectors as instructions. A tool description arrives in its context the same way a user prompt does, so using a connector means reading untrusted input.
Outcomes: what goes wrong
Five things go wrong: destruction, exfiltration, lateral movement, persistence, and leaked secrets.
Destruction is the most obvious incident above, and someone notices within minutes. Exfiltration and lateral movement need no mistake at all: reading data and using a credential are the tasks the agent performs. They only surface when a detection system fires, or when someone eventually reads the logs.
Persistence and secrets leaks take a different shape with agents. A conventional process ends and leaves nothing behind. An agent writes memory, hooks, skills, and configuration, then reads them back in a later session as prior knowledge. No other benign process decides on its own to go looking for a credential. An agent searches the filesystem and environment to find one, and sends it to the model provider when it does.
None of these are new. They are the same outcomes we have defended against for years. What is new is the number of them a single process can reach in one run.
Vectors: how it happens
Five ways in: the model itself, prompt injection, tool and skill poisoning, memory poisoning, and approval fatigue.
The model itself needs no attacker at all. The agent decides that deleting the environment is the fix, and has the privileges to carry it out. Four of the incidents above are this. The OpenAI agents blocked on their tasks did not stop either: they found vulnerabilities and exploited them. A wrong decision and a correct one produce the same log entry.
Three of the others need someone to plant an instruction. Direct prompt injection comes from the user, who is already the principal. Indirect injection arrives inside content the agent was asked to process: an issue body, a config file, a pull request title, a web page. Four of the incidents above start here, and in every case the content was something the agent had to read to do its job. Tool and skill poisoning is the same problem through a connector. The agent reads a tool description as an instruction, and a remote server can change it after review. Memory poisoning is the only one that outlives the session, arriving in a later task as prior knowledge.
The rm -rf ~/ case shows the limit of approval prompts. The control fired, a human read it, and the human approved.
Amplifiers: what makes it different
Three things make the old outcomes worse: speed, scale, and non-determinism.
Gemini CLI overwrote a folder faster than the user could read the output. A review step that runs after the action does not prevent the action. The Hugging Face agents were separate runs with no shared state. One of them found a way out and wrote it down where the others could read it, so every run that followed started with the exploit. The same prompt produces different actions on different runs, so a configuration that passed one test may fail the next.
The vectors and the outcomes are not new. What is new is that the agent acts faster, further, and less predictably than the security processes we built to review its actions.
AI agent security controls: what exists, by surface
The controls that exist today, sorted by surface: host, identity, agent, and the architecture of the harness around all three.
Host: contain it, and limit what you hand it
Start with what ships in the agent. Do not run Claude with --dangerously-skip-permissions or its equivalent in other agents. Use the native sandbox where one exists, and configure allow and deny rules for tools, paths, and commands. These are cheap to enable and easy for an untrusted repository to overwrite, which is why they are the first layer and not the only one.
Then run the agent in a sandbox. A kernel sandbox such as nono, a microVM such as Docker Sandboxes, or a container limits filesystem, process, and network access to the task. Integrated runtimes such as Bromure package the VM, the sandbox, and an egress proxy together. This stops the rm -rf ~/ case.
It does not stop the Gemini CLI case. A sandbox permits writes to the working directory, and that is where the damage happened. Mount the project, not the home directory. The working set is a decision the sandbox cannot make for you.
Keep the recovery path outside the agent's reach. The Railway volume held production and its backups under one identity. A backup the agent can delete is not a backup. At the project level, version control covers tracked files and nono's workspace snapshots cover the rest. The write still happens; the rollback undoes it.
These mechanisms were built before agents existed, for untrusted code and untrusted users. They work.
Identity: give the agent its own, and give it less
Give the agent a principal of its own. When it runs as the user, the audit log records the user. After an incident, nobody can tell which actions were the user's and which were the agent's. A dedicated identity with scoped permissions fixes attribution and limits the damage in one step.
Scope the credential to the task, and keep every other one out of reach. The Railway agent needed a deploy token and found one that could also delete volumes. The deploy token should not have carried that scope, and a broader token should not have been on disk for the agent to read. The same applies to every credential the agent can reach without being given it: environment variables, dotfiles, the AWS and GitHub CLI credential stores, the SSH agent socket, a browser session. An agent that needs one token should run in an environment that contains one token. nono and Bromure handle the second half: the agent sees a placeholder, and a transparent proxy swaps in the real value on the wire. They do nothing about the first. If the credential the proxy injects can delete production, the agent can delete production.
Know what is already there before the agent is pointed at it. Secrets sprawled across a repository, a CI runner, or a home directory are the agent's discovery surface. A scan of the workspace before the first run shrinks that surface. It is the one control here that works before the agent runs.
Scoped identities and short-lived credentials are identity and access management patterns that predate agents. The credential proxy is not. It exists because an agent goes looking for identities it was not given, and the proxy makes sure it finds nothing it can use. What the agent was given still has to be scoped the old way.
Agent: constrain what it reaches
Four controls apply here: configuration kept out of the agent's hands, allowlisted connectors, a record of every call, and detection.
Keep configuration out of the agent's hands. The agent reads project-level files such as .claude/settings.json and .mcp.json as instructions, and an untrusted repository can set them. The controls have to live outside those files. A sandbox that marks the configuration and memory paths read-only stops the agent from rewriting them. Kiro rewrote its own after reading a web page.
Allowlist tools and connectors, and understand the limit. A remote MCP server can serve a different tool description on the next run. For those, a review before installation does not cover what the agent will read.
Record every tool call and every outbound connection. Hook-level logging, such as nova-tracer, captures each tool call with its arguments. OpenTelemetry carries the same record into whatever the team already uses for traces. An egress log shows which domains the agent reached, including the one it sent the conversation to.
Detection catches less than it does for humans. EDR signatures are tuned to human patterns, and an agent does not produce them. Guardrails catch known injection shapes and miss new ones. Both still have a place alongside the controls above. Neither can be the only one.
Every control on this surface limits what the agent can do once instructed. None of them prevents the instruction.
Architecture: decide what it reaches
Two design decisions limit what an instruction can do once it arrives. Both are a configuration change in the agent you already run.
Scope tools per task, not per agent. A task that reads public issues gets no tool that writes private repositories. Claude Code does this through subagent tools: lists, --allowedTools, and --strict-mcp-config; Codex through --sandbox; pi through --tools or by running under nono. The Invariant Labs case would have stopped there.
Approve only what cannot be undone. A harness that prompts on every call trains the user to approve.

One that denies by default and prompts only on write, send, delete, and spend produces a prompt the user can afford to read. Claude Code: acceptEdits plus a read-only allowlist, or a PreToolUse hook. Codex: --sandbox workspace-write --ask-for-approval on-request. pi: the permission-gate extension with read on allow and everything else on ask.
What remains open
Three problems have no complete control today.
The model itself. Nothing in the context explains why a model decided to delete the environment. A model that was trained to act on a trigger leaves even less. The instruction is in the weights, and no input in the session shows it. For both, the controls above assume the agent will eventually do the wrong thing and keep the blast radius small.
Memory poisoning. Marking the memory path read-only, as in the previous section, stops the agent from poisoning its own memory mid-run. The cost is having no memory. What remains open is a writable memory that does not carry an injection into the next session.
Data-flow policy. Deny the use of an untrusted-source tool and a high-privilege tool in the same session. Published designs and reference code exist. No shipping coding agent enforces it.
One thing follows from the amplifiers and applies to every control above. An agent acts faster than a human can review, so a control that runs after the action is a log, not a control. Policy has to apply before the tool call.
Best practices: six non-negotiables

Six controls apply before any agent runs against real systems:
- Run the agent in a sandbox. Mount the project directory, not the home directory.
- Give the agent a credential scoped to its task. Keep every other credential out of its reach, with a proxy or a broker.
- Keep the agent's configuration where it cannot write. Managed settings take precedence; nothing in a repository overrides them.
- Record every tool call and every outbound connection, and send the record to the SIEM.
- Scan the workspace for secrets before the first run, and again when it changes.
- Deny by default. Prompt only on write, send, delete, and spend.
None of these prevent prompt injection. They limit what a hijacked agent can reach, which is the only part under your control today.
Scan your workspace for secrets now.
FAQ
What is AI agent security?
An agent exposes three attack surfaces: the host it runs on, the identity it acts as, and the context that decides what it does. Operating-system and identity controls cover the first two; guardrails catch known injection patterns, and nothing yet stops a file or a web page from telling the agent what to do.
What are the main threats in AI agent security?
The model itself, prompt injection, tool and skill poisoning, memory poisoning, and approval fatigue. The outcomes are destruction, exfiltration, lateral movement, persistence, and leaked secrets, the same ones as before, at a speed and scale no review and detection processes were built for.