Prompt injection: limit what a fooled agent can do
No filter today reliably stops an AI agent being fooled by instructions hidden in what it reads. What you can control is how much a fooled agent is able to do.
By the Webair engineering team8 min read

Prompt injection is when instructions hidden in something an AI agent reads, such as a web page, an email or a support ticket, change what it does. There is no reliable way to prevent it today. The practical defence is to decide in advance what a fooled agent can do: give it only the access its task needs, have a person approve anything that can’t be undone, and test it against attackers who know your defences.
OWASP’s 2026 Top 10 for LLM Applications, published in August, keeps prompt injection in first place. The biggest move is excessive agency, where a system has more functions, permissions or autonomy than its task needs: it rises from sixth to third.1 For the first time, the ranking weighs incident data as well as practitioners’ judgement. A vote by practitioners carries three quarters of the weight, and 6,639 incidents drawn from public vulnerability databases and an AI-harm database carry the rest.
OWASP links the two entries itself. Prompt injection is how an attacker’s instructions get in; excessive agency is what gives them consequences.
What prompt injection is
A language model reads its instructions and the content it’s working on as one stream of text. It has no built-in way to tell your system prompt apart from a line in a customer’s email that says “ignore your previous instructions”.
Direct injection comes from the person typing, for example a user trying to talk the model out of its rules. Indirect injection arrives in content the system reads: a web page it summarises, a document in its knowledge base, the output of a tool. NIST’s taxonomy of attacks on AI systems draws the same line. It describes indirect injection as carried out by controlling a resource the system reads, rather than through what the user types.2
Indirect injection is the one to design for. The attacker doesn’t need to break into your systems. They only need to put text where your agent will read it, and the agent does the rest with its own permissions.
That’s why agents raise the stakes. A fooled chatbot can say the wrong thing. A fooled agent can do the wrong thing, because it can plan a task and use your tools and data to carry it out, as we describe in what an AI agent can do on your website.
Why filters alone won’t stop it
The obvious response is a filter: screen what goes into and comes out of the model, and block anything that looks like an attack. Filters help. But OWASP notes that they can be evaded by rephrasing or encoding an instruction, and concludes that no reliable way to prevent prompt injection exists today.
Research on defences points the same way. In a 2025 preprint, 14 researchers attacked 12 recent defences against prompt injection and jailbreaks, adapting each attack to the defence it faced. They got past most of the defences more than 90% of the time, although most had originally reported attack success close to zero.3
The gap between those two results is the lesson. A defence measured against a fixed set of known attacks can look solid and still fail against someone who studies it.
OWASP’s incident data adds a twist. Ranked on incidents alone, prompt injection would drop out of the top 10. OWASP doesn’t read that as low risk. Teams fight injection hard, so fewer clean exploits reach public databases, and the attack surface is still everywhere a model reads untrusted input.
Adding a filter is the easy part, and worth doing. The harder part, and the part that holds up, is deciding what the agent is allowed to do when the filter misses. OWASP’s advice is to build the defence into the system’s design rather than rely on catching every attack.
Limit what the agent can reach
Start with the question behind OWASP’s excessive agency entry: does the agent have more functions, permissions or autonomy than its task needs?
| Too much | What it looks like | How to narrow it |
|---|---|---|
| Functionality | A tool meant for reading documents can also change and delete them | Use tools that only do what the task needs, and remove any left over from testing |
| Permissions | A tool that only reads a database connects with an identity that can also update and delete | Give each tool its own narrow access, and act with the user’s own permissions where you can |
| Autonomy | The agent deletes a user’s documents without asking | Require confirmation for actions that are high-impact or can’t be undone |
OWASP singles out least privilege as the control that carries the weight for agents. In practice that means two rules. Keep credentials, and anything that changes data, in your application code rather than in the model’s hands. And check every action against a policy in code before it runs, instead of asking the model whether it’s allowed.
In outline:
// The model proposes an action. Your code decides.
const action = await agent.propose(task);
if (!policy.allows(user, action)) {
// checked in code, not by the model
reject(action);
} else if (action.irreversible || action.external) {
// a person sees the exact action
await requestApproval(user, action);
} else {
// with the user’s own permissions
await run(action, { as: user });
}
Take most care with agents that combine three things: they read untrusted content, they can see sensitive data, and they can change something or send something out. OWASP treats that combination as high risk, and recommends that a person approves each action taken by an agent with all three. Removing any one of the three takes away the conditions for the worst outcomes, and it’s often the simplest fix: an agent that summarises incoming email doesn’t need to send any.
The same applies to every system you connect an agent to. Each connection is a set of actions the agent can take on someone’s behalf, and each needs its own limits.
For agents that act on their own, with memory and tools, OWASP keeps a separate list: the Top 10 for Agentic Applications, published in December 2025. Its first entry is agent goal hijack, where injected instructions redirect what the agent is trying to do.4
Put a person before high-impact actions
Some actions shouldn’t run on an agent’s word alone: moving money, deleting data, changing permissions, sending anything outside the organisation. OWASP recommends explicit human confirmation before any action that’s privileged, can’t be undone or is visible outside the system.
Not every action needs a person, and OWASP’s own example draws the line well. A support agent can issue a refund as store credit by itself, because that can be reversed. A payout to an external account goes to a person.
Approval only works if the person can see what they’re approving. Show the exact action: the recipient, the amount, the text that will be sent. A summary written by the agent is itself model output, and OWASP warns that hidden characters can make the action on screen differ from the one that runs.
In practice
Watch how approvals are going, not just that they happen. OWASP warns that approval fatigue wears down judgement at volume. If people approve almost everything without reading it, move the routine actions into rules in code and keep people for the decisions that matter.
Test for it before launch
Test the agent the way an attacker would approach it, not only with a fixed list of known attacks. OWASP’s advice is to test against attackers who have read your defences, and not to trust results from static tests alone. It suggests starting from public benchmarks such as AgentDojo, then red-teaming with the testers told exactly how your defences work.
Cover at least:
- instructions hidden in every kind of content the agent reads: web pages, documents, emails, tickets and tool outputs
- instructions in forms a reviewer won’t see, such as invisible characters, encoded text, other languages and images
- an attempt at every high-impact action, to confirm the limits and approvals hold even when the model is fooled
- the same tests again whenever the model, the prompt or a tool changes
Measure two things separately: how often an attack gets the model to follow it, and how often a fooled model manages to do harm. Your design controls the second number.
These cases belong in the same test set you use to decide whether the feature is ready to ship.
If you sell the product in the EU
If what you sell in the EU is a product with digital elements, the Cyber Resilience Act has required manufacturers since 11 September 2026 to report actively exploited vulnerabilities and severe incidents that affect the security of their products.5
The clock is short. The first step is an early warning within 24 hours of becoming aware, and the second a full notification within 72 hours. Manufacturers report once, through a single platform set up by ENISA, the EU’s cybersecurity agency.
Whether a particular agent is in scope, especially one you run as a hosted service, is a question for your legal team. The engineering point holds either way: 24 hours isn’t long to work out whether an injection was exploited and what it touched. That depends on a record of what the agent read, which tools it called and with what arguments, kept somewhere the agent can’t change.
What to ask your team
Stopping every injection isn’t a realistic goal with today’s models. Knowing what a fooled agent could do, and keeping that small, is.
For each agent you run, three answers show where you stand:
- which tools and permissions it has, compared with what its task uses
- which of its actions need a person, and how often that person says no
- when it was last tested by someone who knew how its defences work
Then the question that matters: if our agent followed the worst instruction it could read today, what could it do, and how soon would we know?
Sources
- OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, 4 August 2026. The ranking gives three quarters of the weight to a practitioner vote and one quarter to 6,639 incidents from public vulnerability databases and an AI-harm database.
- National Institute of Standards and Technology, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025), March 2025.
- Nasr, M. and 13 others, The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections, arXiv preprint, under review, 10 October 2025.
- OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications for 2026, 9 December 2025.
- European Commission, Cyber Resilience Act: Reporting obligations, page updated 11 September 2026. It applies to manufacturers of products with digital elements. Open-source software stewards report from 11 December 2027.


