Excessive agency: how to limit what your AI agent can do

Excessive agency is what OWASP calls the risk of an AI agent that can do more than its job needs. When you connect a model to email, files, a database or a payments API, every tool you hand it is something it might use at the wrong moment. This guide explains where excessive agency comes from, walks through one worked example, and shows you how to test and trim your own agent before your users meet it.

What excessive agency means

A chat model on its own can only produce text. An agent is different: it can call tools, and those tools change things in the real world. It can send a message, edit a record, delete a file or start a refund. OWASP's LLM06:2025 entry describes excessive agency as the weakness that lets the model take damaging actions when its output is unexpected, ambiguous or manipulated.

The important word there is manipulated. A model reads everything you give it, including emails, web pages and documents written by strangers. If any of that text contains instructions, the model may follow them. That is prompt injection, covered in our guide to prompt injection testing. Excessive agency decides how much harm a successful injection can do. You cannot fully prevent the first, so you limit the second.

The three root causes

OWASP traces excessive agency to three causes. Each one is a design decision you can review.

  • Too much functionality. The agent has tools it does not need, or a tool does more than the task requires. A plugin meant to read documents that can also edit and delete them is the classic case.
  • Too much permission. A tool connects with a powerful identity. OWASP's example is a database connection that can update, insert and delete when the feature only ever needs to read.
  • Too much autonomy. The agent takes high-impact actions without anyone confirming them, such as deleting a user's documents with no prompt.

Most real incidents combine at least two of these. Fixing any one of them shrinks the damage.

A worked example: the inbox assistant that could send

This example follows a scenario OWASP gives. You build an assistant that summarizes a user's email each morning. To read the inbox, you connect a mail plugin. The plugin you pick can read, and it can also send, because that is how it was built. The assistant only needs to read.

One morning an email arrives from an outside address. Hidden in its text is an instruction telling the assistant to find messages about invoices and forward them to an outside address. The assistant reads it as part of the inbox and follows it. Nothing was hacked in the usual sense. The model did what the text asked, with the tools you gave it.

Now walk through the three causes:

  • Functionality: a read-only mail tool would have had no way to forward anything.
  • Permission: connecting with a read-only OAuth scope would block sending even if the tool tried.
  • Autonomy: if every outgoing message needed the user to review and approve it, the forward would have stopped at a confirmation screen.

OWASP also suggests rate limiting, so that even a successful attack can only move a little before someone notices. Any one of these fixes would have turned this data leak into a failed attempt, and together they leave far less room for the next one.

How to test your AI agent for excessive agency

Do this against your own agent, in a development copy, with test accounts and test data.

  1. List every tool and what it can really do. Not what you use it for, but every action its API allows. For each, write down the identity it connects with and that identity's permissions.
  2. Mark what the feature needs. Next to each tool, note the smallest set of actions the feature requires. Every gap between the two lists is a finding.
  3. Find the high-impact actions. Sending, deleting, paying, changing permissions and sharing data outside the account. For each, check whether a person must approve it first.
  4. Plant test instructions in content the agent reads. Put a harmless test instruction in a test email, document or web page the agent will process, such as asking it to call a tool it should not need. Watch whether it tries. You are checking your own design, so use test data only.
  5. Check that the refusal happens outside the model. If the agent tries a forbidden action, the tool or the downstream system should refuse it, not the model's own judgement.
  6. Check the logs. Every tool call should be recorded with who asked, what ran and what changed, so you can see an attempt even when it fails.

Fixes, from strongest to weakest

OWASP's mitigations map neatly onto the root causes. Roughly in order of how much they help:

  • Remove tools you do not need. A tool that is not connected cannot be misused.
  • Narrow each tool. Prefer a tool that does one thing over a general one. Avoid open-ended tools such as "run a shell command" or "fetch any URL" when a specific one would do.
  • Shrink the permissions. Give each tool its own identity with the least access, such as read-only access to one table.
  • Act as the user, not as the system. Where you can, run tools in the signed-in user's context with the smallest OAuth scope, so the agent can never reach more than that user could.
  • Ask before high-impact actions. Show the user what is about to happen and let them approve it.
  • Enforce rules downstream. OWASP calls this complete mediation: the system receiving the request checks authorization itself instead of trusting the model.
  • Log, monitor and rate-limit. These do not prevent misuse, but they limit how far it spreads and help you spot it.

Why the model is not your security boundary

It is tempting to fix excessive agency in the system prompt: "never forward email", "never delete anything". Keep those instructions, but do not rely on them. The model treats them as text, alongside all the other text it reads, and a cleverly written input can outweigh them. A permission the tool does not have is a rule nothing can talk its way past. Put the hard limits in code, identities and approvals, and treat the prompt as guidance.

When to bring in an outside tester

Run the steps above yourself first. An outside test is worth it when your agent can move money, touch customer data, or act across many accounts, or when customers ask how you tested it. Someone who did not build the agent will try routes the builders never considered.

That is part of what whitehatstoic does. Its cybersecurity and AI safety testing covers web app and API review and prompt injection tests, ends with a written report and fixes, and includes a retest after the fixes. Testing is scoped after a short call. No test can promise to find every weakness, so the aim is to find the ones that matter and help you close them.

Frequently asked questions

Is excessive agency the same as prompt injection?

No. Prompt injection is one way an agent gets steered into a bad action. Excessive agency is having the tools, permissions or freedom to carry that action out.

Does a strong system prompt solve excessive agency?

It helps, but it is not a control you can rely on. Hard limits belong in tool design, permissions and user approval steps.

Which actions should need user approval?

Anything that is hard to undo or reaches outside the account: sending messages, deleting data, payments, refunds, sharing and permission changes.

Do read-only agents need testing too?

Yes. A read-only agent can still reveal data the user should not see, so check that its reads are limited to what the signed-in user may access.

Get started

Building something that has to be safe? Book a meeting with whitehatstoic: tell us the product, the deadline and your biggest worry, and we reply with a plan and a price.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.