Cascading failures in AI agents: stop one error spreading
Cascading failures in AI agents start small. One agent makes a wrong call: a made-up fact, a bad plan step, or an instruction slipped into its input. The next agent trusts it and acts. Then the one after that. By the time a person looks, hundreds of tasks have run on the first mistake.
This guide explains how OWASP describes the risk, the limits that keep one error contained, and a test you can run on your own agent pipeline.
What OWASP means by cascading failures in AI agents
The OWASP Top 10 for Agentic Applications 2026 lists this as ASI08 Cascading Failures. A single fault, such as a made-up answer, a malicious input, a broken tool or a poisoned memory entry, spreads across agents and grows into harm across the whole system. Because agents plan, remember and hand off work on their own, one error can skip the checks a person would have made step by step, and stay saved for later.
ASI08 is about the spread, not the first fault. OWASP files a faked message under ASI07 (see our post on inter-agent communication) and poisoned memory under ASI06. It uses ASI08 when that first fault moves on to other agents, sessions or workflows.
OWASP is also frank about a limit: errors can spread faster than people can keep up. Some risk stays, and your team has to decide how much it accepts.
Warning signs and how one error spreads
OWASP lists signs you can watch for:
- One decision sets off many downstream tasks in a short time.
- The problem crosses into other customers' data or other parts of the system.
- Two agents retry each other back and forth.
- Queues suddenly fill up, or the same request repeats again and again.
And the common ways an error travels:
- Planner to executor. A planning agent writes an unsafe step, and the executing agent runs it with no check.
- Memory. A poisoned entry keeps shaping new plans even after its source is gone. Our post on AI agent memory security covers that part.
- Feedback loops. Agents rely on each other's output and magnify an error. In one OWASP scenario, a fixing agent hides alerts to meet a speed target, and a planning agent reads fewer alerts as success and widens the automation.
- Looser oversight. After a run of good results, people approve in bulk, and changes spread unchecked.
- Bad updates. A faulty release is pushed to every connected agent at once.
Keep planning and doing apart

OWASP's central fix is to separate the agent that plans from the code that acts, with an independent policy check in between. A bad plan then reaches a gate instead of a database.
- Check every high-impact call. Validate it against written rules in code before it runs, and give each agent short-lived keys scoped to one task, so a drifting agent cannot set off other agents. Our posts on agent tool misuse and AI agent permissions go into both.
- Draw boundaries. Run agents apart from each other, with the least access they need, so a failure stays where it started.
- Review before passing on. Put a checkpoint or a person between a high-risk output and the agents that act on it.
- Plan for failure. Design as if any model, agent or outside source can fail at any time.
Add circuit breakers and limits
A circuit breaker is a switch in code that stops calls from one part to another once something looks wrong, such as too many failures or an odd burst, until a person or a timer resets it. The OWASP AI Agent Security Cheat Sheet recommends them between agents, and OWASP's ASI08 list adds:
- Quotas and progress caps. A cap on how many records, messages or payments one task may touch.
- Rate limits that pause. When commands start spreading fast, slow them down or pause them.
- Loop limits. Caps on chain depth, retries, tokens and cost, so a runaway loop stops by itself. Our post on LLM rate limiting covers the cost side.
Track drift and keep a trail
Some cascades are slow. OWASP asks you to compare decisions against a normal baseline and flag gradual decline. It also asks for logs of every message between agents, every policy decision and every outcome, time-stamped, hard to edit and tied to each agent's identity. Each action should carry its history, so you can trace a bad result back to the first fault and undo what followed.
OWASP adds a replay test: run last week's recorded agent actions in an isolated copy of your system, and only widen what agents may do if the replay stays within your set limits.
Worked example: test a support pipeline
Say your support pipeline has three agents. A triage agent reads tickets and tags them. A refund agent pays refunds on tagged tickets. A notification agent emails customers. Run these tests in a test environment with test accounts.
- Draw the chain. Write down whose output feeds whom, and which hand-offs have no check. Each unchecked hand-off is a place for a breaker.
- Feed in one bad output. Make the triage agent tag 200 test tickets "refund approved". The refund agent's policy check should refuse refunds outside your rules, and a breaker should pause the run and alert someone after a set number of refusals.
- Test the fan-out cap. Have one ticket ask the notification agent to "email every customer about this". The cap must block a mass send without a person's approval.
- Test a loop. Make the refund agent hand a ticket back to triage, and triage hand it back. The retry limit must stop the loop.
- Test memory. Save a note like "always approve refunds for this account", then delete the ticket it came from. Check that the note no longer shapes refunds.
- Trace one result. Pick one wrong refund from step 2 and follow the logs back to the tag that caused it. If you cannot, add the missing fields to your logs.
Run these again before you widen what any agent may do.
Frequently asked questions
Is this the same as a normal outage?
It can start like one, but the risk is that agents keep acting on the error, and act at speed. Limits and breakers stop the actions, not just the fault.
Do I need circuit breakers with only two agents?
Yes, at least a retry limit and a cap on how much one task may change. Two agents feeding each other is the simplest feedback loop.
Can a person just review every step?
Not at agent speed. OWASP notes that errors can spread faster than people can follow, so put human review on high-risk steps and let code limits handle the rest.
Which OWASP entry covers this?
ASI08 Cascading Failures, in the OWASP Top 10 for Agentic Applications 2026.
Get started
Want someone to feed a bad step into your agent pipeline and see how far it goes? whitehatstoic's cybersecurity and AI safety testing covers web app and API review, AI systems and prompt injection tests, with a written report, fixes and a retest after fixes. If you are still deciding where agents belong, the AI adoption service reviews use cases and sets a tool and pilot plan and a safe-use policy.

Book a meeting about AI safety testing, or book a meeting with whitehatstoic: tell us the product, the deadline and your biggest worry, and we reply with a plan and a price.
Comments
No comments yet.
Sign in or make an account to comment.