AI agent audit logs answer the question everyone asks after an agent does something wrong: what exactly did it do, and why was it allowed? Most app logs cannot say. They record a request and a status code. They do not…
AI security
27 posts
Rogue AI agents are hard to catch because each thing they do can look fine. One file read, one API call, one approval: all inside the rules. The harm is in the pattern. An agent that was meant to cut costs starts…
Agent goal hijack is what happens when an AI agent stops working on your task and starts working on someone else's. Nobody breaks into your server. The agent reads a web page, an email or a calendar invite with…
Cascading failures in AI agents start small. One agent makes a wrong call: a made-up fact, a bad plan step, or an instruction slipped into its input. The next agent trusts it and acts. Then the one after that. By the…
Agent tool misuse is when an AI agent does harm with tools it is allowed to use. No one stole a key and nothing broke in. The agent read a customer list with a lookup tool it was given, then emailed that list out with an…
Inter-agent communication is how the agents in your app hand work to each other. A planner sends a task to a research agent. The research agent sends results to a writer. A support agent asks a billing agent for a…
AI agent permissions decide what each agent in your app can reach, and on whose behalf. The easy setup is to give the agent the keys of whoever built it, or to let a main agent pass its full access to every helper it…
Human approval for AI agents is the check where a person says yes before the agent does something that matters: paying an invoice, deleting data, sending an email to a customer. Most teams add a confirm button and call…
Hidden Unicode prompt injection hides instructions in characters that a screen does not draw. A support ticket can look like "Please reset my password" to you, while the model reading it also sees a second sentence…
AI agent memory security is about what your agent remembers between conversations, and who that memory can reach. Memory makes an agent feel helpful: it knows your name, your plan and what you asked last week. It also…
MCP server security is about one question: what happens when you let an AI agent use a tool someone else wrote? The Model Context Protocol (MCP) is a standard way to connect an agent to tools, files and services. Adding…
AI generated code security is the work of checking code that a coding assistant wrote before it reaches your users. The assistant is fast and often right. But it does not know your threat model, it can invent package…
Image prompt injection is when instructions for an AI model are hidden inside a picture instead of typed as text. If your app lets users upload a screenshot, a photo of a receipt or a scanned document, and a vision model…
The OWASP Agentic Top 10 is OWASP's list of the biggest security risks for AI agents: systems that plan, call tools, keep memory and act on their own across several steps. If your app has moved from a chatbot that…
An LLM red team plan is a short, written plan for attacking your own AI feature before someone else does. You pick the feature, list what could go wrong, try to make it go wrong on purpose, and write down what you find.…
The OWASP LLM Top 10 2026 is the newest version of the best-known list of security risks for apps built on large language models. OWASP posted it on its Gen AI Security Project site in August 2026. The order changed more…
Vector database security is about protecting the store behind your AI search, RAG feature or chatbot memory. Teams often treat it as a cache that can be rebuilt, so it gets weaker rules than the main database. But the…
LLM sensitive information disclosure is what happens when an AI feature hands someone data they should never see: another customer's details, an internal document, an API key, or a line from a private chat. The model…
LLM supply chain security is about everything your AI feature depends on that your team did not build: the packages, the model, any adapters, the data, and the provider you send prompts to. A normal web app already has a…
LLM data poisoning is what happens when someone slips bad data into what your model learns from or reads, so it starts giving wrong, biased or harmful answers. If your team fine-tunes a model on its own examples, or…
LLM hallucination testing means checking, on purpose and before launch, how often your AI feature gives answers that sound right but are not. Most teams try a handful of friendly questions, see good replies, and ship.…
LLM rate limiting is what stands between your AI feature and a bill you did not plan for. Every model call costs money, takes time and uses capacity you share with every other user. A login form with no limit lets an…
LLM output handling is the part of AI security most teams skip. You validate what users type into a form. Then the model's reply goes straight into a web page, a database query or a shell command, with no checks at all.…
System prompt leakage is what happens when the hidden instructions behind your AI feature end up in a user's hands. The system prompt is the text your app sends to the model before every conversation: its role, its…
RAG security testing starts with a simple question: can the AI tell one user something only another user should see? Retrieval augmented generation, or RAG, means the app searches your documents for passages related to a…
Excessive agency is what OWASP calls the risk of an AI agent that can do more than its job needs. When you connect a model to email, files, a database or a payments API, every tool you hand it is something it might use…
Adding an AI assistant to a product is now a weekend job. Prompt injection testing, the work of checking what the assistant does when someone tries to turn it against you, usually is not done at all. Prompt injection is…