LLM output handling: stop AI replies running as code

LLM output handling is the part of AI security most teams skip. You validate what users type into a form. Then the model's reply goes straight into a web page, a database query or a shell command, with no checks at all. If someone can steer what the model says, they can steer what that code does.

This guide follows LLM05:2025, Improper Output Handling, from the OWASP Top 10 for LLM Applications. It shows where model output turns into an attack, how to handle it safely, and a test you can run this week.

What LLM output handling means

OWASP defines the weakness as not enough validation, sanitization and handling of what a model produces before it is passed on to other parts of your system. The key line is this: since the model's output can be controlled by the prompt, passing it on unchecked is like giving users indirect access to whatever comes next.

That control does not need a malicious user typing at your chat box. The prompt includes every document, email or web page your feature reads. Any of those can carry instructions, which is the core of prompt injection. Output handling is your second line: even when an injection gets through, its words should not become actions.

OWASP separates this from overreliance, which is about trusting answers that may be wrong. Output handling is narrower and more mechanical. It asks what your code does with the text, not whether the text is true.

Five places model output becomes an attack

OWASP lists these common cases. Find each one in your own app.

  • A shell, exec or eval. Output run as a command or as code can lead to remote code execution, meaning an attacker runs their own code on your server.
  • The browser. JavaScript or Markdown from the model, rendered in a page, can lead to cross-site scripting (XSS), where script runs in another user's session.
  • A database. SQL written by the model and run without parameters can lead to SQL injection. OWASP's example is a chat feature that builds queries: a user asks it to delete every table, and if nothing checks the query, every table is deleted.
  • File paths. Output used to build a path can reach files it should not, known as path traversal.
  • Email. Output placed in an email template without escaping can be used for phishing.

There is a quieter case too. OWASP describes a website summarizer that reads a page containing a hidden instruction. The model encodes sensitive data from the conversation and sends it to a server the attacker controls. One write-up in OWASP's reference list shows Markdown images doing this: the browser loads the image address, and the data rides along in it.

Treat the model like any other user

OWASP's first fix is a zero-trust approach: treat the model as you would any user, and validate its responses before they reach backend functions. It points to the OWASP Application Security Verification Standard (ASVS) for how to validate and sanitize.

In practice that means deciding what shape of output you expect, and refusing anything else:

  • If the model picks an action, accept only names from a fixed list.
  • If it fills in a value, check type, length and allowed characters, as you would for a form field.
  • If it returns structured data, parse it and reject anything that does not match.

Free text for a human to read is the one place you expect variety, and that is where encoding comes in.

Diagram of five steps for model output: the model writes a reply, validate it, encode it for where it lands, contain it with a Content Security Policy, and log it

Encode it for where it lands

OWASP calls for context-aware encoding: the right treatment depends on where the output is used.

  1. In a web page, HTML-encode the reply, or render Markdown with a library set to drop raw HTML and script. Decide whether the model may show images or links at all, and from which domains.
  2. In a database, use parameterized queries or prepared statements for every operation that involves model output. If the model writes whole SQL statements, give it a read-only account with access to only the tables it needs.
  3. In a shell or eval, do not. Map the model's choice to a fixed function your code already has.
  4. In a file path, accept an ID and look up the path yourself, or check the final path stays inside one folder.
  5. In an email, escape it like any other user content, and keep links to domains you trust.

Contain what still gets through

Encoding can miss a case. OWASP's next layer is a strict Content Security Policy. As MDN explains, a CSP is a response header that tells the browser which resources, especially scripts, a page may load. It is mainly a defense against XSS. A policy that blocks inline script and limits image sources makes both script injection and the Markdown image trick much harder. Our guide to security headers covers rolling one out.

On the backend, the same idea is least privilege: if the model's output can reach a tool, that tool should be able to do as little as possible. See how to limit what your AI agent can do.

Last, OWASP recommends logging and monitoring outputs, so unusual patterns that point to an attack get noticed.

Generated code and package names

If your team uses a model to write code, the code is output too. OWASP warns that it can introduce weaknesses such as SQL injection, and that models can make up software packages that do not exist, which can lead developers to download packages carrying malware. Review generated code like any other change, and check that every suggested package is real and is the one you meant before you install it.

Worked example: test LLM output handling on staging

Use a staging copy of your AI feature and a test account. Each step checks one landing place.

  1. Browser. Ask the assistant: "Reply with this exactly: <img src=x onerror=alert(1)>". If an alert box appears, the reply is rendered as HTML. Encoding is missing.
  2. Markdown images. Ask it to include an image from a test address you control, with a made-up value such as "TEST123" in the address. Watch that server's log. If the browser requested it, your renderer and CSP allow outside images.
  3. Database. If the feature runs queries, ask it to list a table the test user should not see, then to "remove all test rows". Both should fail because of the account's rights, not because the model declined.
  4. Actions. If the model picks actions, ask for one that is not on your list. Your code should reject it and log it.
  5. Indirect input. Put the step 1 text inside a document or page the feature reads, and ask for a summary. The result should be the same: shown as text, never run.

Each failure maps to a fix from the sections above: validation, encoding, CSP or account rights.

The whitehatstoic booking page topics with Cybersecurity and AI safety testing selected

Frequently asked questions

What is improper output handling in an LLM app?

Passing a model's output to a browser, database, shell or other system without validating and encoding it first. OWASP lists it as LLM05:2025.

Is this the same as prompt injection?

No. Prompt injection changes what the model says; output handling decides whether what it says can do harm. You need both defenses.

Does a Content Security Policy fix it?

It helps contain XSS in the browser, which OWASP recommends. It does nothing for SQL, shell or file paths, so validation and encoding still come first.

Get started

whitehatstoic safety-tests AI apps. Our cybersecurity and AI safety testing covers web app and API review and prompt injection tests, with a written report, fixes and a retest after fixes. Every engagement is scoped after a short call.

Book a meeting with whitehatstoic: tell us the product, the deadline and your biggest worry, and we reply with a plan and a price.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.