Vector database security: lock down embeddings in RAG apps
Vector database security is about protecting the store behind your AI search, RAG feature or chatbot memory. Teams often treat it as a cache that can be rebuilt, so it gets weaker rules than the main database. But the vectors in it are built from your documents, and OWASP now treats a leak of those vectors as a leak of the documents themselves.
OWASP covers this as LLM08:2025, Vector and Embedding Weaknesses, renumbered LLM09 in the 2026 edition. This guide covers the store itself: what is in it, who can query it, what goes in, and how long it keeps things.
What vector database security covers
An embedding is a list of numbers that stands for a piece of text, an image or code. A vector database stores those lists and finds the ones closest to a question, so the app can hand the matching chunks to the model.
OWASP's 2026 entry says this applies to more than RAG. Agent memory, semantic caches and duplicate detection use the same similarity search. Wherever that search decides what the model sees, the embedding layer is part of your security boundary.
The entry sums up the risks in one line: poisoning makes the system wrong, inversion makes it leak, jamming makes it silent, and weak access control makes it indiscriminate.

Embeddings are copies of your data
It is tempting to think a vector is safe because nobody can read it. OWASP says otherwise: stored embeddings can be turned back into a large part of the original text, and newer methods work without access to the model that made them.
That changes how you handle the store:
- Give the vector database the same access rules as the documents it came from.
- Treat backups and exports of the index at the same sensitivity as the source data.
- Count embeddings you send to an outside service as sending the documents.
- Encrypt the store at rest, with keys kept apart from the app.
OWASP's own scenario describes a leaked vector backup first labelled low risk because "only the embeddings leaked", then treated as a full document leak once the text was recovered.
Filter by tenant inside the search
Many apps keep every customer's vectors in one index and filter results afterwards. OWASP warns that the search then runs across everyone's data before the filter applies. Even when no document is shown, result counts, scores and timing can tell one customer what another has stored.
The fixes it lists:
- Apply the tenant filter inside the index query, and set it on the server from the signed-in user. A scope sent by the browser is a suggestion, not a control.
- Check access per chunk, since one mostly public document can hold a confidential paragraph.
- Use separate indexes per customer for the most sensitive data.
To test this from the chat side, with two accounts and real questions, follow our guide to RAG security testing.
Check what goes into the index
Anyone who can add content to the index can steer answers. OWASP's 2025 entry describes a resume with hidden white-on-white text telling the model to recommend the candidate. The screening tool read the hidden text and followed it.
Before you embed anything:
- Strip hidden text, zero-width characters and lookalike Unicode letters during extraction.
- Record where every embedding came from: the source, when it was added, how much you trust it, and which version of your pipeline made it. Then a bad batch can be found and removed.
- Keep outside content, such as scraped pages or customer uploads, in a separate index from internal documents.
Our guide to LLM data poisoning covers the wider problem of bad data in training and retrieval.
Blocker documents and stale vectors
OWASP describes a quieter attack called retrieval jamming. One planted document, built to match a common question, makes the model refuse or say it has no information. It carries no instructions, so a content scan will not flag it.
Old data causes its own problems. OWASP recommends you:
- Delete a document's vectors within a set time after the document is deleted, and run a regular check that the two still match.
- Re-embed the whole collection when you change embedding models, rather than mixing old and new vectors.
- Rerun your most important questions after each large import, and watch for answers that suddenly turn into refusals.
Logs, limits and keys
OWASP's 2026 entry lists a few small settings that block a lot of probing:
- Do not return raw similarity scores to the browser. They let someone test whether a given document is in your index.
- Rate-limit search and embedding endpoints per customer.
- Keep the embedding service's API key secret. A leaked key lets an attacker run your exact model.
- Keep logs that cannot be edited of every search: the tenant, the query, the IDs returned and their scores.
Worked example: audit your vector store in an afternoon
Say your product has an AI help search over each customer's documents, stored in one shared vector index. Work through these steps in a staging copy:
- Find every copy. List the index, its backups, any exports, and every outside service that receives embeddings. Check who can read each one.
- Read the query code. Confirm the tenant filter is part of the search call and comes from the server session, not a request field.
- Probe across tenants. As tenant A, search for a phrase that exists only in tenant B's documents. Compare result counts, scores and timing with a phrase that exists nowhere.
- Plant a hidden line. Upload a document with white-on-white text and check whether it reaches the index.
- Delete and check. Delete a document, wait the time your policy allows, then confirm its vectors are gone.
- Check the response. Look at what the search API returns to the browser. Remove raw scores if they are there.
Each step that fails becomes a finding with a clear fix.
Frequently asked questions
Are embeddings personal data?
Treat them that way. OWASP says stored embeddings can be inverted to recover the source text, so a leak of embeddings should be handled like a leak of the documents.
Is a metadata filter enough to separate customers?
Only if it runs inside the search query and is set by the server. A filter applied after the search, or one the browser can change, still lets one customer learn about another's data.
Does this apply to semantic caches and agent memory?
Yes. OWASP's 2026 entry covers any feature that uses similarity search to decide what the model sees, including caches, memory and duplicate detection.
Get started
Building or running a RAG feature? whitehatstoic's cybersecurity and AI safety testing covers web apps, APIs and AI systems, with a written report, fixes and a retest after fixes, scoped after a short call. If you are still building the feature, full-stack development covers web and mobile apps from database to deploy.

Book a meeting with whitehatstoic, pick one or more topics, and tell us the product, the deadline and your biggest worry. We reply with a plan and a price.
Comments
No comments yet.
Sign in or make an account to comment.