genai
SECURITY LAB
Learning Path

LLM02: Sensitive Information Disclosure

Prompt injection is how you steer the model; this is what you walk away with. These leaks are rarely clever — a key pasted into a system prompt, a support bot quoting an internal runbook, a cache that hands one customer's answer to the next. We work them against live targets, then the fix that holds: keep the data out of the model's reach, because a prompt that says "never reveal this" is not access control.

OWASP LLM02
Sensitive information disclosure
When an AI reveals data that sits in its context but out of your bounds - a secret, a record field, another tenant's row.
The model can only leak what it can see - and everything you hand it, it can be persuaded to repeat.
genai
SECURITY LAB
Watch what counts as a leak. No break-in, no exploit - just an assistant answering with data that was never theirs to see.
0:00
1:22
The concept
Based on OWASP Top 10 for LLM Applications (2026), LLM02: Sensitive Information Disclosure · Content reviewed August 2026

Sensitive information disclosure happens when the model reveals data that it can reach but the user has no right to see. The data might be a secret in the prompt, or a private field in a customer record. It might also be another customer's row, or restricted text that the app retrieved from a document store.

Six faces, one root cause
Secrets & configAPI key · override codeInternal knowledgepolicy · debug tracePersonal dataPII / PHI over-shareCross-tenant bleedanother tenant's rowRetrieval leakrestricted RAG snippetAggregationinnocuous → wholeIN CONTEXT, OUT OF BOUNDSdata the asker was never entitled to
Each family maps to a section of this module — but every one is the same failure: sensitive data reached the model's context, and the asker was never entitled to it.

The root cause is almost always what is in the model's context window. The context window is all the text the model reads before it replies. It holds the system prompt, the conversation and any records or documents that the app adds. The model can only leak what it can see in this window.

The app might load a full account record, an API key or another customer's data into the window. A persuasive question is then enough to get any of it back out. No special attack technique is needed.

Anything in the window is elicitable
MODEL CONTEXT WINDOWAPI_KEY = sk-live-9f2c4a…account row — name, SSN, flagstenant B — invoice #4471retrieved — salary_band.pdfLLMany of it —asked back out
The model can only leak what it can see — but everything the data layer loads into the context window, it can see. A persuasive question is all it takes to read any line straight back out.

In security, need-to-know means a user sees only the data that their task requires. The context window has no such rule.

A task might need only a customer's name and ticket status. The window may still hold the full record, with an SSN, an internal risk flag and a support override code. The model cannot tell a field it may show from a field it must not show. A routine-sounding request can pull out any of them.

Need-to-know vs. the whole record
What the task needs
name
open-ticket status
✓ minimum context
What the model was handed
name, email, phone
SSN 402-19-8871
internal_risk_flag = HIGH
support_override_code
other_tenant_id
billing_address
✗ every extra field is leakable
Over-disclosure is returning more than the task requires. The handler needs a name and a ticket status — but it hands the model the entire row, and every extra field is now elicitable.

The same failure can cross accounts. A tenant is one customer organisation in an app that many organisations share. Tenant A and Tenant B, for example, might keep their invoices in one accounts table. Some apps limit each tenant to its own rows with an instruction in the prompt.

With only that instruction, one customer's question can return another customer's record. The flaw is missing tenant isolation in the data layer. Better wording in the prompt cannot fix it. Scope every query to the authenticated tenant before the result reaches the model.

Going deeper— Cross-tenant disclosure as an IDOR variant

Cross-tenant disclosure in an LLM app is the AI version of Insecure Direct Object Reference, or IDOR (CWE-639). In classic IDOR, the attacker changes a URL or API parameter to reach another user's resource. In an LLM app, the "parameter" is a natural-language reference that the model passes to a tool or query.

The model becomes an intermediary that does not know it is being misused. It turns a request such as "pull up account 88213" into a function call. The developer never expected that call to need an authorisation check, because the prompt said "only use the current customer's data".

Broken Object Level Authorization (BOLA) ranked #1 on the OWASP API Security Top 10 (2023) for this same pattern. The LLM layer makes it worse in two ways.

First, a user can socially engineer the model into building queries the developer never planned for, such as "check on my colleague's order". Second, the model has no concept of an authenticated session. A scoping rule in the prompt is therefore easy to get around.

The only durable fix is the one BOLA requires. Bind the tenant filter to the authenticated session at the query layer. The model's output then becomes an untrusted hint about what to look up. It never decides whose data comes back.

Cross-tenant bleed is an isolation failure
User · Tenant A"show my invoices"App queryno tenant filterACCOUNTS TABLETenant Ainvoice #A-2231Tenant Binvoice #B-4471both rowsAssistant leaksTenant B to Tenant A
Tenant scoping belongs in the data layer. Without a WHERE tenant_id = A filter, the query returns Tenant B's row too — and the model faithfully hands it to Tenant A. No clever prompt required.

Retrieval widens the blast radius. A RAG (retrieval-augmented generation) pipeline searches a document store and adds the best-matching snippets to the context. The search can pull in a restricted snippet that the user could never open directly.

Check every retrieved snippet against the asker's access rights before it enters the context. A relevance score shows only how well the text matches the question. It says nothing about who may read the text.

Aggregation is a second risk. The model can combine pieces that are each harmless into a fact that was meant to stay split apart.

Going deeper— The mosaic effect in LLM-assisted aggregation

Privacy research calls this the "mosaic effect", also known as a "composition attack". Sweeney (2000) showed that 87% of the U.S. population could be uniquely identified from just ZIP code, birth date and sex. Most people would consider each of those three fields harmless on its own.

In an LLM app, the risk grows. The model can do the aggregation within a single session. It can link answers from different tool calls or retrieved snippets. The user does not need to write any code.

Differential privacy (Dwork, 2006) was designed to limit this kind of cumulative leakage by adding calibrated noise. It applies to statistical queries over datasets. No differential-privacy equivalent exists for verbatim text retrieval, such as a RAG pipeline that returns exact snippets.

Current mitigations include rate-limiting sequential queries and monitoring for result sets that keep widening. Another mitigation is to partition retrieval indexes by clearance level. No principled theoretical bound exists for aggregation risk in free-text retrieval systems.

Retrieval pulls it in — aggregation assembles it
RETRIEVALuser queryvector storerestricted chunkuser can't open it directly→ into contextAGGREGATIONallowed: dept = Cardiologyallowed: shift = nightsallowed: badge photo on file+restricted wholethree allowed answers identify one patient
RAG can hand the model a chunk the user could never open directly. And several answers that are each fine to give can combine into a fact that is not — restriction has to hold on the derived whole, not just the parts.
Where to put the controls

The system around the model has to keep data confidential. Scope and redact data before it reaches the model. Never ask the model to keep a secret that it can read.

Shrink what the model can leak
1
Minimize
give only the fields the task needs
90% still leakable
2
Scope
tenant / row filter in the data layer
58% still leakable
3
Redact
strip secrets & PII on the way in
32% still leakable
4
Filter
block known patterns on the way out
12% still leakable
You cannot leak what is not in context. Each layer removes more of the leakable surface before data ever reaches — or leaves — the model. The bar is what an attacker could still pull out.
Key principles

Secrets and config values in the prompt are ordinary context. The model can be talked into repeating anything it can see.

Over-disclosure means returning the full record when the task needed only three fields. Minimise what enters the context.

Leaks between tenants, or between users' chat histories, are isolation failures. Scope rows in the data layer. Never rely on the prompt to scope them.

Retrieval and aggregation can assemble a restricted fact from pieces that are each harmless on their own.

The breach that makes it real
Breach replay
ChatGPT exposed other users' chat titles and billing info · 2023

A caching bug briefly let some ChatGPT users see other users' conversation titles and, for a subset of subscribers, names, email and payment addresses, and the last four digits of payment cards - data that should have stayed isolated per account.

You'll make an assistant surface data scoped to someone other than the asker.
Check yourself
Knowledge check
Which of these is an example of sensitive information disclosure?
Go deeper
Get the next part

New parts ship regularly. Leave your email and I’ll send each one — no spam, unsubscribe anytime.

© 2026 GenAI Security Lab. All rights reserved. You may read, quote, and link to this material with attribution. Copying, republishing, redistribution, resale, or use to train models or build competing products is prohibited without prior written permission.