LLM02: Sensitive Information Disclosure
Prompt injection is how you steer the model; this is what you walk away with. These leaks are rarely clever — a key pasted into a system prompt, a support bot quoting an internal runbook, a cache that hands one customer's answer to the next. We work them against live targets, then the fix that holds: keep the data out of the model's reach, because a prompt that says "never reveal this" is not access control.
Sensitive information disclosure happens when the model reveals data that it can reach but the user has no right to see. The data might be a secret in the prompt, or a private field in a customer record. It might also be another customer's row, or restricted text that the app retrieved from a document store.
The root cause is almost always what is in the model's context window. The context window is all the text the model reads before it replies. It holds the system prompt, the conversation and any records or documents that the app adds. The model can only leak what it can see in this window.
The app might load a full account record, an API key or another customer's data into the window. A persuasive question is then enough to get any of it back out. No special attack technique is needed.
In security, need-to-know means a user sees only the data that their task requires. The context window has no such rule.
A task might need only a customer's name and ticket status. The window may still hold the full record, with an SSN, an internal risk flag and a support override code. The model cannot tell a field it may show from a field it must not show. A routine-sounding request can pull out any of them.
The same failure can cross accounts. A tenant is one customer organisation in an app that many organisations share. Tenant A and Tenant B, for example, might keep their invoices in one accounts table. Some apps limit each tenant to its own rows with an instruction in the prompt.
With only that instruction, one customer's question can return another customer's record. The flaw is missing tenant isolation in the data layer. Better wording in the prompt cannot fix it. Scope every query to the authenticated tenant before the result reaches the model.
Going deeper— Cross-tenant disclosure as an IDOR variant
Cross-tenant disclosure in an LLM app is the AI version of Insecure Direct Object Reference, or IDOR (CWE-639). In classic IDOR, the attacker changes a URL or API parameter to reach another user's resource. In an LLM app, the "parameter" is a natural-language reference that the model passes to a tool or query.
The model becomes an intermediary that does not know it is being misused. It turns a request such as "pull up account 88213" into a function call. The developer never expected that call to need an authorisation check, because the prompt said "only use the current customer's data".
Broken Object Level Authorization (BOLA) ranked #1 on the OWASP API Security Top 10 (2023) for this same pattern. The LLM layer makes it worse in two ways.
First, a user can socially engineer the model into building queries the developer never planned for, such as "check on my colleague's order". Second, the model has no concept of an authenticated session. A scoping rule in the prompt is therefore easy to get around.
The only durable fix is the one BOLA requires. Bind the tenant filter to the authenticated session at the query layer. The model's output then becomes an untrusted hint about what to look up. It never decides whose data comes back.
Retrieval widens the blast radius. A RAG (retrieval-augmented generation) pipeline searches a document store and adds the best-matching snippets to the context. The search can pull in a restricted snippet that the user could never open directly.
Check every retrieved snippet against the asker's access rights before it enters the context. A relevance score shows only how well the text matches the question. It says nothing about who may read the text.
Aggregation is a second risk. The model can combine pieces that are each harmless into a fact that was meant to stay split apart.
Going deeper— The mosaic effect in LLM-assisted aggregation
Privacy research calls this the "mosaic effect", also known as a "composition attack". Sweeney (2000) showed that 87% of the U.S. population could be uniquely identified from just ZIP code, birth date and sex. Most people would consider each of those three fields harmless on its own.
In an LLM app, the risk grows. The model can do the aggregation within a single session. It can link answers from different tool calls or retrieved snippets. The user does not need to write any code.
Differential privacy (Dwork, 2006) was designed to limit this kind of cumulative leakage by adding calibrated noise. It applies to statistical queries over datasets. No differential-privacy equivalent exists for verbatim text retrieval, such as a RAG pipeline that returns exact snippets.
Current mitigations include rate-limiting sequential queries and monitoring for result sets that keep widening. Another mitigation is to partition retrieval indexes by clearance level. No principled theoretical bound exists for aggregation risk in free-text retrieval systems.
The system around the model has to keep data confidential. Scope and redact data before it reaches the model. Never ask the model to keep a secret that it can read.
Secrets and config values in the prompt are ordinary context. The model can be talked into repeating anything it can see.
Over-disclosure means returning the full record when the task needed only three fields. Minimise what enters the context.
Leaks between tenants, or between users' chat histories, are isolation failures. Scope rows in the data layer. Never rely on the prompt to scope them.
Retrieval and aggregation can assemble a restricted fact from pieces that are each harmless on their own.
A caching bug briefly let some ChatGPT users see other users' conversation titles and, for a subset of subscribers, names, email and payment addresses, and the last four digits of payment cards - data that should have stayed isolated per account.
© 2026 GenAI Security Lab. All rights reserved. You may read, quote, and link to this material with attribution. Copying, republishing, redistribution, resale, or use to train models or build competing products is prohibited without prior written permission.