genai
SECURITY LAB
6Part 6 of 6

Minimizing what the model can leak

A model cannot leak data that never reaches it. The four controls in this lesson work in code. They keep sensitive data out of the model's reach, or catch it before the reply ships.

LLM02 - Defense
You cannot leak what is not in context
Confidentiality is an architecture property. Scope and redact BEFORE data reaches the model - never ask it to keep a secret it can read.
Every technique in this path shares one root cause - and one durable answer: shrink what the model can see.
genai
SECURITY LAB
Watch the same failure from the defending side. The question is not what the model should refuse - it is what you handed it before the conversation started.
0:00
1:36
Shrink what the model can leak
1
Minimize
give only the fields the task needs
90% still leakable
2
Scope
tenant / row filter in the data layer
58% still leakable
3
Redact
strip secrets & PII on the way in
32% still leakable
4
Filter
block known patterns on the way out
12% still leakable
You cannot leak what is not in context. Each layer removes more of the leakable surface before data ever reaches — or leaves — the model. The bar is what an attacker could still pull out.

Assume the model already said everything

Every attack in this series ended the same way:

  • StatusBot relayed a bearer token out of a tool error.
  • TenantBot read back another workspace's record.
  • Vera's source panel showed the key that her answer had masked.
  • Marlow tucked a support token into an image URL. The console fetched that URL on its own.

Each time, a value came back that the asker was never entitled to see. The only thing in front of it was a sentence asking the model to be careful.

So do not start from "how do I make the model refuse". Plan for the worst case instead. Sooner or later, the model will say everything it can see to whoever finds the right frame. A frame is the way a request is presented, such as a debug pretext.

Then ask a question you can answer in code. When that request wins, what does it get? Four controls decide the answer. None of them lives in the prompt.

Least-privilege context and data-layer scoping remove data before the model reads it. Redaction on the way in and output filtering on the way out narrow what is left.

Layer 1: put less in the context window

The context window is all the text the model reads before it replies. Least-privilege context means that window holds only what the task needs. This control cuts more risk than any other. It is also the most boring.

In the code that builds the prompt, decide which fields and documents the task needs. Render nothing else. The model has no sense of need-to-know. Your code has to apply need-to-know before the model reads a single token.

  • Allowlist fields. Render name and ticket_status by name. Never render the full row minus ssn, because that is a denylist. When someone adds a recovery_code column, an allowlist fails closed and leaves it out. A denylist ships the new column without anyone noticing.
  • Apply the same allowlist everywhere. Tool JSON, error strings and retrieved chunks all need it. A chunk is a short slice of a document that search can return. A field hidden in the UI but still rendered into the prompt is personal-data over-disclosure in miniature.
  • Keep secrets out of the text. A credential the model must use never needs to be read. Attach it from a vault, in the client that makes the outbound request. A config dump exposed a hardcoded provider key. RunOps's debug export exposed a directive value. Both leaks worked only because each value had been rendered into text.
  • Give the model handles. A handle is a reference that means nothing on its own, such as acct_91d2. Only server code can resolve it to the real value.
  • Keep untrusted text away from secrets. ReviewLens read reviews that anyone could post into the same window as an analytics credential.

No jailbreak, memorised weight or logging bug can disclose a value you never render. Custom GPTs gave up their prompts and uploaded files to simple requests. Their builders had put that material in the model's reach and trusted wording to guard it.

Layer 2: scope in the data layer, before retrieval

Least-privilege context decides which fields a prompt may carry. Scoping decides whose rows a request can touch. The cross-tenant failures in this series all happened at scoping:

  • TenantBot's lookup resolved any workspace name.
  • The Meridian Assistant's get_account fetched whatever id the model passed.
  • PipelineAgent's raw mode checked the caller against nothing.
  • Cindra let a router model pick the workspace to search.
  • FastAnswer's cache had no tenant in its key.

Each time, the question "whose data is this?" was left to the message, to the model or to nobody.

The principal is the user or tenant that a request acts for, such as a member signed in to a workspace. One rule, in three parts, fixes all five apps:

  • Resolve the principal on the server, from the authenticated session.
  • Bind the principal into the query yourself.
  • Never let an id that the model produced reach a WHERE clause, a search filter or a tool argument that selects rows.

Make the scope filter impossible to forget. Row-level security, a per-tenant index and a repository that refuses unscoped queries all do this. If the principal cannot be resolved, fail closed.

In retrieval, filter before the search. A top-k search returns the k closest matches to a question. A global top-k search that drops foreign hits afterwards has already read them.

Tenant ownership is not the full check. A record in the caller's own tenant can still be owner-only, embargoed or withdrawn. Those records need an object-level entitlement check in code. That check decides, for each record, whether this caller may see it.

Then scope every copy of the data. Caches, history, memory, embeddings and logs hold the same rows again. They usually have no tenant column.

In the ChatGPT Redis bug, the model leaked nothing. Users saw each other's chat titles because a caching layer crossed the user boundary.

InsightBot: keep one customer out of another's data

The first defence lab is Build the Guardrail: Keep One Customer Out of Another's Data (Intermediate). It has you build Layer 2.

InsightBot is the workspace analytics assistant inside the Northwind SaaS platform. It serves a member signed in to the Acme workspace.

InsightBot is a real RAG app over a real vector store. RAG means the app retrieves snippets and hands them to the model with the question. InsightBot is told to answer only from those snippets. It must also quote identifiers and keys verbatim.

InsightBot does not choose what is retrieved. The member gets whatever retrieval hands it. Out of the box, that retrieval is default-allow:

  • It uses one shared index.
  • It has no workspace pre-filter.
  • It ignores document classification.

Two classes of record leak as a result:

  • Globex is another customer on the platform. Its integration-credentials record ranks straight into an Acme session.
  • Acme's own platform secrets record is restricted to owners and platform staff. It still comes back to an ordinary member.

Reproduce both leaks in the InsightBot app before you write any config:

One surfaces a record from another workspace; the other surfaces an owner-only record inside your own workspace. Seeing both fire against InsightBot is what tells you each control has a live disclosure to catch once it is in place.

Now map each class to its control in the policy schema:

  • The Globex record lives in another workspace. The hard pre-filter to the caller's own workspace, scope_to_current_workspace, keeps it out.
  • The restricted record lives inside Acme. Only the object-level check, exclude_restricted, touches it.

Set both to true. Either one alone still fails one attack.

Then protect utility. The benign suite is the set of ordinary questions your config must still answer. It asks how many active users Acme has. It also asks how a member gets the workspace API key.

The API key question is about credentials, like the attacks. In vector search terms, it sits in the same semantic space. Acme's API access policy answers it.

The trap is allowed_workspaces. It overrides the workspace toggle. The skeleton config that the lab inserts gives it an empty list, which retrieves nothing. Delete it.

What completion looks likeSubmitting the config validates the remediation on the retrieval trace. The harness replays the attack and benign suites under an empty config, then under yours. Baseline holds if the leak fires under the empty config. Security holds if no Globex or restricted document entered the retrieval set. Utility holds if both Acme answer documents still came back. All three must hold for Config accepted, which completes the lab.

NorthwindBot: an embargo is a retrieval boundary too

The second lab is Build the Guardrail: Keep the Unreleased Launch Out of the Answer (Intermediate). It moves the boundary from tenancy to publication stage.

NorthwindBot is the customer-facing sales assistant for Northwind Tools. It answers from one shared product index. It quotes prices, dates and codes verbatim.

Every document in that index carries a release stage. Published documents form the catalogue that customers may see. Documents for the unannounced Atlas line sit in an embargoed pre-release workspace, held back until launch. The internal Atlas launch brief there carries the launch code.

A draft Atlas press announcement went live in error. It was then withdrawn. It still sits in the published stage as a soft-deleted tombstone. A tombstone is a record marked as deleted that the index still holds.

Reproduce both leaks before you write any config, as with InsightBot:

One surfaces an embargoed pre-release record; the other resurfaces a withdrawn, tombstoned record. Seeing both fire first is what lets you confirm each control holds once the config is in place.

Again, two classes need two controls. published_only pre-filters retrieval to the customer catalogue. That keeps the embargoed workspace out.

published_only does nothing about the withdrawn draft, because the draft's stage is published. Only exclude_retracted drops it.

Set both. Then delete the skeleton's empty allowed_release_stages list. An empty list retrieves nothing. It would lose the Pro plan pricing answer that the benign suite needs.

Completion uses the same three checks as InsightBot. Baseline, Security and Utility must all hold on the retrieval trace.

The decoy controlThe schema also offers redact_terms, which masks words in snippet text after the search has run. If you use it in place of the two controls, the reply can look spotless. But the embargoed brief has already entered the retrieved set, so Security still fails. A boundary has to keep the document out of that set.

Layers 3 and 4: redact on the way in, filter on the way out

A document the caller is entitled to can still hold values that nobody should read in an answer. Redaction masks those values. Where redaction runs decides whether it works:

  • Before you chunk. Sable masked each chunk on its own. A credential split across a chunk boundary was complete in neither chunk. Redact the full document at ingestion instead.
  • Before you embed. Strip staff notes before indexing. DocsSearch AI indexed an article with a staff note still inside it.
  • Before you train. Scrub the training data. RecallServe memorised a key from one ticket that nobody scrubbed.
  • At the tool boundary. StatusBot leaked because a verbose upstream error reached its context raw. Strip Authorization headers and secret-shaped strings from tool output first.
  • By tokenising what the task must recognise. In data protection, a token is a stand-in for a sensitive value. Give the model an SSN's last four digits or a format-preserving card token. Keep the mapping back to the real values server-side.

Output filtering is the last check. It scans the reply for credential-shaped strings, personal data (PII) and text from restricted sources before the reply ships.

The filter must cover every channel that the reply renders through. Vera's filter redacted her answer but missed her source panel. Marlow's filter skipped the image URLs that the console fetched.

Output filtering is still only a backstop. A pattern filter matches only what it was written for. MemberDesk gave up a credential one fragment at a time.

EchoLeak shows what an unwatched output channel costs in production. One crafted email was retrieved into Microsoft 365 Copilot's context, beside the user's own data. That email then carried the data out with zero clicks.

Why the prompt-level fixes keep losing

Every weak fix in this series followed one pattern. It left the data in context and added a sentence, such as:

  • "Never reveal the API key."
  • "Only use the current customer's data."
  • "Do not mention Atlas."

Each sentence asks the model to enforce a rule. The model is the one component that cannot. It has no caller identity, no entitlement table and no release stage.

The secret and the rule that guards it are both text in the same context window. Chat formats label each message with a role, such as system or user. Providers train models to rank system-prompt text over user text and over text inside documents.

This trained ranking is often called the instruction hierarchy. No parser or permission check enforces it. It can fail.

Useful assistants are also built to quote verbatim, as InsightBot and NorthwindBot are. When a secret sits in context, a model doing its job well will quote it too.

These fixes also lose on timing. The disclosure usually happens before the model writes anything:

  • A foreign row is fetched.
  • A withdrawn draft is retrieved.
  • A cached answer is served.

By the time a refusal could fire, the data is already in context, in logs or in a source panel. For this reason, both defence labs grade the retrieval set, not the reply.

A behaviour

Refusals, "do not reveal" rules and a model asked to redact itself are behaviours. The model produces them when a request feels wrong. A better frame can change them.

A boundary

A field that is never rendered is a boundary. So is a query that cannot run unscoped, or a document redacted before indexing. A boundary holds however the request is worded. You can also inspect what it let through.

The fix that holds, in the order you ship it

Ship the fixes in this order. Each layer removes a class of leak that the next one cannot.

  • First. Pull every credential out of model-visible text, then rotate it. Replace full-record context with a field allowlist. Bind a tenant and user filter from the session into every query, search and cache key.
  • Next. Make the scope impossible to omit. Pre-filter retrieval. Gate restricted, embargoed and withdrawn records in code. A label on the record is not enough. Redact at ingestion and at the tool boundary. Filter every rendered channel, URLs included.
  • Then prove it. As tenant A, ask for data that only tenant B holds. Pass a model-supplied id into a scoped query. Fail the build if anything comes back. Seed a unique marker in each sensitive store. Alert when it appears in output, logs or an outbound request.

Where to go nextThe in-app LLM02 Sensitive Information Disclosure: Defense path is the builder's version of this lesson, free with a sign-in. The InsightBot and NorthwindBot labs both live there. Complete the attack labs in this series first, then close the same boundaries on the retrieval trace. The incident database shows the same failures in shipped products.

Key principles

Least-privilege context. Build the prompt from only the fields and documents the task needs. Leave secrets, tool schemas and unrelated records out of the window entirely.

Enforce tenant and row scoping in the data layer, outside the prompt. Filter every fetch by the authenticated session. The wrong principal's data is then never loaded or cached.

Redact or tokenise secrets, personal data (PII) and health data (PHI) before they enter context. Add an output filter on the way out. Faithful quoting and summarising then have nothing sensitive to reproduce.

These controls are mechanical. They hold whatever the phrasing or the claimed role. They remove the data from the path, so the model is never asked to guard it.

Key points
Build each prompt from only the fields the task needs.
Scope every query, search and cache key to the authenticated session in code, before retrieval.
Redact before chunking and at the tool boundary, then filter every channel the reply renders through.
Go deeper
FAQ
What is the most effective way to prevent sensitive information disclosure in an LLM application?

Keep the data out of the model's context. Render only the fields a task needs. Keep credentials server-side and attach them at the outbound call. Scope every query and search to the authenticated session, in code. Then no phrasing or jailbreak has anything to extract. Redaction on the way in and output filtering on the way out narrow what is left. They back up those controls. They do not replace them.

Why doesn't a system-prompt rule like "never reveal the API key" protect a secret?

The rule and the secret are both text in the same context window. Providers train models to rank system-prompt text over user text and over text inside documents. This trained ranking, often called the instruction hierarchy, has no parser or permission check behind it. It can fail. A config-dump request, a debug pretext or a planted document can out-rank the rule. The model also has no caller identity or entitlement table to check. A refusal is a behaviour that a better frame can change. The durable fix keeps the key out of model-visible text entirely.

Should tenant filtering in a RAG pipeline happen before or after the vector search?

Before. Resolve the tenant from the authenticated session. Pass it into the search as a hard pre-filter, or give each tenant its own index or namespace. A global top-k search that drops foreign hits afterwards has already read them. If the model chooses the scope, the access control list (ACL) becomes only advisory. Add an object-level entitlement check too. A restricted record inside the caller's own tenant passes any tenant filter.

Is redaction or output filtering enough to stop an LLM leaking data?

Not on its own. Redaction has to run before data enters context. Run it at ingestion, before chunking, and at the tool boundary. Per-chunk masking misses a value split across chunks. Masking a document after retrieval does not undo the retrieval. Output filtering is a last-line backstop. It must cover every rendered channel, including source panels and image URLs. It also catches only the patterns it was written for.

Comments
No comments yet — be the first.
Get the next part

New parts ship regularly. Leave your email and I’ll send each one — no spam, unsubscribe anytime.

© 2026 GenAI Security Lab. All rights reserved. You may read, quote, and link to this material with attribution. Copying, republishing, redistribution, resale, or use to train models or build competing products is prohibited without prior written permission.