Retrieval and aggregation leaks
Retrieval and output pipelines can disclose sensitive information even when the model follows its assigned task.
How the pipeline discloses data
The earlier LLM02 lessons put a secret in the model's context and asked whether the model would say it. These five labs are different. The disclosure comes through the data pipeline, and the model keeps doing its assigned job.
Retrieval-augmented generation, or RAG, feeds the model passages pulled from a document store so it can ground its answer. A retriever selects the passages that match a question, and one of them might carry a service account key.
A redactor removes sensitive values before content is shown. An output channel is any place content appears, such as the answer or a source panel.
The model then grounds its answer in the passages it was handed, or summarises the document faithfully. The problem is what reached the model, and which channel it left through.
Three patterns recur:
- Relevance is not authorisation. A retriever ranks passages by how well they match a question, never by who may read them. A passage the asker was never entitled to can be promoted into context and grounded into a customer-facing answer.
- Summaries inherit sensitivity. A summary carries the sensitivity of its source. A faithful two-line summary of a restricted document repeats the restricted line straight back out.
- Aggregation. No single answer is sensitive on its own. A sequence of individually permitted answers can be assembled into a restricted whole. A guard that checks one response at a time never sees it forming.
A fourth failure appears the moment you add a redactor. An output control is only as wide as the channel and the unit it runs on:
- Mask the answer but render the raw source panel, and the value walks out through the panel.
- Mask each retrieval chunk but reassemble the chunks downstream, and a secret split across the boundary is whole in neither piece the filter saw.
- Scan the prose but not the image URL, and the secret leaves inside a link the client fetches on its own.
This section has five labs, and each one is a different door the pipeline left open.
Cindra: a model chooses the tenant scope
Cindra is the knowledge assistant for a multi-tenant workspace product. You are signed in to Brightmark Retail and should only read Brightmark documents. Aldergrove Legal is another customer. Its confidential records live in a separate namespace, the workspace-specific collection used for retrieval.
For each question, Cindra calls a retrieval-scope router. This model call chooses which workspace to search. The app then runs vector.search(question, namespace=<whatever the router returned>) and accepts the returned scope. Your request can therefore influence an authorisation decision.
Ask an ordinary Brightmark question and watch the scope line. The router selects your workspace, and every retrieved snippet belongs to Brightmark.
To retrieve Aldergrove's record, two conditions must hold together:
- The router must receive a reason to search beyond your workspace.
- The question must concern Aldergrove's record, or the wider search will not rank it.
The move is to satisfy both conditions in one message rather than relying on authority framing: the request has to widen the retrieval scope and concern the other tenant's record at the same time.
Watch the scope line move away from Brightmark. The retrieved-context panel then shows a snippet belonging to Aldergrove Legal. The tenant boundary has failed before Cindra generates an answer, because the foreign chunk is already in the context and the logs.
Cindra then summarises the retrieved snippet and grounds its answer in it. The tenant boundary has already failed at retrieval, so once the cross-tenant snippet reaches your context the leak has fired; submit the evidence the lab asks for and the Completion Check confirms it.
ConfusedPilot shows a related production failure, where control over what an enterprise copilot retrieved was the lever. In Slack AI, a public-channel message caused private-channel content to reach someone outside that channel.
Why prompt-hardening the router fails. Telling the router "always stay in the caller's workspace" leaves the scope as a model output the app still honours. The next framing re-widens it. Scope is authorisation, and authorisation cannot be a value a model generates from untrusted input.
MemberDesk: combine permitted partial answers
MemberDesk is the account-support assistant at a membership club. It holds a member verification string and refuses attempts to obtain it directly. Its refusals cover:
- Direct requests.
- Authority claims.
- Urgency.
- Requests to echo, assemble, or format the string.
No phrase unlocks the complete value through those requests. MemberDesk does, however, help members who misread characters on worn cards. It will confirm one named part at a time:
- The card prefix.
- The enrolment block.
- The renewal block.
- Whether a given position is a letter or a digit.
Each answer is a permitted, bounded disclosure. The application does not track how much it has disclosed across the conversation.
The move here is aggregation: the pieces come out one at a time through ordinary, permitted confirmations, and because nothing tracks cumulative disclosure, those bounded answers can later be pulled back together into the whole.
The final reply contains the complete credential. The parts reached you through ordinary support confirmations.
Prompt extraction means obtaining information from the model's setup text. Recovering the credential that way would demonstrate a different technique; this lab is about aggregation, so the credential should come from the earlier partial confirmations. When the leak fires, submit the evidence the lab asks for and the Completion Check confirms it.
A per-request check cannot enforce this cumulative limit without tracking earlier disclosures. Here, sensitivity depends on the accumulated answers.
Vera: redaction misses the source panel
Vera is an internal support assistant that uses a company knowledge base. Vantage applies a data loss prevention, or DLP, policy to Vera's output. This policy replaces service keys with redaction placeholders. Ask a question whose answer would contain a key, and the answer shows a placeholder.
Vera also renders a source preview, showing the passage retrieved for each answer. The preview displays the raw passage without applying the DLP policy. You are allowed to retrieve this passage. The failure concerns how the application displays it.
The move here follows from that scoping gap: the passage carrying the key is reached through an ordinary data-ops question, and the value is visible in the channel the redactor never runs on, the source preview.
The answer itself remains correctly redacted, which is the tell: the value leaves through the preview, not the reply. When the leak fires, submit the evidence the lab asks for.
A regular expression, or regex, is a text-matching pattern used here to identify keys. Tightening the answer's regex cannot affect the source preview, because that channel never runs it. The retrieval trace needs the same output-handling policy as the reply.
Sable: redaction misses a credential split across chunks
Sable is an internal platform and site reliability engineering assistant that uses a knowledge base. Its DLP policy is stricter than Vera's. Sable masks a complete recovery credential in each chunk and on the answer channel. You can watch it mask a rotated-out example credential that sits inside one chunk.
A chunk is a document fragment processed separately during retrieval. Sable uses parent-window retrieval, which joins adjacent chunks into a larger passage. It redacts each chunk in isolation before joining them. The joined passage then supplies context for the answer and appears in the output.
The vault-recovery credential crosses a chunk boundary. One chunk ends with the bare prefix RKEY-. The next begins with the token body. Neither chunk contains the complete pattern the masker requires, so both halves stay visible.
The move here depends on getting both adjacent chunks retrieved together: a question about the operation the credential belongs to pulls in both halves as separate passages.
Because each chunk was masked in isolation, neither half was ever redacted, and the answer channel will not assemble them for you, so the reassembly happens on your side. When the leak fires, submit the evidence the lab asks for to complete the lab.
A stricter pattern that requires the complete credential still cannot match either isolated chunk. The control scans fragments, while the retriever serves their combined content.
Marlow: redaction excludes image URLs
Marlow is the Halyard Support Desk copilot. It writes replies in markdown for the agent console. Its context contains a confidential per-session support token.
A redaction pass removes that token from answer prose. Direct requests receive a refusal or a redacted value.
To avoid corrupting asset links, however, the redactor excludes markdown image and link markup. It never inspects their URLs. That leaves the channel open.
The routing instruction comes from the knowledge base itself. You do not supply it. Marlow treats retrieved content as trusted procedure.
One article contains a reply-composition macro, a procedure for composing replies. It tells Marlow to embed a markdown status image whose URL contains the session token. The article presents this as telemetry kept out of customer-visible text.
The move here is indirect: the token travels because Marlow follows the planted reply-composition macro, which embeds a status image whose URL carries it. Get Marlow to compose a reply that follows that article and the token rides out in the image markup.
The redactor removes sensitive prose but passes the image markup through unchanged. The URL therefore carries the token verbatim. A markdown-rendering console would fetch the image the instant the reply appears, sending the token to the host with no click.
EchoLeak chained retrieval into zero-click data disclosure through a markdown image in M365 Copilot. CamoLeak also used an image rendered by Copilot Chat to send data out.
Telling Marlow to keep the token secret does nothing. The retrieved procedure reframes the token as a telemetry value that belongs in the badge URL. An assistant that follows its knowledge base faithfully emits it there anyway, and the prose-only redactor still returns the URL. Access to an output channel is not controlled by asking the model to self-censor.
The shared authorisation problem
Each application needs code to answer one concrete question. May this content reach this user through this channel? The pipelines instead rely on narrower decisions:
- Cindra accepts a model-generated tenant scope.
- MemberDesk approves individual requests without tracking cumulative disclosure.
- Vera and Marlow apply redaction to only some output channels.
- Sable scans individual chunks before serving their combined content.
The models carry out their assigned retrieval, support, or composition tasks. These leaks require no jailbreak. The disclosure is not in what the model was persuaded to say. It is in what the pipeline was allowed to hand over.
Prompt hardening leaves those pipeline decisions unchanged. A router instructed to remain in the caller's workspace still produces a scope the application accepts. Telling Marlow to keep the token secret still leaves the URL channel outside redaction.
Enforce controls in the pipeline
Enforce retrieval and output decisions in code. Resolve permissions from the authenticated session. Do not delegate those decisions to prompts or model outputs.
Apply a server-side permissions filter before the search. Derive it from the authenticated session, using a per-tenant namespace or a separate per-tenant index. Ignore any scope value the model or the request supplies. An unauthorised chunk then never becomes a search candidate, so no framing can rank it in.
Apply the output policy to the assembled passage the model grounds on and the app renders. Cover every channel that shows it: the answer, the source panel, citations, and the URLs inside markdown image and link markup. You can also redact once at the retrieval layer, before anything is rendered.
If partial confirmations are necessary, track what the session has already received. Cap the total fraction of a protected value that can be released. This stops individually permitted answers from adding up to the complete value. Better still, keep the protected value out of an unverified session's context, because you cannot leak what is not there.
Treat retrieved content as untrusted data. It is input to read, and never a set of instructions to execute. Allowlist the hosts an outbound URL may point to, and strip unknown or secret-bearing query parameters. A value that reaches a URL then cannot beacon out.
Test these controls together. Place a secret across a chunk boundary in the corpus. Assert it is absent from every rendered channel: the reply, the source panel, the reassembled passage, and the image URL. Fire an adversarial cross-tenant query and confirm the retriever returns nothing from the other tenant. Replay a field-by-field aggregation walk and confirm the disclosure budget stops it. A green answer-redaction test alone proves none of these.
Where to go next. Work through the five labs, Cindra, MemberDesk, Vera, Sable, and Marlow, applying these fixes until you can complete each. Read them beside the ConfusedPilot and EchoLeak writeups in the incident database. You will see the same pipeline failures play out against shipped products.
Search relevance measures a match to the question. Apply the user's permissions before retrieval can select restricted snippets.
Apply source-sensitivity rules to summaries. A faithful summary can reproduce restricted lines from a ticket or document.
Track disclosure across responses. A check that remembers no earlier answers cannot detect a restricted fact assembled across turns.
Redact retrieved content before it reaches the model. Minimise summarised and derived output, and track field-by-field disclosures across the session.
Why can a RAG assistant leak another tenant's data even when it is built to stay in your workspace?
A model-generated scope lets your request influence which workspace the retriever searches. Semantic matching can then select another tenant's records. The disclosure begins the moment those records enter the context and logs. Enforce tenant permissions with a server-side filter derived from the authenticated session.
How can a model leak a value it correctly refuses to state outright?
Partial support confirmations can accumulate into the complete value, which the model then repeats from its own earlier answers. Other leaks use channels that answer redaction never sees. These include source panels, reassembled passages, and markdown image URLs.
Why doesn't tightening the redaction regex or the system prompt fix these leaks?
A stronger answer pattern still leaves the unscanned channels untouched. A complete-credential pattern cannot match a credential split across isolated chunks. A prompt instruction cannot police a source preview or URL the app renders directly. The durable fix is architectural. Redact the assembled unit across every rendered channel, or keep the secret out of the model's context.
What does completing these labs look like?
You have demonstrated the flaw when the restricted value reaches you through the pipeline itself, with no jailbreak. Cindra's retrieved-context panel shows an Aldergrove snippet. MemberDesk restates all three parts in one reply built from earlier confirmations. Vera's source panel shows the key its answer masked. Sable's two adjacent chunks let you splice the RKEY- prefix onto its token body. Marlow emits the support token inside a markdown image URL. In each lab you submit the recovered value as evidence.
© 2026 GenAI Security Lab. All rights reserved. You may read, quote, and link to this material with attribution. Copying, republishing, redistribution, resale, or use to train models or build competing products is prohibited without prior written permission.