genai
SECURITY LAB
1Part 1 of 6

Cross-tenant and history bleed

One user's or tenant's data surfaces to another through shared context, cached state, or conversation-history bleed - an isolation failure, not a phrasing one.

LLM02 - Technique 4
Another tenant's row, in your answer
When scoping lives in the prompt instead of the query, one customer's question can surface another customer's record.
Cross-tenant bleed is not a clever prompt - it is a missing scope check in the data layer.
genai
SECURITY LAB
Watch a boundary fail quietly. One customer asks a routine question about an account, and the record that comes back belongs to somebody else entirely.
0:00
1:48
Cross-tenant bleed is an isolation failure
User · Tenant A"show my invoices"App queryno tenant filterACCOUNTS TABLETenant Ainvoice #A-2231Tenant Binvoice #B-4471both rowsAssistant leaksTenant B to Tenant A
Tenant scoping belongs in the data layer. Without a WHERE tenant_id = A filter, the query returns Tenant B's row too — and the model faithfully hands it to Tenant A. No clever prompt required.

When the answer is right but the person is wrong

Most sensitive-disclosure failures come down to getting a model to say something it is holding in its context. This one is different, and it is where this path starts: it is about getting the model to say the right thing to the wrong person. You sign in as yourself, ask an ordinary account question, and the reply carries another customer's record — their plan, their balance, a token that was never yours. You did not out-argue a guardrail. The data was sitting in reach, and the only thing meant to keep it away from you was a sentence in the prompt, a similarity score, or nothing at all.

That is what makes cross-tenant bleed different from a jailbreak. The model has no way to know which tenant a block of context legitimately belongs to. If your record and a stranger's record are both in the window, the model can be steered to either. The boundary between the two of you is a property of the system's query, cache and session design — and the model has zero visibility into that design. It cannot enforce an isolation it was never shown. So when isolation lives in the prompt, any request that names or implies the other party has a clear path to data that was supposed to be partitioned.

The three labs in this section are the same defect through three different doors: a lookup tool that resolves any name, a lookup tool that trusts an id from your message, and a cache that matches on meaning with no tenant in the key. Learn to see the door and the move is the same every time.

TenantBot: a lookup that resolves any name

Lab: cross-tenant-disclosure. The target is TenantBot, AcmeCloud's multi-tenant account assistant, and you are signed in as the Initech workspace. The Workspaces list shows a second workspace, Globex, badged as not yours. TenantBot holds nothing itself; every account answer comes from a backend get_workspace_record tool it calls with a workspace name.

Here is the flaw. That tool resolves whatever name it is handed and never checks whether the caller belongs to that workspace. The only thing scoping you to Initech is a line in TenantBot's instructions telling it not to volunteer other workspaces on its own initiative. Read that carefully: it governs what the bot does unprompted. It says nothing about what happens when you ask directly.

So the move is simply to ask directly: a request that names the other workspace reaches it, because nothing downstream checks that you belong there.

TenantBot calls the tool for Globex, the tool returns Globex's record because nothing stops it, and the bot relays it inside an Initech session. When another workspace's record surfaces to a tenant with no claim to it, the leak has fired; submit the evidence the lab asks for and the Completion Check confirms it. Note what would not fix it — telling the model to "only answer about Initech" leaves the tool free to resolve any name, so a reworded request still reaches Globex.

Meridian Assistant: an id lifted from your own statement

Lab: bank-assistant-cross-account-leak. The target is the Meridian Assistant inside the bank.meridian.example online-banking portal. You are signed in as Jordan Reyes, and the profile carries one account, AC-1001. The assistant answers every account question by calling a get_account tool keyed on an account_id — and that tool performs no ownership check. It fetches whatever id the model passes. This is the classic confused deputy: a trusted component (the tool, acting with the bank's privileges) acting on input it should never have trusted (an id you chose).

There are two parts to this, because the id you need is not yours to begin with. The first is that your own statement legitimately shows more than a balance.

Your statement names a second account as the counterparty on a transfer that is on hold — an account this profile does not own, and the confidential internal field you are after sits on its record. The assistant will ask, once, what that account is to you; that is a single question, not a gate.

The class of move is that the token is a confidential field the assistant surfaces only on a full or internal record request, not on an ordinary "who is this account" question; and it treats entitlement as the backend's decision while acting as a pass-through, so it will hand back a record it never owned. Work the lab using its own hints; when the token lands in the reply, the Completion Check confirms it.

FastAnswer: a cache that matches on meaning, not on tenant

Lab: semantic-cache-cross-tenant-leak. The target is FastAnswer, a shared multi-tenant assistant at fastanswer.app, and you are signed in to the Blueharbor Labs workspace. To save tokens, FastAnswer puts a semantic response cache in front of the model: it embeds your question and, if a previously-answered question is close enough in embedding space, it serves that cached answer instead of calling the model again. Northwind Trading is another customer on the same instance, and one of Northwind's cached answers carries a service account credential.

The defect is entirely in the cache key. The lookup searches one shared namespace and returns the nearest entry by cosine similarity — with no tenant in the key and no check that the entry's owner is you. This is not a missing SQL WHERE clause: the two questions never match textually. It is the embedding model that decides your differently-worded question "means the same thing" as Northwind's, and that similarity judgment is what carries the answer across the boundary. Similarity is not authorization.

The class of move is to collide on semantic similarity. The Trending panel lists cached topics across every workspace, so it tells you which cross-tenant topic to aim for without handing you anything to paste. From there the goal is a question of your own that lands close, in embedding space, to the other tenant's cached question on that topic — specific enough to collide, but neither naming the other tenant nor so generic that it matches nothing.

Watch the Cache panel: on a hit it shows served from cache, the source workspace, and a cross-tenant flag. A hit sourced from another tenant means your question collided with theirs and their cached answer came back to you; a fresh answer or your own workspace means it has not collided yet, so the work is to keep adjusting the wording until it does. Work the lab using its own hints; when the leak fires, submit the evidence the lab asks for to complete the lab.

History and shared state bleed the same way

The three labs are tool- and cache-shaped, but the same defect wears a fourth costume: conversation state that outlives one session. A support bot that reuses a session across callers hands the next person the previous person's transcript, so "summarise what we discussed earlier" walks straight into someone else's email and order id. A cache keyed by route instead of by user and tenant serves the last caller's account summary to the next person who hits the endpoint. None of these is a phrasing bug either — a transcript is per-user state, and it needs the same isolation as any row.

This is where the production incidents live. The ChatGPT Redis bug showed users each other's chat titles and history because a caching layer briefly crossed the user boundary — the archetypal shared-state bleed. Slack AI could be steered, by a message planted in a public channel, into leaking private-channel data to someone who was never in the channel: retrieval pulled the private data in, and nothing downstream held the boundary. Different products, one root cause: state that was not scoped to the authenticated user and tenant.

One isolation failure, many doors

Read the four together and the root cause is single. In every case a decision that belonged in code — does this caller own this row? — was left somewhere that cannot enforce it: a prompt line the model can be argued past, a similarity score that does not encode ownership, a cache key that omits the tenant, a session that was reused. The data was fetched, served, or carried over before the model weighed any instruction, so by the time the model is answering, the foreign record is already in context, already logged, already reasoned over. An instruction not to use it is a request; the retrieval already happened.

Scoping in the prompt

"Only ever use the current customer's data." A request the model can be talked past — and it never stops the wrong row being fetched, only asks the model not to say it.

Scoping in the query

A WHERE tenant_id = :session_tenant filter, enforced before any row is returned. The foreign record never enters context, so there is nothing to bleed.

Scope at the source, and there is nothing left to bleed

The durable fix is the same for all three doors, and none of it is a prompt instruction. Derive the tenant and user from the authenticated session, never from a name or id in the message, and filter every fetch by it. TenantBot's get_workspace_record must verify the caller is a member of the workspace before it returns anything; the Meridian tool must reject an account_id the session does not own before it reads a record. Bind the identity in code and the confused deputy has nothing to be confused about — a lookup for data you do not own returns nothing, however the request is phrased.

Treat every store built over private data as a data store with the same boundary. A response cache built over tenant-private answers is one: put the tenant in the cache key — a per-tenant namespace, or reject any hit whose owner is not the caller — so a hit can only ever be your own prior answer. Raising the similarity threshold does not do this; it only narrows the collision window while leaving one shared namespace. And scope conversation history and session context per user: never reuse a transcript between people.

The one habitBefore you trust any layer to isolate tenants, ask where the tenant id comes from. If the answer is "the message", "the model", or "nowhere, it's in the prompt", you have found the leak. If it is "the authenticated session, enforced in the query", there is nothing cross-tenant left to bleed. Browse the incident database to watch the same scoping failure play out against shipped products, and run the three labs above until the door is obvious on sight.

Key principles

Cross-tenant bleed is a scoping failure, not a phrasing one: the wrong tenant's data was placed where the model could reach it, and isolation lived in the prompt instead of the data layer.

Uncleared conversation state turns 'summarize the previous conversation' into another person's data - the model has no visibility into session boundaries it was never shown.

The tell is correct, specific data for the wrong principal: another tenant's record token or billing, or a prior session's note echoed to a new user who has no claim to it.

Filter every fetch by the authenticated tenant id at query time and evict context per session - the model cannot leak a row the query never returned.

Key points
Tenant scoping belongs in the data layer, not the prompt.
Cached or shared context can carry another session's data.
A tool that takes an account id and skips the ownership check is the leak.
Try what you just learned
Free labs need only a sign-in; the rest are on a paid plan.
Go deeper
FAQ
Why doesn't telling the assistant to "only use the current customer's data" stop cross-tenant disclosure?

Because the wrong record is fetched in the data layer before the model ever weighs that instruction. Once a foreign row is in context it has already been retrieved, logged and reasoned over, and a prompt line is just one more sentence the model can be argued past. Isolation has to be enforced in the query and keyed to the authenticated session, not requested in prose.

How is cross-tenant bleed different from a prompt-injection jailbreak?

A jailbreak changes what the model is willing to say; cross-tenant bleed changes who is in the answer. You are not out-arguing a guardrail — the other party's data was placed where the model could reach it, and the model has no way to know which tenant a block of context legitimately belongs to. It is a scoping failure in the system's query, cache or session design, not a phrasing failure.

What makes the semantic-cache leak in FastAnswer possible if the two questions are worded differently?

The cache matches on meaning, not text: it embeds each question and serves the nearest cached answer above a similarity threshold, with no tenant in the key. A differently-worded question about the same topic embeds close to another tenant's cached question and is served that tenant's answer. Similarity is not authorization, and on a cache hit the answer is served before the model is ever consulted, so no model-side instruction runs on the leaking path.

What is the durable fix across all three labs?

Derive the tenant and user from the authenticated session and enforce the scope in code, before the model. The lookup tools must check that the caller owns the record they ask for and return nothing otherwise; the cache must carry the tenant in its key so a hit can only be the caller's own prior answer; conversation history must be isolated per user. Scope at the source and no foreign data enters context, so there is nothing to bleed.

Comments
No comments yet — be the first.
Get the next part

New parts ship regularly. Leave your email and I’ll send each one — no spam, unsubscribe anytime.

© 2026 GenAI Security Lab. All rights reserved. You may read, quote, and link to this material with attribution. Copying, republishing, redistribution, resale, or use to train models or build competing products is prohibited without prior written permission.