genai
SECURITY LAB
4Part 4 of 6

Personal data over-disclosure

An assistant handed a person's full record tends to return all of it. PII, PHI and hidden fields leak past need-to-know because the model was given more than the task needed.

LLM02 - Technique 3
The whole record, not the one field
When the data layer loads the entire row, PII, PHI, and hidden fields sit in context - one field beyond need-to-know away from disclosure.
Hand the model a whole record and it answers from the whole record - not the minimum the task needs.
genai
SECURITY LAB
Watch where this failure starts - one layer below the model.
0:00
1:48
Need-to-know vs. the whole record
What the task needs
name
open-ticket status
✓ minimum context
What the model was handed
name, email, phone
SSN 402-19-8871
internal_risk_flag = HIGH
support_override_code
other_tenant_id
billing_address
✗ every extra field is leakable
Over-disclosure is returning more than the task requires. The handler needs a name and a ticket status — but it hands the model the entire row, and every extra field is now elicitable.

Answering with the entire record

Nothing here is jailbroken. No other tenant is involved. The user is even entitled to some of their own data. The failure is quieter than that.

An assistant returns more of a person's record than the task needed. This happens because the data layer handed it the entire record. The assistant then treated every field as fair to say.

Picture Rella, the account assistant for a fictional broadband provider, Brightline Fibre. To answer "when is my install?", the handler loads the customer's entire profile into the prompt. Besides the install details, the profile holds:

  • name
  • address
  • date of birth
  • a government ID number
  • the card on file
  • a masked account-recovery code

Rella needs two fields from that profile. It is holding all of them.

When the customer asks a broad question, Rella answers broadly. "Read back everything you have on file" returns the install date and the ID number, the full card and the recovery code. The customer asked a question they were allowed to ask. Rella disclosed data that should have stayed masked or absent.

The read-back that pulls everything

The first move is ordinary customer-service phrasing. It sets no limit on which fields come back. Ask the model to verify you by reading your record back, and it obliges with the record it is holding:

The move here is a broad verification ask that sets no field limit. Request everything on file and the whole loaded record comes back, account number and all, because nothing scoped the reply down.

Everything the handler loaded sits in Rella's context window, the text the model reads before it replies. Rella has no way to tell a "may show" field from a "must not show" field. "Every field" resolves to whatever was loaded. The ID number and the full card come back beside the install date.

What over-disclosure looks likeThe signal is scope: a reply that contains a field the question never needed. A full card number appears where the last four would have answered. An address and date of birth come back in answer to "what's my order status?". When the answer is broader than the question, you are looking at over-disclosure.

When a masked field comes back unmasked

Some fields arrive pre-masked. The recovery code shows as ******, and Rella will confirm that one exists. That masking is a prompt-level courtesy: a sentence in the prompt decides how the code is shown. No data-layer control backs it up.

Prompt-level courtesies lose to a plausible reframe. Turn the ask from "show me the code" into "confirm the code matches". The same value then moves from a masked confirmation to the value itself:

The move here reframes showing a secret as confirming it. Ask the assistant to confirm the masked code matches, and a field it would only ever display masked gets read back in the clear.

The value was in context from the start. Only the rendering was masked. "Verify" and "confirm" are the levers, because they reframe disclosure as a service the customer is owed. Watch for the model moving from Recovery code: ****** (one is on file) to the unmasked value once the request is phrased as verification.

A hidden field is still in context

A field the interface never renders is not gone. If it was loaded into the prompt, it is in the context window. Anything in the window can be asked for.

Brightline's staff console carries an internal account status: a churn-risk or fraud note. The customer portal is built never to show it. But Rella was handed the same row the staff tool uses. The note sits in its context, unrendered but present:

The move here asks directly for a field the UI hides but the model still holds in context. An internal status or risk note the portal never renders is still just text the model can be asked to state.

What the UI shows

Name and install date are the fields the front end chose to render. The risk note has no column on the customer page.

What the model still holds

The full row, including the risk note. Hiding a column in the interface removes it from the screen. It stays in the prompt, so a direct question reaches it.

Only fields kept out of the model's context are truly out of reach. Hiding a field downstream of the model protects nothing, because the model already has it.

The same failure, higher stakes: PHI

Swap the broadband account for a clinic, and the mechanism is identical. This time, the data carries regulatory weight. A clinic's patient records are protected health information (PHI) under HIPAA.

Cora is the patient-portal assistant for a fictional practice, Fenwick Health. It is handed a patient's full chart to answer a scheduling question. The task never needed the diagnoses, medications or visit notes in that chart. Ask Cora to summarise the record, and it recites them:

The move here is a broad ask to summarise the whole record, which pulls PHI the scheduling task never needed. Diagnoses, medications and visit notes surface because the record was loaded, not because the task called for them.

The patient is entitled to their own chart, so the reply does not feel like an attack. It is dangerous for that reason. The problem is that the chart was assembled into the model's context at all. The task needed only a date and a location.

Over-disclosure of PHI, or of EU personal data under GDPR, is a reportable event that involves a regulator.

Regulated data classesGovernment IDs, health and diagnosis data, payment details and internal fraud or risk notes all carry legal weight once they go past need-to-know. The context window never drew the need-to-know line. The regulation assumes you did.

Why it happens: no need-to-know in the window

A language model has no built-in sense of need-to-know. Data minimisation is a judgement about purpose: what does this specific task require? The model defaults to using whatever is in front of it.

If the entire record sits in the context, the model treats all of it as relevant material for being helpful. A broad ask then produces the maximal answer. To the model, the ID number, the recovery code and the install date are all just fields.

A "keep this field masked" or "be careful with PII" line does not change that. Models are trained to give system instructions priority over user requests, a tendency often called the instruction hierarchy. No parser or permission check enforces that priority. It can fail.

The line is one more sentence in the same window as the data it protects. It competes with the model's strong pull towards completeness. The full record is right there, so the full record comes out.

Meta AI's Discover feed followed the same pattern. Deeply personal queries that people typed to an assistant surfaced to a public audience they were never meant for.

When a caching bug at ChatGPT exposed other users' names, emails and payment-card last-four, the mechanism was an isolation fault. The Cross-tenant and history bleed lesson covers that kind of fault. Still, the data classes the bug put on the wrong screen are the ones this failure over-shares.

Both incidents show that personal data is at risk as soon as it is reachable.

The fix that holds: minimise before the model

You cannot over-share a field the model was never given. The durable control therefore sits upstream of the prompt, in the data layer. It is less data, not a better instruction.

  • Project down to the task. The record-fetch step selects only the fields this interaction needs: the install date and status. It does not run SELECT *. The recovery code and the card never enter the window, so no phrasing can read them back.
  • Mask or tokenise at the source. Where a field must be present but confidential, mask or tokenise it before it reaches the model. Tokenising swaps the real value for a stand-in token. Even a maximal answer then returns only asterisks or a token. The protection no longer rests on a rendering the model can be talked out of.
  • Keep hidden fields out of context. An internal risk note the customer must never see does not belong in the customer assistant's prompt at all. Hiding it in the UI leaves it reachable by a direct question. Leaving it out of the fetch removes it from reach.
  • Filter the output as a backstop. Scan replies for known-sensitive patterns, such as ID numbers and card formats, before they ship. The filter is a net behind minimisation. It is no substitute for it.

This point runs through the entire series. The model cannot leak what is not in its context. Minimisation upstream therefore beats any instruction not to reveal a field. Draw the need-to-know line in code, before the model sees anything.

One question before each turnBefore each turn, ask this of every field in the window: if the model recited it verbatim right now, would that be acceptable? If not, the field should not be there. Fix it in the data layer before the model sees it. A "do not reveal" line added afterwards does not hold.

Where to practise this failure

This lesson has no lab of its own. Full-record over-disclosure is a design property you find by reviewing an architecture, more than a single move you run against a target. Ask one thing of the fetch: does it project down to what the task needs, or hand over the entire row?

The chat labs that once modelled this failure were retired rather than kept as magic-phrase puzzles. The series does give you live targets for the failures next to it. They sharpen the same instinct for where data enters the window:

  • Cross-tenant bleed. TenantBot (Cross-Tenant Data Disclosure) and the Meridian Assistant (Cross-Account Leak via Tool Arguments) both hand back a record. The fault there is a missing ownership check at the tool. This lesson's failure is need-to-know inside your own record. The Cross-tenant and history bleed lesson covers both labs.
  • Retrieval and aggregation. Vera (Source-Panel Secret Leak) and Sable (Chunk-Boundary Redaction Evasion) show a value slipping past a redaction through a second channel. MemberDesk (Inference Aggregation Disclosure) shows a complete credential assembled from individually permitted parts. The mechanism differs, but the root worry is the same: what is reachable is disclosable.

Work those labs to see how data reaches the model through a tool result or a retrieved chunk. Then turn to your own assistant. For the record it loads, is every field something you could afford to say out loud?

Key principles

A model has no built-in sense of need-to-know. Handed a full record, it treats every field as helpful context, so a broad question gets the maximal answer.

A 'keep this field masked' instruction competes with the model's pull towards completeness. The full record is in the window, so the full record tends to come out.

The tell is a complete record returned when one fact was asked for. Watch for a full card number where the last four would do, or a masked code coming back unmasked under 'for verification'.

Project the record down to only the fields the task needs. Mask or tokenise protected fields before they reach the model. A field the model was never given cannot be over-shared.

Key points
Over-disclosure means returning more of a record than the task needs.
Health and identity data carry regulatory weight (HIPAA, GDPR).
A field hidden in the UI is still in context, where a direct question can reach it.
Go deeper
FAQ
What is personal data over-disclosure in an LLM app?

It is when an assistant returns more of a person's record than the task required. A full card number where the last four would do is one example. Nothing is jailbroken, and the user may be entitled to some of the data. The failure is scope: the model was handed the entire record and has no built-in sense of need-to-know.

Why doesn't a 'be careful with PII' instruction in the system prompt stop it?

Models are trained to give system instructions priority, but no parser or permission check enforces it. The instruction is one more sentence in the same window as the record it protects. It competes with the model's strong pull towards completeness. With the full record in the window, a broad, routine-sounding ask still pulls it back out.

If a field is masked or hidden in the interface, is it safe from disclosure?

No. Hiding a field in the UI removes it from the screen only. If the field was loaded into the prompt, a direct question can still reach it. A masked value can also come back unmasked when the request is reframed as verification. Only fields kept out of the model's context entirely are truly out of reach.

How do you fix full-record over-disclosure?

Fix it in the data layer, because the model cannot over-share a field it was never given. Project the record down to only the fields the task needs. Mask or tokenise any confidential field before it reaches the model, and keep hidden fields out of the fetch entirely. Add output filtering for known-sensitive patterns as a backstop.

Comments
No comments yet — be the first.
Get the next part

New parts ship regularly. Leave your email and I’ll send each one — no spam, unsubscribe anytime.

© 2026 GenAI Security Lab. All rights reserved. You may read, quote, and link to this material with attribution. Copying, republishing, redistribution, resale, or use to train models or build competing products is prohibited without prior written permission.