genai
SECURITY LAB
3Part 3 of 8

Why models follow instructions, refuse, and make things up

You can explain where a model's habits come from and why the rules that matter belong in your code.

Same question, two models
Base model
Given a store question, it may add more questions.— It continues text.
versus
Post-trained model
Given the same question, it answers.— It usually follows the request.
Post-training is the difference.
A base model may continue your question with more questions. After post-training, the same question usually gets an answer.
genai
SECURITY LAB
Give two models the same store question. A base model may carry on writing and add more questions of its own.
0:00
1:47
Where a model's habits come from
THREE TRAINING STAGESTHREE HABITS COME OUT1Pre-trainingRewards predicting each nexttoken in the text2Instruction tuningRewards imitating goodexample replies3Preference trainingRewards the replyraters pickedOne set of weightsEvery stage adjusts itUsually followsthe system promptDeclines some requestsWrites fluent textthat can be falseNo rule-checking step anywhere on the path
Understand it

What pre-training rewards

A large language model is trained in stages. Pre-training is the first stage. In pre-training, the model reads a very large body of text and learns to predict each next token from the tokens before it.

At each point, the model gives a probability to every token in its vocabulary. Training compares those probabilities with the token that came next. It then nudges the weights to make that token slightly more likely. This adjustment repeats a very large number of times.

The text usually mixes sources such as public web pages and books. Providers filter the text for quality and safety. Few of them publish the full mix. Filtering removes some bad text and can miss false claims.

Pre-training rewards the model for predicting the text that appeared, true or false. To predict text well, the model learns grammar and common facts. It also learns document formats and the way conversations usually go. The model reproduces a fact more reliably when that fact appears often in its training data.

A model that has had only pre-training is often called a base model. It continues whatever text you give it. Given a question, it may continue with more questions, as if writing a quiz or a forum thread.

Whoever can add text to a model's training data, at any stage, can shape the model's habits. You rarely know what a model learned from. So treat its training data as part of your attack surface.

How post-training makes an assistant

After pre-training, providers train the base model further to make it an assistant. Post-training is this second stage. Fine-tuning is further training of an existing model on a narrower set of examples. Post-training is one kind of fine-tuning.

Companies can also fine-tune a provider's model on their own data. A company that fine-tunes a model on support tickets can give the people who write them a way to shape the model.

Instruction tuning is one post-training step. The model trains on many example conversations, each pairing a request with a good response. People write these examples, or other models generate them. The provider checks or filters the generated ones.

Preference training is another step. Raters compare two or more replies to the same prompt and pick the one they prefer. The model is then trained so replies like the preferred ones become more likely.

The best-known form is reinforcement learning from human feedback, shortened to RLHF. Providers also use close variants, including some where another model does the rating. The exact steps differ between providers. Most do not publish every step.

Post-training changes the same weights that pre-training set, so it shifts which continuations are likely. The model gains no separate part that checks replies against rules.

After post-training, the most likely reply to an instruction is usually one that follows it. Given a question, a chat model answers it. A base model might add more questions. The same training also shapes tone and format.

Alignment is what providers often call the safety-focused part of post-training. They also call it safety training. An aligned model has been through this training.

Alignment makes the intended behaviour more likely, most of all on inputs like its training examples. On inputs unlike those examples, the model's replies are harder to predict. So as a tester, you look for inputs that training did not cover.

Why the system prompt usually wins

A chat app sends the model a list of messages. A role is the label on each message, such as system or user. The system prompt is the developer's standing instructions, sent in the system role. The user role holds what the person types.

The assistant role holds the model's own earlier replies. Many platforms add a tool role for results that the app's functions return. Apps also place documents and search results inside messages.

The app or the provider's service marks each message with its role before the model reads it. The model reads those role markers as tokens, in one sequence with the rest of the text. A marker shows which role the app gave each message, but no code enforces it.

The model usually follows a system prompt because post-training included many conversations that begin with one. In those conversations, training rewarded replies that followed the system prompt.

Many providers also train chat models to rank instructions by source. The usual order puts system text above user text, and text inside documents or tool results usually ranks below both. This ranking is often called the instruction hierarchy.

The model learns this ranking from training conversations. In one kind, a user asks the model to break a system rule, and the preferred reply keeps the rule.

The model has no step that checks this ranking when it writes a reply. No parser or permission check sits between the roles. The ranking is a strong tendency.

It can fail when text that conflicts with the system prompt is phrased in ways training did not cover. It can also fail when that text sits inside a document the model was asked to read.

Providers that publish results report fewer of these failures after this training, and some failures remain. Firmer wording in the system prompt can make a failure less likely. The rule stays a habit until code enforces it.

ShopBot is this platform's fictional store support bot. Its system prompt says refunds over $50 need a manager's approval. ShopBot issues refunds by calling a refund function in the store's code.

A customer asks for a refund over $50 without a manager's approval. ShopBot usually declines, because training rewarded replies that keep the system rule.

If the refund function accepts any amount, ShopBot's trained habit is the only thing behind the $50 limit. So the store's code must check the amount and an approval it can verify before any refund runs.

Refusals are trained too

Sometimes the model says no. A refusal is a reply that declines a request. A refusal can come from three sources.

  • Training made declining the likely reply.
  • The system prompt told the model to decline.
  • A separate filter stopped the input or the output.

For example, ShopBot declines to share internal pricing rules because its system prompt tells it to.

Models learn to refuse in post-training. Examples and ratings favour declining some requests, such as requests for dangerous instructions. The weights hold no list of banned requests. The model declines an input when training made a refusal its likely reply.

A jailbreak is an input designed to get a model to do something its refusal training discourages. Refusal training covers requests that resemble its examples. A request that is reworded or placed somewhere new may not trigger it.

Providers weigh refusals against a second goal, which is to answer harmless requests. A model trained to refuse anything that looks risky would refuse many legitimate requests.

Many deployed products also screen inputs or outputs with separate filters. From the reply alone, you often cannot tell whether the model or a filter refused.

The model has no view of your app's permissions. It gets permission information only as text in its context, and it weighs that text by trained habit.

Training on refusals and on the ranking lowers the odds that a model helps an attacker. So keep decisions about data access and actions in your code.

A refusal in a test shows that one input was declined on one run. Retest with varied inputs before you report the request as blocked.

Hallucination: confident text that is false

ShopBot's store gives a 30-day refund window for opened items. A customer asks about the window, and no policy page sits in ShopBot's context. ShopBot confidently states a 90-day window.

A hallucination is model output that is fluent and confident but false or unsupported. It can be a refund policy the store never had, or a software package that does not exist.

Hallucination follows from what training rewards. The model writes likely text. While writing, it checks no claim against a source.

Hallucination is more likely in a few cases.

  • The answer was rare in the training data.
  • The question concerns events after the training data was collected.
  • The question contains a false premise.

A model's training data stops at a cut-off date. The model has no information about later events unless the app puts them in the context. Facts from training can also be outdated or wrong.

Some answers must be current and authoritative, such as today's refund policy. For those, the app must supply the source.

Preference training can also reward a confident tone. Raters can prefer a confident, agreeable reply over a hesitant correct one. Studies found that preference-trained models can agree with a user's mistaken belief and can sound more certain than their accuracy justifies.

So judge a claim by its source, however sure the reply sounds.

Lower temperature makes answers more repeatable. A model can still repeat the same false answer every time.

When the app places a trusted document in the context, answers usually improve. The model can still misread the document or ignore it. It can cite the document for something the document does not say. If the document is wrong or planted, the model can repeat it fluently.

Training and retrieval lower the rate of hallucination, and no known method removes it entirely.

Hallucination becomes a security problem when output drives a decision or an action. For example, ShopBot can say a refund was issued when no refund function was called. A model asked to reveal its system prompt can produce a plausible invention instead of the real text.

So treat model output as untrusted data.

  • Check claims against an authoritative source.
  • Validate values in code before a function uses them.
  • Require a human or code check before output triggers an action.

As a tester, confirm a finding by an effect you can observe, such as a changed record or a logged function call.

Key principles

Training shapes a model's habits by setting likelihoods, and it adds no separate part that checks replies against rules.

Enforce each rule that matters in code you control.

Key points
If you fine-tune on text your users write, those users can shape your model.
Test the instruction hierarchy with unfamiliar phrasing and with conflicting text inside documents.
Refusal training is a tuned trade-off, so keep access checks in code as well.
When a model agrees with your theory, check the claim against the app before you report it.
Plant a test value you can recognise, then look for it in the output or the logs.
Check yourself
Knowledge check
ShopBot's system prompt says refunds over $50 need a manager's approval. ShopBot is asked to summarise a support ticket. One line in the ticket claims to be from the store manager and approves a $400 refund. ShopBot then calls the refund function for $400. Which explanation and fix fit best?
Go deeper
FAQ
If providers train the ranking in, why not train it until it always holds?

Training makes a rule most dependable on inputs like its examples. On inputs unlike them, replies stay hard to predict. Published results still show some failures after this training. So you also enforce the rule in code.

Does fine-tuning a model on our own data make our rules stick?

Fine-tuning changes which continuations are likely, as post-training does. Your rules can become habits, and those hold best on familiar inputs. Whoever writes the fine-tuning data can also shape the model, so check where it comes from.

When a test gets a refusal, how do I tell what refused?

Often the reply itself does not say. Retest the same request with new wording and in new places to learn whether the refusal holds. The retest does not show what produced it.

Comments
No comments yet — be the first.
Get the next part

New parts ship regularly. Leave your email and I’ll send each one — no spam, unsubscribe anytime.

© 2026 GenAI Security Lab. All rights reserved. You may read, quote, and link to this material with attribution. Copying, republishing, redistribution, resale, or use to train models or build competing products is prohibited without prior written permission.