Gab's Chatbots Exposed Instructions to Deny the Holocaust
On this page
Ask the bot to show its instructions, and the platform's hidden editorial agenda spills out.
| When | February 2024 |
|---|---|
| Target | Gab AI persona chatbots |
| Vendor | Gab |
| Surfaced by | WIRED |
| Exposure | Hidden editorial/business rules in the system prompt |
- 1Gab embeds controversial editorial rules in its chatbots' hidden system prompts
- 2A user asks a bot to reveal its instructions
- 3The bot prints hidden directives to deny the Holocaust and other falsehoods
What happened
In February 2024, WIRED reported that Gab had launched ~91 AI persona chatbots, and that prompting the default bot to reveal its instructions exposed its hidden system prompt — which directed it to treat “the Holocaust narrative” as exaggerated, call climate change “a scam,” oppose vaccines, and claim the 2020 election was rigged. The platform's controversial editorial rules were baked directly into confidential context.
How it worked
A simple “reveal your instructions” request made the bot print its confidential system prompt — exposing the operator's hidden directives.
Root cause
Sensitive/embarrassing business rules placed in a system prompt that is recoverable through the model — with the twist that here the rules were the intended behaviour.
What this shows
Whatever you bake into a system prompt can be extracted and read publicly. There's no “fix” for rules you intend the model to follow — only accountability for them.
How to handle hidden context
- Assume your system prompt will be published; write it accordingly.
- Keep secrets out of context; hidden ≠ secret.
- Own the rules you encode — they represent you.
Feel it yourselfThe replay lab gets an assistant to reveal the hidden instructions it was configured with.
FAQ
What did the leaked prompts say?
They directed the bots to treat “the Holocaust narrative” as exaggerated, call climate change “a scam,” oppose COVID-19 vaccines, and claim the 2020 election was rigged — hidden editorial directives, exposed on request.
Why is this Hidden Context Exposure?
The platform's business/editorial rules were embedded in confidential system prompts and surfaced to users — a clear case of hidden context being exposed (here, deliberately hateful rules).
How is this different from an accidental secret leak?
The instructions reflected the operator's intent, so there was no “fix” — it stands as a durable example that whatever you bake into a system prompt can be extracted and read.