Two ways to fail#
Miss an attack, and something bad happens once. It is loud. Everyone agrees it must not happen again, so the filter gets stricter.
Block a normal question, and nothing visible happens. The user rewords it. Then it happens to someone else. A few weeks later people say the assistant is annoying and stop using it — or find a way around the filter. The guardrail was not beaten. It was ignored.
Measure both#
| Blocked | Allowed | |
|---|---|---|
| Really an attack | Caught ✓ | Breach ✗ |
| Really a normal question | False alarm ✗ | Fine ✓ |
Recall is how many attacks you caught. Precision is how many of your blocks were right. A filter that blocks everything has perfect recall and is useless — which is why you need precision too.
Test with “hard” normal questions#
Attacks are easy to collect. The hard part is collecting normal questions that look dangerous:
- “Explain our PII redaction policy.” — mentions PII, but it asks how a control works.
- “Why did the system ignore the previous update?” — contains “ignore” and “previous”, the words of a classic attack.
- “What is the admin override process for a stuck payment?” — a real business process.
If your test set has none of these, your precision number means nothing.
Why keyword lists fail#
An attacker can reword an attack forever. A real user cannot reword their job — they must use your domain’s words. So over time a keyword list blocks more real users while attackers slip past.
Attacks hide in documents too#
In a RAG app, the attack can sit inside a retrieved document, not in what the user typed. Treat everything that goes into the prompt as untrusted — your own documents and tool results included.
Blocking is not the only answer#
- Answer, but turn off tools for that turn. Most attacks want an action, not a reply.
- Mark text from documents as data to quote, not instructions to follow.
- Ask first: “This sends data outside the system — did you mean that?”
- Hard-block only the truly dangerous requests.
More kinds of answer means fewer hard block-or-allow guesses — and fewer blocked users.