The problem
Guardrail work is usually described, not shown. The hard part — catching an attack without blocking a normal question — is easier to play than to explain.
How it works
- 01Prompts fall
- 02Catch or let pass
- 03Score
- 04Precision + recall
What was hard
- Blocking a normal question is the costly mistake, so it loses points.
- Some attacks hide inside retrieved documents, not in what the user typed.
- The slower round scores both kinds of mistake, the way a real guardrail should be judged.
The result
A game that teaches the real trade-off in AI safety filters.