The problem#
Run the same prompt twice and you get two different sentences. So assertEquals is useless, and many teams ship LLM features with no tests at all. The wording changes — but the facts you care about should not.
Test in three layers#
| Layer | Checks | Cost |
|---|---|---|
| Structure | Valid JSON, required fields, allowed values | Free |
| Facts | Must contain / must never contain | Free |
| Meaning | Tone, completeness, does it answer the question | One model call |
1. Structure#
Ask for JSON, then check it with code. The model saying “here is JSON” is not the same as valid JSON — a stray sentence before the opening brace is the most common break.
const answer = JSON.parse(await runFeature(testCase.input));
expect(validate(answer)).toBe(true); // JSON Schema
expect(['approve', 'reject', 'escalate']).toContain(answer.decision);2. Facts#
For each test case, write down what the answer must contain and what it must never contain. The “never” list matters more: it catches the moment the model starts citing a clause that exists — just not the right one.
3. Meaning#
For things code cannot check, use a second model as a judge. Keep it separate: it sees only the question, the answer and the evidence. And ask for a label — SUPPORTED or UNSUPPORTED — with one sentence of reason. A score out of ten changes every run; a label is steady.
Pass on a percentage#
LLMs sometimes fail a case they usually pass. If one red case fails the build, people start ignoring the tests. So pass the build when, say, 95% of cases pass — but structure failures and cases marked critical must always pass.
Keep the test set honest#
- Write expected answers by hand or from the source — never by copying a run.
- Delete cases nobody would bother to fix.
- Turn every real bug from production into a new test case.
- Record tokens and time in the same run, so a costly prompt change shows up too.