‹ All projects
DesignLLM testing

Prompt regression tests

Golden test cases that run in CI, so a prompt or model change cannot quietly break a right answer.

The problem

LLM output changes wording every run, so normal “equals” tests fail — and teams end up with no tests at all.

How it works

  1. 01Golden set
  2. 02Run
  3. 03Check
  4. 04Compare
  5. 05Gate

What was hard

  • Check facts and structure, not exact wording.
  • Pass on a percentage of cases, with some cases that must always pass.
  • Expected answers written by a person, never copied from a run.

The goal

A prompt change that hurts accuracy fails the build, with the exact cases it broke.

This is a design I worked out on paper — the problem, the approach and the trade-offs. It is not a shipped product.

Built with

TestNGGolden datasetsCI/CD