LLM test gate
Test an LLM feature the way the post says: structure, facts, then meaning. Compare two prompts and see which one the gate lets ship.
How it works
- 1Exact match fails even correct answers, because the wording changes every run.
- 2Structure and facts are plain code checks — edit any answer and they re-run as you type.
- 3Meaning uses a judge model’s label. This page has no model, so those labels are written with the cases, and an edited answer is marked “not re-judged”.
- 4The build passes on a percentage, but structure checks and critical cases must always pass.
More demos
All demosQueryLite: a SQL database
A SQL database written from scratch — parser, B+ tree indexes, a query planner, joins and transactions — running live on 22,000 rows in your browser.
Raft consensus, live
The algorithm that keeps etcd and Kubernetes consistent. Five servers elect a leader and replicate a log — crash them and split the network while it runs.
Live collaborative editor
Real-time editing with no server. Three devices share a note — take one offline, edit everywhere, reconnect, and they merge. Open a second tab and it syncs live.