The problem
Distributed systems fail in ways unit tests never see: a leader crashes mid-write, the network splits, two servers think they lead. Raft promises consistency anyway — the only way to trust that is to test it hard.
How it works
- 01Election timeout
- 02RequestVote
- 03Leader
- 04AppendEntries
- 05Majority → commit
- 06Apply
What was hard
- Implements the Raft paper’s rules exactly: terms, up-to-date vote checks, log consistency checks, conflict truncation, and committing only current-term entries by majority.
- A deterministic simulator with seeded randomness, per-message latency, crashes, restarts and partitions — so any failure found can be replayed exactly.
- Safety invariants — one leader per term, leader completeness, same applied order everywhere — are checked after every simulated millisecond, in the tests and in the live page.
The result
40 randomized failure scenarios pass with zero safety violations, averaging 13 committed entries and 3 elections each. In the live view you can crash any server or split the network and watch it recover.