The problem
Every prototype called a model provider directly, with its own key and retry logic. Nobody knew what a feature cost.
How it works
- 01Request
- 02Route
- 03Cache
- 04Call + fallback
- 05Meter cost
What was hard
- Switching providers mid-stream without repeating text the user already saw.
- Cache keys include the prompt version, so a prompt change is not hidden by old answers.
- Hitting the budget moves to a cheaper model instead of failing.
The goal
Cost and speed per feature in one place, and an outage at one provider slows things down instead of breaking them.
This is a design I worked out on paper — the problem, the approach and the trade-offs. It is not a shipped product.