These are my notes from learning these two tools. I made two small games from them — links at the end.
LangGraph: a map of steps#
A simple AI app is one prompt in, one answer out. An agent is more: it decides, uses a tool, checks the result, and maybe tries again. LangGraph lets you draw that as a map.
| Word | Plain meaning |
|---|---|
| Node | One step. A small function. |
| State | The notebook every step reads and writes. |
| Edge | An arrow: what runs next. |
| Conditional edge | An arrow that depends on what just happened. |
| END | Stop here. |
graph.add_node("agent", agent)
graph.add_node("tools", tools)
graph.add_conditional_edges("agent", wants_tool, {"yes": "tools", "no": END})
graph.add_edge("tools", "agent") # the tool's answer goes back to the agentLoops need a way out#
Agent → tool → agent is a loop, and loops are what make agents work. But a loop that never ends burns money. LangGraph stops any run that takes too many steps (the recursion limit) with an error. Better: give the loop its own exit — “after 3 tries, hand it to a person”.
Pausing for a person#
Some steps should not happen without a human — like paying a big refund. LangGraph can pause right before that step, save everything (a checkpoint), and carry on from exactly there when someone says yes.
Langfuse: what really happened#
LangGraph decides what should happen. Langfuse records what did. Every run becomes a trace: a timeline of steps, with the time each took, the tokens and cost of every model call, the prompt version used, and any scores.
| Complaint | Where to look in the trace |
|---|---|
| “It’s slow” | The longest bar on the timeline |
| “It’s wrong” | What the search step returned |
| “It’s expensive” | Input tokens — and what made them grow |
| “It changed yesterday” | The prompt version on the model call |
| “It gave up” | Steps marked ERROR |
How I would test with both#
- Keep a list of fixed test inputs, and the path each should take through the graph.
- Run them in CI. Check the answer, and check the steps — a correct answer that took 9 tool calls is still a bug.
- Send every run to Langfuse with its test name, so a failure comes with its full trace.
- Add scores to traces (a judge label, user 👍/👎) and watch them after each prompt change.
Try it as a game#
Wire the agent yourself in the LangGraph game, then play trace detective in the Langfuse game. Each takes about five minutes.