LangGraph
/AI/4 min read
The point is not that graphs are more flexible than lines. It is that the graph's position is stored alongside its data, so a run can be paused, resumed after a crash, or held for a human — none of which a straight-line pipeline can do.
LangGraph describes an application as a graph: nodes that do work, edges that say what runs next, and a state object that every node reads from and writes to.
The usual explanation is that graphs allow branching and loops where a straight pipeline does not. True, and not the interesting part. The interesting part is that the run's position is data.
State, and what is in it
A node is a function from state to a partial update:
def classify(state):
return {"category": pick_category(state["message"])}It receives the whole state, returns only the fields it changed, and the framework merges them. Nodes never call each other; they only leave things in the state for whatever runs next.
Edges decide what that is. A plain edge always goes to the same node. A conditional edge is a function that inspects the state and returns the name of the next node — which is how a routing decision becomes part of the structure rather than something buried in a node's body.
Because an edge can point backwards, the graph can contain cycles. An agent loop — call the model, run a tool, call the model again — is a cycle between two nodes, and a pipeline cannot express it at all: a pipeline is a line, and a line has no way back.
Who calls the tool
Worth stating plainly, because it is the most common misunderstanding: the model does not execute anything.
A node asks the model what to do. The model replies with a request — this tool, these arguments. A different node runs it and writes the result into the state. Then the loop goes round and the model sees the result as text.
The graph makes this separation structural rather than conventional. The tool-executing node is a visible thing you can put a guard in front of, which matters as soon as the tools do anything real.
The part that only works because position is data
Run state and current node are both persisted. A checkpointer writes them after every step, under a thread identifier.
This one property is what the framework is actually for.
A crash resumes where it stopped. Suppose each step fails 2% of the time — a timeout, a rate limit, a bad parse. A workflow with no saved position has to start over, and the longer it is the worse that gets:
| steps | finishes with no retry | steps executed, restart from scratch | checkpointed |
|---|---|---|---|
| 9 | 83.4% | 10.0 | 9.2 |
| 25 | 60.3% | 32.9 | 25.5 |
| 60 | 29.8% | 118.3 | 61.2 |
| 120 | 8.9% | 516.8 | 122.4 |
At nine steps, restarting costs 9% extra — barely worth the machinery. At 120 steps it costs 4.2× the work, because a run that has to be perfect end to end almost never is. A 2% step failure rate means a 120-step workflow completes cleanly 8.9% of the time.
The saving is not linear in length; it accelerates. Which tells you exactly when this framework earns its complexity, and when it does not.
A run can pause for a person. This is the feature that is simply impossible without stored position. Stop before the node that sends the email, write the state down, return. Hours later a human approves, the state is loaded, and execution continues from that node with everything intact.
Nothing was held in memory across that gap. The process that resumes need not be the process that paused — which is what makes approval workflows deployable rather than a demo.
Conversations get memory for free. Reusing a thread identifier loads the previous state, so prior turns are already there. Memory is not a separate component; it is the same persistence read back.
When not to reach for it
A single model call does not need a graph. Neither does a fixed three-step pipeline that either works or fails as a unit — a straight chain expresses that, with less to understand.
The cost of the graph is real: every piece of shared data has to be named in a state schema, control flow lives in edge functions rather than in the reading order of a file, and debugging means inspecting state snapshots instead of a stack trace.
That cost buys branching, cycles, and the ability to survive interruption. If a workflow has none of those needs, it is paying for nothing.
What to take away
Every framework of this kind is really answering one question: where does the run's progress live?
A pipeline keeps it on the call stack, which means the run cannot outlive the process. LangGraph moves it into storage, and everything that makes the framework worth using — resuming, pausing for a human, retrying just the failed step, remembering a conversation — is a consequence of that single move.