All posts

Plan-and-execute agents

/AI/4 min read

Instead of deciding one step at a time, the agent writes the whole plan first and then works through it. Cheaper and steadier when the path is knowable, and useless when it is not.

Some agents decide what to do next after every result. A plan-and-execute agent decides everything up front: it writes a numbered plan, then works down it.

Reconsidering only happens when something comes back other than what the plan assumed.

The parts

A planner reads the task and produces the steps. This is the expensive call and the one that determines whether anything else works.

An executor runs the steps one at a time. Each step is narrow — call this tool with these inputs — so this is usually a smaller, cheaper model than the planner.

Tools, memory holding the task and the plan and everything produced so far, and a re-planner that looks at what actually happened and decides whether the remaining plan still makes sense.

The re-planner is technically optional and practically essential. Without it the agent will follow a plan that stopped being correct three steps ago.

The shape

PLAN ONCE, THEN WORK DOWN IT 1 · read the lockfile 2 · look up advisories 3 · find fixed versions 4 · write the table step 1 · done step 2 · done step 3 · running step 4 only when something surprises it
The plan is written once. The dashed path is the exception, not the loop.

One running

List every dependency with a known vulnerability, and the version that fixes it.

The planner writes four steps: read the lockfile, look up advisories for each package, find the fixed version for anything that matches, and write the table.

stepwhat happened
1read the lockfile — 214 packages. Also found a second lockfile in another workspace
—re-plan: the original plan assumed one lockfile. Steps rewritten to cover both
1bread the second lockfile — 96 more packages, 310 in total
2advisory lookup across all 310 — 4 matches
3fixed versions found for 3; the fourth has no released fix
4table written, with the unfixed one flagged

The re-plan is the interesting row. Nothing failed at step 1 — it ran, it returned a valid result, and it returned more than the plan expected. A planner that never revisits its plan would have finished confidently and reported on 214 of 310 packages, which is a wrong answer that looks exactly like a right one.

Step 3 is worth noting too: it found no fix for one package and said so, rather than treating a missing result as an error.

Against deciding step by step

The alternative is to think again after every observation. The two are not ranked; they suit different tasks.

one step at a timeplan first
when it thinksevery turnonce, then on surprises
model callsone per step, all on the big modelone big call plus cheap ones
adaptingimmediateonly at a re-plan
long tasksdrifts, and costs grow with the transcriptsteadier and cheaper
open-ended tasksgoodpoor — the plan is guesswork
visibilityyou see each thought as it happensyou see the whole intent up front

Planning first is cheaper because the expensive reasoning happens once. It is steadier because the goal is written down rather than reconstructed from a growing transcript at every turn.

It fails when the path genuinely cannot be known in advance. If step two depends on what step one returns, a plan written before step one ran is a guess.

There is also a real benefit that has nothing to do with the model: a plan is legible. You can read the four steps before any of them run, and stop it if they are wrong. An agent that decides as it goes only shows you the decision after it has acted on it.

What goes wrong

The plan is bad. Everything downstream inherits it, and the executor will faithfully carry out a wrong plan. Give the planner the real tool list and demand concrete steps — a step reading "gather the relevant data" cannot be executed, only reinterpreted.

The executor drifts. It does something defensible that was not what the step meant. Steps should name the tool and the inputs, not describe an intention.

It never re-plans. A step returns something unexpected and the agent carries on regardless. This is the failure in the trace above, and it produces confident wrong answers rather than visible errors.

It over-plans. Thirty steps for a three-step job. Ask for the shortest plan that does the work.

It re-plans forever. Each rewrite produces a slightly different plan and no progress. Cap the re-plans, and stop when two consecutive rewrites come out substantially the same.

The short version

  • The planner writes all the steps up front; the executor works down them.
  • The executor can be a smaller model, because each step is narrow.
  • A re-planner checks after each step whether the rest of the plan still holds.
  • Re-planning matters most when a step succeeds but returns something the plan did not expect.
  • Cheaper and steadier than re-deciding every turn, when the path is knowable.
  • Useless when each step depends on what the last one returned.
  • A written plan can be read and stopped before anything runs.