All posts

The agent loop

/AI/4 min read

The loop is about twenty lines of ordinary code. Its cost is not ordinary — because the whole history is re-sent every turn, tokens grow with the square of the turn count.

The loop is the part of an agent that is not a model. It calls the model, executes whatever the model asked for, appends the result, and calls again.

Written out it is unremarkable:

messages = [system_prompt, goal]
 
for step in range(MAX_STEPS):
    reply = model(messages, tools=TOOLS)
    messages.append(reply)
 
    if reply.final_answer:
        return reply.final_answer
 
    for call in reply.tool_calls:
        messages.append(run(call))
 
raise StepLimitReached()

One thing in there is worth stating plainly: the model never runs a tool. It emits a request, and the loop executes it. Everything the agent can actually do is defined by what that loop is willing to run.

The cost is quadratic

A model keeps nothing between calls, so the loop keeps the history and re-sends all of it every turn. That is what makes an agent coherent across steps — and it means the cost does not grow the way people assume.

If each turn adds about 800 tokens, turn 10 is not sending 800 tokens. It is sending 8,000.

turnsif it were linearactually sentratiocost
54,00012,0003.0×$0.030
108,00044,0005.5×$0.110
2016,000168,00010.5×$0.420
4032,000656,00020.5×$1.640
6048,0001,464,00030.5×$3.660
WHAT IS ACTUALLY SENT linear guess actual 10.5x here 0 20 60 turns tokens
Turn n sends n times the per-turn size, so the total goes as N².

The total is T⋅N(N+1)/2T \cdot N(N+1)/2. Doubling the turn limit roughly quadruples the bill.

This is the number that makes a step limit an economic decision rather than a safety afterthought. Raising MAX_STEPS from 20 to 40 sounds like giving the agent twice as long. It is closer to four times the cost.

Trimming the history

The fix is not to re-send everything. Keep the recent turns in full and compress the older ones:

turnsfull historyrecent 5 kept, older summarisedsaved
1044,00033,80023%
20168,00086,40049%
40656,000227,60065%
601,464,000416,80072%

The saving grows with the length of the run, because it is the old turns — the ones there are most of — that get compressed.

What you lose is detail from earlier steps, and it matters which detail. The goal must survive intact, along with anything the agent might still need to refer back to. Summarising away a constraint the user gave at the start produces an agent that confidently violates it at turn 30.

Stopping

Two exits, and both are required.

The model says it is finished. It returns an answer rather than a tool call. This is the intended one.

The loop gives up. A step counter runs out. This is not a failure mode to be engineered away — it is the only exit that is guaranteed to happen, because the first one depends on the model's judgement.

Production loops usually add two more, for the same reason: a time budget and a cost budget. A step limit does not bound spend if a single step can retrieve a very large document.

What goes wrong

It never stops. Tool after tool, no answer. The step limit is the backstop; detecting a repeated identical call and telling the model it already tried that is the graceful version.

It repeats itself. The same call with the same arguments, getting the same result, expecting something different. Worth detecting explicitly, because it is invisible in aggregate metrics — the run completes, it just costs ten times what it should have.

It overflows. Given the quadratic growth above, this is not an edge case on a long run; it is the default outcome. Trimming has to be built in, not added when it first breaks.

It stops too early. Answers from partial information because nothing defined what finished means. That belongs in the instructions as a checkable condition.

The short version

  • The loop is ordinary code: call the model, run what it asked for, append the result, repeat.
  • The model never executes anything — the loop does, which is where all the real control lives.
  • Because the whole history is re-sent each turn, cost grows with the square of the turn count.
  • Twenty turns sends 10.5× what a linear estimate suggests; doubling the limit quadruples the bill.
  • Keeping the last few turns in full and compressing the rest saves 65% on a forty-turn run.
  • Two exits are needed: the model finishing, and the loop giving up. Only the second is guaranteed.
  • Add time and cost budgets, since a step limit does not bound spend.