All posts

Recursive language models

/AI/4 min read

Rather than reading a huge input, the model writes code that reads it — calling other models on the pieces and seeing only what comes back. What that buys is not fewer tokens.

There is a limit to how much text can usefully go into one model call. Not just the hard context limit — well before that, quality falls off. Material in the middle of a very long input gets used less reliably than material at the ends.

A recursive language model works around that by never putting the input in front of the main model at all.

The input becomes a variable

The full text sits in a code environment as a variable. The main model gets a short prompt saying that the variable exists, roughly what it contains, and that it has two things available: the ability to run code, and a function that calls another language model.

Then it writes code.

tickets = context.split("\n---\n")
 
findings = [
    call_llm(f"What is the underlying cause here? One line.\n\n{t}")
    for t in tickets
]
 
print(findings)

Each call_llm sends one ticket to a separate model instance and gets back a single line. The main model never sees a ticket. It sees four hundred one-line answers, decides whether that is enough, and either answers or writes more code.

That last part is what makes it recursive rather than a fixed pipeline: the model chooses how to split the input, and it chooses differently depending on what is being asked. A question about one specific customer would produce a filter before the fan-out. A question about overall themes would produce grouping.

What it actually buys

Four hundred support tickets, about 600 tokens each — 240,000 tokens of source.

one callfan out per tickettwo levels
largest single context240,000660660
what the root model sees240,00016,060860
total tokens processed240,000296,060298,860
model calls1401421
ONLY THE LEAVES TOUCH THE SOURCE root sees 860 tokens group group group 660 each the 240,000 tokens of source exist only at the bottom row
Each call reads a little. Nothing in the tree ever reads all of it.

The number that changed is the first row. No call has to handle more than 660 tokens — 364 times smaller than the original problem — and every one of them is operating in the range where models are reliable.

The number that did not improve is the third row. Total tokens went up, by about 23%, because each sub-call repeats a prompt and produces output. This is not a technique for processing fewer tokens. It is a technique for never processing many at once.

When it is cheaper, and when it is not

That distinction decides the economics, and the answer depends entirely on which model runs the leaves.

cost
one frontier call over the whole input$0.606
leaves on a small model, root on the frontier$0.067
every call on the frontier model$0.898

Nine times cheaper, or half again more expensive, from the same structure.

The saving does not come from recursion. It comes from the fan-out moving almost all the token volume onto a cheap model, and leaving only the small synthesis step on the expensive one. Run the whole tree on your best model and you pay a premium for the privilege of splitting the work up.

Which is the right way to think about it: recursion is what makes it possible to use a small model on a large problem, and that is where the money is.

Against retrieving instead

Search picks the relevant pieces before the model sees anything. Its weakness is that it has to decide relevance from the question alone, which fails when what matters cannot be recognised by resemblance — counting, comparing across everything, or finding what is absent.

A recursive model visits everything and decides as it goes, and can change strategy after seeing partial results. That costs far more, since it touches the whole input rather than a retrieved slice.

Roughly: search when the answer is in a few places, recurse when the answer is about all of it.

What it costs you

Latency. Four hundred calls fan out in parallel, but the levels are sequential, and the slowest leaf holds up its group.

Errors compound. A leaf that misreads its ticket returns a confident wrong line, and the root cannot tell — it never saw the source. Nothing downstream can catch it.

It is hard to debug. The model wrote the code that split the input. When an answer is wrong, the fault may be in a leaf, in the synthesis, or in a split that put related material in different groups. The transcript shows what happened without showing why.

It needs real infrastructure. A sandboxed environment that runs model-written code, with limits on time, calls and spend.

The short version

  • The input stays in a code environment; the main model gets the variable name, not the contents.
  • It writes code that slices the input and calls other models on the pieces.
  • Only short summaries come back, so the root's context stays small.
  • On 400 tickets: largest single context 240,000 → 660 tokens, a 364× reduction.
  • Total tokens go up about 23% — this does not reduce work, it redistributes it.
  • It is 9× cheaper only if the leaves run on a small model; all-frontier is more expensive than one big call.
  • Search when the answer sits in a few places; recurse when it is about the whole input.
  • The costs are latency, uncatchable leaf errors, and debugging model-written code.