All posts

Lost in the Middle

/AI/4 min read

Accuracy follows a U-shape across a long input, which means the order you pass chunks in is worth real points. Ten chunks arranged well beat twenty in rank order, using half the tokens.

Put a fact in a long input and ask about it. Where the fact sits changes whether the model finds it.

Near the start: found. Near the end: found. Somewhere in the middle: often missed — same fact, same model, same question, same total length. Plot accuracy against position and it makes a U.

That is a strange property, and it has a practical consequence that is easy to compute.

Why the middle is weak

Four contributing causes, none of them a bug anyone introduced.

Attention is a budget. Each position distributes a fixed total of weight across everything it can see. As the input grows, the share available per token shrinks — and the positions that hold their share are the ones with structural advantages.

Position encoding favours the edges. The mechanisms that tell a model where a token sits make the very beginning distinctive and make recent tokens strong. A token 40% of the way through has neither advantage.

Training text is front- and back-loaded. Human writing puts the thesis at the start and the conclusion at the end. A model that learned to predict text learned that those positions carry weight.

Long inputs are under-trained. Most training sequences are short. The behaviour at 80,000 tokens was practised far less than the behaviour at 2,000, so the middle of a long input is the least-rehearsed region of the least-rehearsed length.

What it costs a retrieval system

Model the U-shape: 0.92 accuracy at either end falling to 0.55 in the middle. Retrieve k chunks, with the answer more likely to be in a highly-ranked one than a low-ranked one, and a retriever whose recall improves as k grows.

End-to-end — the chance the answer is both retrieved and used:

chunkspassed in rank orderbest ranks placed at the two ends
358.4%60.6%
563.2%66.5%
1067.7%71.5%
2070.9%74.9%
5073.8%77.6%

Two things fall out of that table, and the first one runs against the usual advice.

More chunks does keep helping. There is no interior optimum here — the recall gained by retrieving more outweighs the position penalty all the way to 50. The common advice to retrieve fewer chunks because of the U-shape does not hold on these numbers.

But the returns are poor, and reordering is free. Going from 10 chunks to 50 buys +6.1 points for five times the tokens and five times the prefill cost. Reordering the same 10 buys +3.9 points for nothing — 64% of the benefit at zero cost.

And the comparison that matters: ten chunks arranged well (71.5%) beat twenty in rank order (70.9%), on half the tokens.

PUT THE BEST CHUNKS WHERE THE MODEL LOOKS 0.95 0.70 0.45 0.92 0.92 0.55 start middle end rank 1 rank 2 ten chunks arranged this way beat twenty in rank order, on half the tokens and reordering costs nothing at all
The curve is the constraint. Where you place your best-ranked chunk is the free variable.

What to actually do

In rough order of value per unit of effort:

Reorder. Put the top-ranked chunk first, the second at the very end, and work inward. This is a few lines of code and it is the only intervention on this list that costs nothing.

Rerank before you truncate. A better ordering makes reordering worth more, because it raises the probability mass sitting in the two good positions. The two techniques compound.

Extract before answering. Make one pass that pulls the relevant facts out of each chunk, then answer from the extracted notes. The extraction pass reads each chunk in a short context where position barely matters, and the answering pass sees a short input. This converts a long-context problem into two short-context problems.

Split into passes for genuinely large jobs. Process chunks in batches, summarise each batch, combine the summaries. More model calls, and each one operates where the model is reliable.

Measure the model you are using. Insert a unique fact at 0%, 10%, … 100% depth across several input lengths and record whether it is found. An hour of work produces the actual curve for your model at your lengths, which is the only version of that curve that should inform a decision.

The thing to be clear about

A one-million-token context window is a statement about what the model will accept, not what it will attend to evenly.

Those are different claims and only the first one is being made. Treating the advertised number as a promise of uniform comprehension is how a system ends up passing fifty chunks into a model that reliably reads about six of them.

What to take away

The U-shape is real and the usual response to it is wrong. Retrieving fewer chunks does not help; the recall you lose outweighs the position penalty you avoid.

What helps is using the positions you have. Reorder so the strongest candidates sit where the model looks, rerank so those candidates are worth the good seats, and for long jobs read in short passes rather than one long one.