All posts

Sliding Window Attention

/AI/4 min read

Limiting each token to a 4,096-position window cuts the attention cost 32× at 131k tokens — and cuts the memory a sequence holds from 17.18 GB to a fixed 537 MB, which is 3 concurrent requests becoming 122.

In ordinary attention every position looks at every earlier position. A sequence of n tokens produces n² pairs of comparisons, which is fine at a thousand tokens and ruinous at a hundred thousand.

Sliding window attention limits each position to a fixed number of recent ones. A window of 4,096 means position 90,000 looks at positions 85,905 through 90,000 and nothing before that.

What that saves in arithmetic

sequence lengthfull attentionwindowed at 4,096ratio
8,1920.07B pairs0.03B2×
32,7681.07B0.13B8×
131,07217.18B0.54B32×

The cost stops being quadratic and becomes linear — n × w instead of n² — so the ratio is exactly n / w and grows without bound.

What it saves in memory, which matters more

The arithmetic is the advertised benefit. The memory is the one that decides what you can deploy.

With full attention, every generated token adds a set of keys and values that must be kept for the rest of the sequence. At roughly 128 KB per token for a mid-sized model:

sequence lengthfull attentionwindowedratio
8,1921.07 GB537 MB2×
32,7684.29 GB537 MB8×
131,07217.18 GB537 MB32×

Look at the windowed column. It does not change. Past 4,096 tokens, positions falling out of the window are discarded, so the cache reaches a fixed size and stays there — a sequence of any length costs 537 MB.

Turn that into concurrency. On 66 GB of free GPU memory at a 131,072-token context:

concurrent sequences
full attention3
4,096-token window122

Three requests becoming a hundred and twenty-two. That is not an efficiency improvement; it is the difference between a demo and a service.

THE WINDOWED CACHE STOPS GROWING 17.2 GB 8.6 GB 0 full attention — 17.18 GB flat from 4,096 onward — 537 MB, whatever the length 0 65,536 131,072 tokens on 66 GB: 3 concurrent sequences becomes 122 because positions leaving the window are discarded rather than kept the arithmetic saving is 32×; the deployment saving is the same number
The flat line is the whole argument. Length stops being a memory variable.

Can it still see far?

Yes, indirectly, and the reach is exactly computable.

Layer 1 lets a token see 4,096 positions back. Layer 2's input at that position already summarises 4,096 before it, so layer 2 reaches 8,192. Stacking multiplies:

receptive field = layers × window

A 32-layer model with a 4,096 window can in principle propagate information across 131,072 tokens — which is exactly the context length in the tables above, and not a coincidence.

But "in principle" is doing work. Information travelling k hops has been mixed with everything else at every hop. If each hop preserves a fraction of one specific signal:

hopssignal remaining, r = 0.75r = 0.6
175.0%60.0%
256.3%36.0%
431.6%13.0%
810.0%1.7%
161.0%0.03%

Eight hops — 32,768 tokens away — arrives at 10% strength in the optimistic case and 1.7% in the pessimistic one. The path exists; it is not the same as looking directly.

So the honest statement is: a windowed model has a long reach and a short grip. Something 100,000 tokens back is reachable and will not be attended to sharply.

Which is why it is usually mixed

Few production models use a window everywhere. The common arrangement is to give most layers a window and a few layers full attention.

The full-attention layers provide the direct long-range paths — exact retrieval of a distant fact, with no hop attenuation. The windowed layers provide the cheap local mixing that most of the work actually needs. Since only a minority of layers keep a full cache, the memory stays close to the windowed figure.

That is a better design than either extreme, and it follows from the attenuation table: you need some direct paths, and you do not need thirty-two of them.

What to take away

Two numbers, and they are the same number.

Windowing cuts attention cost by n / w, and it cuts the cache a sequence holds by the same factor — except the cache saving is better than it looks, because the windowed cache is bounded rather than merely smaller. Length stops being a variable in the memory budget.

What you give up is sharpness at distance, which decays with the number of hops rather than the number of tokens. Which is why a handful of full-attention layers is the arrangement that actually ships.