Continuous batching
/AI/4 min read
Waiting for every request in a batch to finish leaves most of the GPU's slots idle most of the time. Refilling each slot the moment it empties is a scheduling change, and it roughly doubles what the same hardware serves.
A GPU serving one request at a time is mostly idle. The arithmetic for a single token uses a fraction of what the hardware can do, so requests are processed in batches — several at once, at almost the same cost as one.
The question is when a batch begins and ends, and the obvious answer is the expensive one.
The problem with fixed batches
Take a batch of four. Run all four together, one decode step at a time, until every one of them is done. Then start the next four.
The trouble is that requests are not the same length. One reply runs to fifteen tokens and another to two hundred and forty. The short one finishes at step fifteen and its slot then sits empty for two hundred and twenty-five steps, because the batch does not end until the longest request does.
Meanwhile new requests wait in a queue for a slot that is empty and unusable.
Refilling as you go
Continuous batching changes one thing: when a request finishes, its slot is given to the next queued request immediately, in the middle of the batch.
Nothing about the model changes. It is purely a decision about scheduling — the batch stops being a group of requests that start and end together, and becomes a set of slots that are kept full.
What that is worth
Twelve requests, four slots, and reply lengths of 240, 30, 90, 15, 60, 45, 20, 180, 25, 35, 120 and 40 decode steps. That is 900 steps of actual work.
| static | continuous | |
|---|---|---|
| steps to clear everything | 540 | 255 |
| at 50 ms a step | 27.0 s | 12.8 s |
| slot utilisation | 41.7% | 88.2% |
| throughput | 33 tokens/s | 71 tokens/s |
Same requests, same model, same hardware. 2.12× the throughput, from changing when a slot is refilled.
The utilisation figure is the honest way to see it. Under fixed batching the slots do useful work 42% of the time; the rest is waiting for stragglers. Continuous batching takes that to 88%, and the remaining gap is the tail — near the end there is not enough queued work left to keep four slots busy.
What it does to individual requests
Throughput is the headline, but the effect on waiting is larger and less obvious:
| reply length | finishes at, static | finishes at, continuous |
|---|---|---|
| 20 steps | 260 | 95 |
| 45 steps | 285 | 75 |
| 60 steps | 300 | 75 |
| 25 steps | 445 | 115 |
| 240 steps | 240 | 240 |
The short requests gain most. A twenty-step reply that was queued behind a batch finished at step 260 under fixed batching and at 95 under continuous — and almost all of that wait was queuing, not generating.
The longest request is unchanged. Nothing sped it up; it was never the one waiting.
What it does not fix
It is scheduling only. The model, the attention, the arithmetic per token are all identical — this is a change to when work is admitted, nothing more.
The batch size is still capped by memory. Every request in flight holds a KV cache proportional to its length, so the number of slots is limited by how much cache fits, and keeping slots full means holding more caches at once. Continuous batching increases memory pressure precisely because it succeeds at keeping the machine busy.
Which is why it is normally paired with a scheme that stops the cache being allocated in wasteful contiguous blocks. Better scheduling exposes the memory limit; the memory work is what raises it.
The short version
- Requests are batched because a GPU running one at a time is mostly idle.
- With fixed batches, a slot stays empty from when its request finishes until the longest one in the batch does.
- Continuous batching refills each slot the moment it empties.
- On twelve mixed-length requests across four slots: 540 steps to 255, and 42% slot utilisation to 88%.
- Short requests benefit most; the longest request is unaffected.
- It changes scheduling only — no change to the model.
- Keeping slots full raises memory pressure, which is why it pairs with better KV cache allocation.