Multi-head attention
/AI/4 min read
Running several attention operations side by side costs nothing extra — the parameters and the cache are identical to one big head. What you actually spend is the width each head gets to work in.
One attention operation produces one weighting over the sequence. Every token ends up with a single opinion about which other tokens mattered to it.
That is limiting, because a token relates to its neighbours in several ways at once — grammatically to one, semantically to another, positionally to a third. One set of weights has to average all of that into a single answer.
Multi-head attention runs several attention operations in parallel, each with its own projections, so each can settle on a different notion of relevance.
How the split works
Take the model width and the number of heads . Each head works in a narrower space of dimensions.
Each head has its own , and , projecting the full-width input down into its own . Each runs ordinary attention in there and produces a -wide output. The outputs are concatenated back to full width, and a final matrix mixes them.
It costs nothing
This is the part that is usually stated as a footnote and is actually the point.
With :
| heads | Q, K, V, O parameters | KV cache per token per layer | |
|---|---|---|---|
| 1 | 512 | 1,048,576 | 1,024 |
| 8 | 64 | 1,048,576 | 1,024 |
| 32 | 16 | 1,048,576 | 1,024 |
| 64 | 8 | 1,048,576 | 1,024 |
Every row identical. The projections total , and since that is just , plus for the output matrix — regardless of .
The cache behaves the same way: per token per layer is , whatever the head count.
So "more heads" is not a cost-versus-capability trade in the way it sounds. It is free.
What it actually costs
Not memory. Width.
| heads | dimensions per head |
|---|---|
| 1 | 512 |
| 8 | 64 |
| 32 | 16 |
| 64 | 8 |
At 64 heads, each one has eight numbers in which to encode what it is looking for and what each token offers. Attention scores become dot products of 8-dimensional vectors, and there is only so much a head can distinguish with that.
That is the real trade: many narrow specialists against few broad generalists, at fixed total capacity. Too few heads and each is forced to average several kinds of relationship together. Too many and each is too thin to represent one properly.
The original design used 8. Current models use 32 or 64, because the model width grew alongside — a 64-head model at still gives each head 64 dimensions.
What the heads do
Nothing assigns them roles. Each head has its own randomly initialised projections, and whatever it ends up specialising in is a product of training.
What is observed afterwards is that heads do differentiate — some attend mostly to adjacent tokens, some track syntactic dependencies at a distance, some appear to do something with no clean description. Many are close to redundant and can be removed with little effect.
The specialisation is a side effect of having several separate parameter sets that all reduce the same loss. Nothing coordinates them, and nothing guarantees they will divide the work sensibly.
Where it sits
Every attention block in a transformer is multi-head. In an encoder, all heads attend across the whole input. In a decoder's self-attention, all heads are masked so nothing sees the future. Where a decoder consults an encoder, all heads attend across to it.
The head count is per block, so a 32-layer model with 32 heads runs 1,024 attention operations for one forward pass — all of them in parallel, which is why hardware likes this shape.
The short version
- Several attention operations run side by side, each with its own projections.
- The model width is divided among them: .
- Parameter count is whatever the head count — adding heads is free.
- The KV cache is likewise independent of .
- What changes is how many dimensions each head gets: 64 heads at width 512 leaves eight each.
- The trade is many narrow specialists against few broad generalists at fixed capacity.
- Roles are never assigned; specialisation emerges, and some heads end up doing very little.