Why attention divides by √dₖ
/AI/3 min read
The scaling factor in scaled dot-product attention is not a tuning knob. It is the one number that keeps the attention scores at a usable size no matter how wide the vectors are.
The attention formula has a division in the middle of it:
That is not a hyperparameter someone tuned. It is the exact value needed to keep the scores at a workable size, and the reason is a short piece of statistics.
What goes wrong without it
An attention score is a dot product: one query vector against one key vector. The wider those vectors are, the more terms get summed, and the larger the score tends to be.
That matters because the next step is softmax, which exponentiates. Feed it large numbers with large gaps between them and it produces something close to a one-hot vector — one token takes almost everything, the rest get nothing.
That is bad in two ways. The token gets almost no information from anywhere else. And because softmax is nearly flat once it saturates, the gradients through it shrink towards zero, so the model barely learns from that step.
The size of a dot product
Here is the useful fact. Suppose the elements of the query and key vectors each have mean 0 and variance 1, which is roughly what normalisation gives you.
The dot product is a sum of products. Each product has mean 0 and variance 1, and they are independent, so the variances add:
The variance of a score is the width of the vectors. Nothing else.
That is easy to check rather than take on trust. Drawing 400,000 random query and key pairs at each width:
| measured variance | standard deviation | |
|---|---|---|
| 16 | 15.93 | 3.99 |
| 64 | 63.75 | 7.98 |
| 128 | 127.94 | 11.31 |
| 512 | 512.13 | 22.63 |
The variance tracks every time. And the standard deviation — the typical size of a score — is : 4, 8, 11.31, 22.63.
So at , scores routinely land tens of units apart. Softmax on numbers that far apart is effectively a hard maximum.
Why the square root specifically
Now the choice of divisor falls out.
Dividing a quantity by a constant divides its variance by . So dividing the scores by gives:
We want that to be 1 — scores with a variance of 1 sit in the range where softmax is well behaved. Set it and solve:
That is the whole derivation. is the only divisor that returns the score variance to 1 regardless of how wide the vectors are.
Measured at , dividing by different constants:
| divisor | predicted | measured |
|---|---|---|
| 1 | 128.0 | 128.10 |
| 4 | 8.0 | 8.01 |
| 1.0 | 1.00 | |
| 20 | 0.32 | 0.32 |
Divide by too little and the scores stay huge. Divide by too much and they all squash together, which flattens the attention into a near-uniform average and loses the signal. Only lands on 1.
What it looks like in practice
Take four tokens with scores 26, 21, 24 and 17, at .
Run softmax on them as they are, then again after dividing by :
Unscaled, one token takes 0.876 of the attention and the weakest gets 0.000 — it may as well not be in the sequence. Scaled, the same four scores give 0.341, 0.219, 0.286, 0.154.
The ordering is identical in both. Scaling changes nothing about which token wins. What it changes is how brutally that win is enforced — and whether the other three tokens contribute anything at all.
The short version
- Attention scores are dot products, and their variance equals .
- Wider vectors therefore mean bigger scores, purely as a side effect of width.
- Softmax on big, widely spread scores collapses to nearly one-hot, starving the other tokens and flattening the gradient.
- Dividing by divides the variance by , so is what brings it back to 1.
- Scaling does not change which token wins, only how much room the rest are left.