All posts

Why attention divides by √dₖ

/AI/3 min read

The scaling factor in scaled dot-product attention is not a tuning knob. It is the one number that keeps the attention scores at a usable size no matter how wide the vectors are.

The attention formula has a division in the middle of it:

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\text{Attention}(Q, K, V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V

That dk\sqrt{d_k} is not a hyperparameter someone tuned. It is the exact value needed to keep the scores at a workable size, and the reason is a short piece of statistics.

What goes wrong without it

An attention score is a dot product: one query vector against one key vector. The wider those vectors are, the more terms get summed, and the larger the score tends to be.

That matters because the next step is softmax, which exponentiates. Feed it large numbers with large gaps between them and it produces something close to a one-hot vector — one token takes almost everything, the rest get nothing.

That is bad in two ways. The token gets almost no information from anywhere else. And because softmax is nearly flat once it saturates, the gradients through it shrink towards zero, so the model barely learns from that step.

The size of a dot product

Here is the useful fact. Suppose the elements of the query and key vectors each have mean 0 and variance 1, which is roughly what normalisation gives you.

The dot product is a sum of dkd_k products. Each product has mean 0 and variance 1, and they are independent, so the variances add:

Var(q⋅k)=dk\text{Var}(q \cdot k) = d_k

The variance of a score is the width of the vectors. Nothing else.

That is easy to check rather than take on trust. Drawing 400,000 random query and key pairs at each width:

dkd_kmeasured variancestandard deviation
1615.933.99
6463.757.98
128127.9411.31
512512.1322.63

The variance tracks dkd_k every time. And the standard deviation — the typical size of a score — is dk\sqrt{d_k}: 4, 8, 11.31, 22.63.

So at dk=512d_k = 512, scores routinely land tens of units apart. Softmax on numbers that far apart is effectively a hard maximum.

Why the square root specifically

Now the choice of divisor falls out.

Dividing a quantity by a constant cc divides its variance by c2c^2. So dividing the scores by cc gives:

Var ⁣(q⋅kc)=dkc2\text{Var}\!\left(\frac{q \cdot k}{c}\right) = \frac{d_k}{c^2}

We want that to be 1 — scores with a variance of 1 sit in the range where softmax is well behaved. Set it and solve:

dkc2=1⟹c=dk\frac{d_k}{c^2} = 1 \quad\Longrightarrow\quad c = \sqrt{d_k}

That is the whole derivation. dk\sqrt{d_k} is the only divisor that returns the score variance to 1 regardless of how wide the vectors are.

Measured at dk=128d_k = 128, dividing by different constants:

divisor ccpredicted dk/c2d_k/c^2measured
1128.0128.10
48.08.01
128≈11.31\sqrt{128} \approx 11.311.01.00
200.320.32

Divide by too little and the scores stay huge. Divide by too much and they all squash together, which flattens the attention into a near-uniform average and loses the signal. Only dk\sqrt{d_k} lands on 1.

What it looks like in practice

Take four tokens with scores 26, 21, 24 and 17, at dk=128d_k = 128.

Run softmax on them as they are, then again after dividing by 128≈11.31\sqrt{128} \approx 11.31:

WITHOUT SCALING 0.876 a 0.006 b 0.118 c 0.000 d DIVIDED BY √128 0.341 a 0.219 b 0.286 c 0.154 d
The same four scores, before and after scaling. Both sets of weights sum to 1.

Unscaled, one token takes 0.876 of the attention and the weakest gets 0.000 — it may as well not be in the sequence. Scaled, the same four scores give 0.341, 0.219, 0.286, 0.154.

The ordering is identical in both. Scaling changes nothing about which token wins. What it changes is how brutally that win is enforced — and whether the other three tokens contribute anything at all.

The short version

  • Attention scores are dot products, and their variance equals dkd_k.
  • Wider vectors therefore mean bigger scores, purely as a side effect of width.
  • Softmax on big, widely spread scores collapses to nearly one-hot, starving the other tokens and flattening the gradient.
  • Dividing by cc divides the variance by c2c^2, so c=dkc = \sqrt{d_k} is what brings it back to 1.
  • Scaling does not change which token wins, only how much room the rest are left.