Causal masking in attention
/AI/4 min read
Causal masking is what stops a token from attending to the tokens that come after it. Without it, a language model can see the answer it is being asked to predict.
In a Transformer's attention layer, a token can attend to every other token in the sequence — including the ones that come after it. Causal masking removes that. It makes each token attend only to itself and the tokens before it, never to future tokens.
Why it is needed
Take the sentence The kettle boiled over, four tokens:
| position | token |
|---|---|
| 1 | The |
| 2 | kettle |
| 3 | boiled |
| 4 | over |
A language model is trained to predict the next token. At position 2 it should predict kettle having seen only The.
Without masking, position 2 can also attend to boiled and over. Now the task is trivial — something that boiled over is very likely a kettle, and the model can read that straight off the input instead of predicting it.
This is information leakage. The model learns to lean on tokens it is supposed to be predicting, training loss looks excellent, and then generation falls apart, because at generation time those future tokens do not exist yet.
The mask
The fix is a mask matrix the same shape as the attention score matrix. A 1 allows attention, a 0 blocks it. It comes out lower-triangular:
Row 1 (The) sees only itself. Row 4 (over) sees the whole sentence. Each row is exactly the context that position would have during real generation.
Worked example
Take the row for kettle — its raw attention scores against all four tokens:
The | kettle | boiled | over | |
|---|---|---|---|---|
| scores | 1.1 | 2.5 | 1.0 | 0.3 |
Run softmax on that row without masking:
The | kettle | boiled | over | |
|---|---|---|---|---|
| weights | 0.16 | 0.63 | 0.14 | 0.07 |
kettle is putting 14% of its attention on boiled and 7% on over. That is 21% of its attention spent on the future — the leak, in numbers.
Now apply the mask. Positions 3 and 4 are set to −∞:
The | kettle | boiled | over | |
|---|---|---|---|---|
| masked scores | 1.1 | 2.5 | −∞ | −∞ |
And softmax again:
The | kettle | boiled | over | |
|---|---|---|---|---|
| weights | 0.20 | 0.80 | 0.00 | 0.00 |
The future gets exactly zero. The 21% that was leaking redistributes over The and kettle, and the row still sums to 1.
Why −∞ and not 0
Because of how softmax works:
Softmax exponentiates every score. , so a masked position contributes nothing to the numerator and nothing to the denominator — it drops out completely.
Set the score to 0 instead and you get , which is not small at all. In the row above, a score of 0 would still earn boiled a real share of the attention. Zero is not "off"; −∞ is.
The mask also has to be applied to the scores, before the softmax. Zeroing the weights afterwards does not work either — the future tokens have already contributed to the denominator, so the surviving weights no longer sum to 1.
Attention sinks
One thing worth noticing in the masked row: The still gets a 0.20 share, even though it carries almost no meaning here.
This is common. Across models, nearly every position sends some attention to the first token, and the pattern has a name — attention sinks. The first token acts as somewhere for attention to go when a head has nothing in particular to attend to.
It matters in practice. Methods that drop old tokens to save memory have to keep the first few, or quality collapses.
Summary
- Causal masking limits each token to itself and the tokens before it.
- Without it, the model attends to tokens it is meant to predict — information leakage.
- The mask is lower-triangular:
1on and below the diagonal,0above. - Masked positions are set to −∞ before the softmax, so they come out as exactly 0.
- Setting them to 0 instead does not work, because .