All posts

Self attention

/AI/4 min read

The word self means the queries, keys and values all come from the same sequence. The consequence is that a word's representation is built from its neighbours, so the same word comes out different in different sentences.

Every token in a sequence looks at every other token, decides which ones matter to it, and rebuilds itself as a blend of what it found.

The self is the important word. The things being looked at are the other tokens in the same sequence — not a separate input, not a stored document. A sentence is examined against itself.

Why that is necessary

A word's embedding is fixed. charge gets the same vector every time, whatever the sentence, because the lookup table does not know what sentence it is in.

But charge in a courtroom and charge in a battery are different words wearing the same spelling. Something has to make the representation depend on the surroundings, and self attention is that something.

Three roles per token

Each token produces three vectors, all from its own embedding through three learned matrices:

Query — what this token is looking for. Key — what this token offers to anyone looking. Value — what it actually contributes if chosen.

A token's query is compared against every token's key, including its own. The comparison is a dot product, divided by the square root of the vector width to keep the numbers in a sane range, and passed through softmax so the results are weights that sum to 1. The output is those weights applied to the values.

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V

All three come from the same sequence. That single fact is what makes it self attention.

The same word, twice

Take four dimensions and pretend they mean something: legal, electrical, generic, physical. Give charge an ambiguous embedding, sitting equally in the legal and electrical directions:

charge = [0.5, 0.5, 0.6, 0.2]

Now run it in two sentences.

SAME INPUT VECTOR, DIFFERENT OUTPUT …dropped by the judge .76 .18 L E G O …went flat in the battery .18 .76 L E G O L legal · E electrical · G generic · O physical — the output vector for "charge"
Nothing about the word changed. Its context did, and the output followed.
legalelectricalgenericphysical
…dropped by the judge0.7580.1790.4390.168
…went flat in the battery0.1750.7620.4370.365

The legal and electrical dimensions swap almost exactly. The cosine similarity between the two outputs is 0.59, from an input that was byte-identical in both cases.

That is the whole mechanism. The token did not become legal or electrical by itself — it absorbed a share of its neighbours' values, and the neighbours were different.

Notice too that the attention weights were nearly flat, around a third each. Attention did not need to pick a winner. The output changed because what was there to average over changed, which is a quieter and more common way for it to work than the dramatic one-token-attends-strongly-to-another picture suggests.

Why it replaced what came before

Everything happens at once. Previous sequence models read left to right, each step waiting on the last. Self attention computes all positions in parallel, which is what makes training on large data practical.

Distance stops mattering. A token at position 400 reaches a token at position 3 directly, in one operation. In a recurrent model that information had to survive 397 sequential updates, and it usually did not.

Words stop having one meaning. As above.

Several at once

One set of Q, K and V matrices learns one way of relating tokens. Running several in parallel, each with its own matrices, lets different heads pick up different relationships — one tracking grammatical agreement, another nearby modifiers, another something with no clean name.

Each head produces its own output, the outputs are concatenated, and a final matrix mixes them back to the expected width.

The one variant that matters

In a model that generates text, a token must not see what comes after it — during training the whole sentence is present, so without intervention position 3 could read position 5 while learning to predict position 4.

The fix is to set those scores to negative infinity before the softmax, so their weights come out as exactly zero. Same mechanism, with the upper triangle switched off.

Encoders skip the mask and let every token see the whole sequence, because there is nothing to predict in order.

The short version

  • Every token attends to every token in the same sequence, itself included.
  • Self means queries, keys and values all come from that one sequence.
  • Each token compares its query against all keys, softmaxes the scores, and blends the values.
  • Because the blend is over its neighbours, a fixed embedding produces different output in different sentences — 0.76 legal in one, 0.76 electrical in the other.
  • Weights are often fairly flat; what changes is what there is to average.
  • All positions compute in parallel, and distance costs nothing.
  • Multiple heads learn different relationships and are concatenated.
  • Masking future positions turns it into the version generation needs.