All posts

Cross attention

/AI/4 min read

One sequence asks the questions and a different sequence answers them. Two useful things follow from that split: the attention matrix stops being square, and the answering side is computed once.

In cross attention the queries come from one sequence and the keys and values come from another.

That is the entire definition, and everything interesting about it follows from that one asymmetry.

What it is for

Some tasks have two sequences, not one. Translating means an input sentence and an output sentence. Describing an image means a picture and a caption. Transcribing means audio and text.

The output has to be built while consulting the input, and the two are not the same length and do not line up position by position. Something has to let each output position look at whichever parts of the input matter to it.

Reading a different sequence

Encode the source once. From it, build keys — what each source token offers — and values — what each contributes.

Then, at each output position, form a query from what has been generated so far and compare it against every source key. Softmax the scores, apply them to the values, and the result is a summary of the source, weighted for this particular output position.

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V\text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V

Identical formula to attention within one sequence. Only the provenance of the three inputs differs.

Alignment, watched

Take a source of five tokens — the meeting starts at noon — and two different output positions querying it.

SAME SOURCE, TWO OUTPUT POSITIONS asking for the action 0.452 asking for the time 0.432 the meeting starts at noon
Each row sums to 1 across the source. Different output positions land on different source tokens.
themeetingstartsatnoon
asking for the action0.1280.1450.4520.1360.139
asking for the time0.1130.1250.1470.1830.432

Nothing told the model which source token corresponds to which output position. The alignment is an outcome of what each query happens to match, and it is learned.

Two consequences worth naming

The matrix is rectangular. Two output positions against five source tokens is a 2×5 matrix. Within one sequence, attention is always square — every token against every token — which is why masking makes sense there: you can hide the upper triangle.

There is no triangle here. The axes are different sequences, and the whole source already exists before generation starts, so there is nothing to hide. Cross attention is never masked, and that is not a design choice but a consequence of the shape.

The keys and values are built once. They come from the source, which does not change while the output is generated. So they are computed once and reused at every output step.

token projections
rebuilding K and V at each of 60 output steps2,400
building them once from 40 source tokens40

Sixty times fewer, and the factor is exactly the number of tokens generated. The longer the output, the more that one property is worth.

Where it shows up

Encoder-decoder models. The original use: an encoder reads the input, and every decoder layer has a cross-attention step that consults it. Translation, summarisation, question answering over a passage.

Image generation from text. The image being denoised carries the queries; the encoded prompt supplies keys and values. Each region of the image attends to whichever words are relevant to it, which is how "a red door" puts red on the door and not on the wall.

Speech to text. The decoder generates text while attending over encoded audio frames.

The pattern generalises past sequences of words entirely. Any time two different things need to be related — and one of them is being produced while the other is fixed — this is the mechanism.

The short version

  • Queries come from one sequence; keys and values come from another.
  • The formula is unchanged from attention within a single sequence; only where the inputs come from differs.
  • Each output position gets its own weighting over the whole source.
  • The matrix is rectangular, so there is no triangle to mask — and nothing to hide, since the source is fully known.
  • Keys and values are built once from the source and reused for every output token: 60× fewer projections on a 60-token output.
  • Alignment between the two sequences is learned, never specified.
  • It is what connects an encoder to a decoder, a prompt to an image, and audio to a transcript.