All posts

Autoregressive models

/AI/4 min read

Generate one token, append it to what you have, and use that to generate the next. The whole design follows from one identity in probability, and so does its main weakness.

An autoregressive model predicts the next item in a sequence from everything before it, then treats its own prediction as part of the history and does it again.

Auto because it feeds on its own output. Regressive because each prediction is made from past values.

Where the design comes from

It is not an arbitrary choice. It falls out of an identity that is true of any joint probability:

P(x1,x2,…,xn)=P(x1) P(x2∣x1) P(x3∣x1,x2)⋯P(xn∣x1,…,xn−1)P(x_1, x_2, \ldots, x_n) = P(x_1)\,P(x_2 \mid x_1)\,P(x_3 \mid x_1, x_2)\cdots P(x_n \mid x_1, \ldots, x_{n-1})

The probability of a whole sequence is the product of the probability of each item given the ones before it.

That decomposition is the useful part. Modelling the probability of every possible sentence directly is hopeless. Modelling one conditional — given this prefix, what comes next — is a single manageable problem, and the identity says that solving it gives you the whole distribution for free.

So the training objective is next-token prediction, and it is not a simplification. It is the full thing, factorised.

Generating

THE PREFIX GROWS BY ONE EACH STEP rain fell 0.55 rain fell all 0.40 rain fell all night 0.72 rain fell all night <end> 0.68 each row is one full forward pass through the model
Four tokens, four passes. The right-hand column is the only new information each time.

Multiply the conditionals and you get the sequence probability, with P(rain)=0.12P(\text{rain}) = 0.12 to start:

0.12×0.55×0.40×0.72×0.68=0.012925440.12 \times 0.55 \times 0.40 \times 0.72 \times 0.68 = 0.01292544

Why nobody uses that number

A sequence probability is a product of numbers below 1, so it shrinks as the sequence grows — regardless of how good the sequence is.

per-token probability10 tokens50 tokens200 tokens
0.93.5e−15.2e−37.1e−10
0.72.8e−21.8e−81.1e−31
0.59.8e−48.9e−166.2e−61

A model that is 90% confident at every step still assigns a 200-token passage a probability of about 10−910^{-9}. Comparing that against a 10-token passage tells you which is shorter, not which is better.

So the useful quantity is the average log probability per token, which does not depend on length. For the sequence above that is −0.8697, and exponentiating its negative gives a perplexity of 2.39 — roughly, the model was as uncertain as if choosing between 2.4 equally likely options at each step.

The cost that cannot be removed

Look at the diagram again. Each row is a complete forward pass, and row four cannot start until row three has produced its token.

A 500-token answer is 500 sequential passes. Not 500 units of work that could be spread across hardware — 500 steps that must happen in order, because each one's input includes the previous one's output.

This is why generation feels slow in a way that is unrelated to how much compute you have. Adding GPUs does not shorten a chain of dependencies.

There is a related cost that can be removed. Recomputing the attention keys and values for the entire prefix at every step is enormous waste — the prefix has not changed, only grown by one. Storing them and computing only for the new token turns quadratic redundant work into linear.

Training on whole sentences

At generation time the future does not exist, so nothing can leak. Training is different: whole sentences are present at once, which is what makes training parallel and fast.

That creates a problem. If position three can attend to position five while learning to predict position four, it has read the answer.

The fix is to mask: before the softmax, set the attention scores from each position to every later position to negative infinity, so their weights come out as zero. Every position then sees only what precedes it, exactly as it will at generation time, while all positions are still processed in one pass.

What the trade buys

Coherence. Every token is chosen with the entire preceding text available. Nothing is committed to before its context exists.

A simple objective. Next-token prediction on ordinary text. No labels, no special construction.

Any length. There is no fixed output size — the model emits a stop token when it is finished.

Against that: it is sequential, latency grows linearly with output length, and an early mistake conditions everything after it. The model cannot revise a token it has already emitted; it can only continue from it.

The alternative is producing many positions at once, which is faster and gives up the guarantee that each token saw the ones before it. Output tends to be less coherent — repetitions and contradictions across positions that never conditioned on each other.

The short version

  • Predict the next item from all previous ones, append, repeat.
  • It follows from the chain rule: any joint probability factorises into conditionals.
  • That is why next-token prediction is the whole objective and not an approximation.
  • Sequence probabilities shrink with length, so per-token log probability and perplexity are what get reported.
  • Generation is a chain of dependent forward passes — more hardware does not shorten it.
  • Caching keys and values removes the redundant recomputation, which is the part that is waste.
  • Causal masking makes parallel training behave like sequential generation.
  • The trade is coherence and simplicity against sequential speed and unrevisable mistakes.