All posts

Mixture of experts in LLMs

/AI/4 min read

A mixture-of-experts layer holds many small networks and sends each token through only two of them. The model gets to be large without every token paying for all of it.

A mixture-of-experts layer holds several small networks side by side and sends each token through only two of them. The rest do nothing for that token.

Why it is needed

In an ordinary language model, every token passes through every parameter. A 30-billion-parameter model runs all 30 billion of them, whether the token is a comma or a rare technical term.

That is expensive, and a lot of it is wasted. Most tokens are easy and do not need the whole model.

Mixture of experts breaks the link between how large a model is and how much work each token costs. The model can hold a very large number of parameters while any single token touches only a slice of them.

What an expert is

An expert is a small feed-forward network — the same kind that already sits inside every transformer layer. That is all it is.

The name oversells it. Nobody hands an expert a subject. They all start out the same shape and the same blank, and whatever division of labour appears is one the model worked out during training. The splits tend to be mundane — punctuation, word endings, the shape of a token — rather than tidy topics like law or biology.

A layer usually holds 8, 16, 64 or 128 of them.

Where it sits

A transformer layer does two things: attention, then a feed-forward network. Mixture of experts swaps out the feed-forward network for a group of experts plus a router. Attention is left exactly as it was.

Every layer gets its own router, trained separately. A 24-layer model has 24 of them, and the same token can be sent to a different pair of experts at each one.

The router

The router is a small linear layer followed by a softmax. It reads the token's vector and gives one score per expert. The scores add up to 1.

Take a layer with 12 experts, and one token arriving at it:

ROUTER SCORES · 12 EXPERTS · TOP-2 KEPT 1 2 3 0.189 4 5 6 7 8 0.400 9 10 11 12
One token's router scores. Two experts are kept; the other ten never run.

The two highest scores are expert 9 at 0.400 and expert 4 at 0.189. Those two run. The other ten are skipped.

Those two scores do not add up to 1 by themselves, so they are rescaled until they do:

w9=0.4000.400+0.189=0.68w4=0.1890.400+0.189=0.32w_9 = \frac{0.400}{0.400 + 0.189} = 0.68 \qquad w_4 = \frac{0.189}{0.400 + 0.189} = 0.32

The layer's output is then those two experts' outputs, blended in that ratio:

output=0.68⋅E9(x)+0.32⋅E4(x)\text{output} = 0.68 \cdot E_9(x) + 0.32 \cdot E_4(x)

Ten of the twelve experts contributed nothing to this token. That is where the saving comes from.

Total parameters and active parameters

Because most experts sit out, an MoE model has two sizes worth quoting.

Total parameters is everything stored in the weights. Active parameters is what actually runs for one token.

Take a model with 24 layers, 12 experts in each, half a billion parameters per expert. Attention, embeddings and normalisation come to 9 billion, and those are shared — they are not copied per expert.

workingcount
experts, all of them24×12×0.5B24 \times 12 \times 0.5\text{B}144B
shared layersattention + embeddings + norms9B
total153B
experts that run, per token24×2×0.5B24 \times 2 \times 0.5\text{B}24B
shared layerssame, always run9B
active per token33B

So each token costs about what a 33-billion-parameter dense model would cost — 21.6% of the weights — while the model holds 153 billion.

Note which parts got replicated. Only the feed-forward half is copied across experts. Attention and the embeddings are shared, which is why adding more experts grows the total count without changing what one token pays.

Keeping the experts busy

Left alone, the router does something unhelpful. It picks a few experts early, those experts get more training and improve, so the router picks them even more. The rest are barely trained and end up as dead weight.

The fix is a second loss term added during training, alongside the usual one. It penalises uneven routing, pushing the router to spread tokens across all the experts over a batch rather than crowding a favourite few.

There is also a cap. Each expert is given a fixed number of tokens it can take in a batch, and once it is full the extra tokens are either dropped or sent to the next expert on the list.

What it costs

  • Memory. All 153 billion parameters have to be in GPU memory, even though only 33 billion run. Memory follows the total; speed follows the active count. Sparsity buys compute, not RAM.
  • Communication. Experts are usually spread across several GPUs, so tokens have to be shipped to whichever device holds their expert and the results shipped back — twice per layer.
  • Fine-tuning. On a small dataset the router can shift its habits and fall back onto a handful of experts, undoing the balance that training worked to establish.

The short version

  • A mixture-of-experts layer replaces one feed-forward network with many, plus a router that chooses between them.
  • The router scores every expert, keeps the top two, and rescales their weights to sum to 1.
  • The output is a weighted blend of just those two; the rest never run.
  • Experts are not given subjects — the division of labour is learned, and it is usually mundane.
  • Memory follows the total parameter count, compute follows the active count.
  • A balancing loss stops the router collapsing onto a few favourites and wasting the others.