All posts

RMSNorm

/AI/4 min read

Layer normalisation does two things: it recentres values and it rescales them. RMSNorm drops the recentring, keeps the rescaling, and modern language models have almost all switched to it.

Stack enough layers and the numbers flowing through them drift. They grow, or they shrink towards nothing, and training becomes unstable either way.

Normalisation is the fix: at points along the network, pull the values back into a sensible range. RMSNorm is a stripped-down version of how that is usually done.

What layer normalisation does

Layer normalisation takes a vector and performs two separate corrections.

First it recentres: subtract the mean, so the values sit around zero. Then it rescales: divide by the standard deviation, so the spread is consistent.

LayerNorm(x)=γ⋅x−μσ2+ϵ+β\text{LayerNorm}(x) = \gamma \cdot \frac{x - \mu}{\sqrt{\sigma^2 + \epsilon}} + \beta

γ\gamma and β\beta are learned, so the network can undo or adjust the normalisation if it turns out to need to.

Dropping half of it

RMSNorm rests on a claim: of those two corrections, only the rescaling was doing much work.

So it keeps the scaling and throws the centring away. There is no mean to subtract, which means no variance to compute either — variance is defined around the mean. What replaces it is the root mean square, which measures magnitude directly, without reference to any centre:

RMS(x)=1n∑ixi2\text{RMS}(x) = \sqrt{\frac{1}{n}\sum_i x_i^2} RMSNorm(x)=γ⋅xmean(x2)+ϵ\text{RMSNorm}(x) = \gamma \cdot \frac{x}{\sqrt{\text{mean}(x^2) + \epsilon}}

Note what is missing. There is no β\beta — with no recentring step there is nothing for a shift parameter to correct — so RMSNorm carries half the learned parameters of LayerNorm.

Watching both run

Take a vector with a mean that is clearly not zero: x=[3,−1,5,7,−2,0]x = [3, -1, 5, 7, -2, 0].

Its mean is 2, its standard deviation is 3.2660, and its RMS is 3.8297.

resultmeanRMS
input3, −1, 5, 7, −2, 02.00003.8297
LayerNorm0.3062, −0.9186, 0.9186, 1.5309, −1.2247, −0.61240.00001.0000
RMSNorm0.7833, −0.2611, 1.3056, 1.8278, −0.5222, 0.00000.52221.0000

Both land the magnitude in exactly the same place — an RMS of 1. That is the part that stabilises training, and RMSNorm gets there without the mean.

Where they differ is the offset. LayerNorm has moved the whole vector so it straddles zero. RMSNorm has left it where it was, only smaller: the mean of 2 became 2 ÷ 3.8297 = 0.5222.

SAME MAGNITUDE, DIFFERENT CENTRE LAYERNORM 0 mean 0.00 · rms 1.00 RMSNORM 0 mean 0.52 · rms 1.00
The same six numbers after each. Both reach an RMS of 1; only one of them moves the centre.

Why leaving the offset alone is fine

The obvious worry is that a drifting mean will cause the trouble normalisation was supposed to prevent.

In a transformer it mostly does not, because the mean is not left unattended. Every sub-layer is followed by a linear projection, and a linear layer can shift its output by whatever its bias says — so any offset the network actually minds can be corrected there, using parameters that already exist. Recentring inside the normalisation step was doing a job something else was already able to do.

What could not be handled elsewhere is the magnitude, since that is what runs away across dozens of layers. RMSNorm keeps exactly that.

What it buys

Not much per call, but the calls are everywhere.

Computing a mean and then a variance means two passes over the vector, and the second depends on the first. Root mean square needs one pass and one square root. That removes a dependency as well as some arithmetic, which matters more than the operation count suggests — it is one fewer point where the whole vector has to be reduced before anything can continue.

The parameters shrink too. For a model with d=4096d = 4096 and 32 layers, normalising before attention, before the feed-forward network, and once at the end:

parameters
LayerNorm (γ\gamma and β\beta)532,480
RMSNorm (γ\gamma only)266,240

That is a rounding error against billions of weights. The speed is the reason it was adopted, not the size — normalisation runs twice per layer for every token of every batch across a training run of trillions of tokens, and a cheaper version of something that frequent is worth having.

Where it sits

In current models the normalisation goes before each sub-layer rather than after it — once before attention, once before the feed-forward network, and once more at the end before the output projection.

The residual path skips around it. That is the arrangement that makes deep stacks trainable: the untouched signal flows straight through, while each sub-layer sees an input that has already been brought back to a stable magnitude.

The short version

  • Deep stacks need normalisation or their values explode or vanish.
  • LayerNorm recentres and rescales; RMSNorm only rescales.
  • It divides by the root mean square, which measures magnitude without needing a mean.
  • With no centring step there is no β\beta, so it has half the learned parameters.
  • On the same vector both reach an RMS of 1; RMSNorm just leaves the offset in place.
  • That is safe because the following linear layer's bias can shift the mean anyway.
  • One pass instead of two, twice a layer, for trillions of tokens — which is why it took over.