All posts

LoRA

/AI/4 min read

Fine-tuning normally means updating every weight in the model. LoRA freezes them all and learns a pair of thin matrices alongside, which turns out to be enough.

Fine-tuning a model normally means updating every weight it has. LoRA leaves all of them frozen and learns a small pair of matrices next to each one instead.

The frozen model still does the work. The pair adds a correction on top.

Why the normal way is expensive

The parameter count is not the problem on its own — it is what training drags along with it.

Every trainable parameter needs a gradient, and an optimiser like Adam keeps two running averages per parameter on top of that. In fp32, with a master copy of the weights, that comes to roughly 16 bytes for every parameter you intend to train.

For a model of 3.62 billion parameters that is 54 GiB before a single activation is stored. The weights themselves are the small part.

Then there is what you are left with. Fine-tuning produces a whole new model, so a second task means a second 6.75 GiB file, and ten tasks means ten of them.

The bet

LoRA rests on one claim: the change a model needs in order to learn a task is much simpler than the model itself.

The weights are a d×dd \times d matrix. The update you would apply to them is also d×dd \times d, but it need not be a full-rank one — the useful part of it may live in a far smaller space. If that is true, the update can be written as a product of two thin matrices and nothing important is lost.

ΔW=BA\Delta W = BA

AA is r×dr \times d and BB is d×rd \times r, where rr is small — 8, 16, 32. Multiplying them gives something back at full d×dd \times d size, but built from only 2rd2rd numbers.

W + B × A = ΔW 3072 × 3072, frozen 9,437,184 3072×16 and 16×3072, trained 98,304 full size again
The update comes out the same shape as the weights, but there are 96 times fewer numbers behind it.

How it trains

Freeze WW. It never receives a gradient and never changes.

Initialise AA randomly and BB to zeros. That detail matters: BABA is zero at the start, so before any training the model behaves exactly as it did before, and the fine-tune begins from the pretrained behaviour rather than from a jolt.

The forward pass runs both paths and adds them:

h=Wx+αr(BA)xh = Wx + \frac{\alpha}{r}(BA)x

The α/r\alpha/r term is a scale factor. It exists so that changing the rank does not silently change how strongly the adapter speaks — pick α=2r\alpha = 2r and the scale stays at 2 whatever rank you choose.

Only AA and BB get gradients. Everything else is along for the ride.

What it saves

Take a 32-layer model with d=3072d = 3072, applying LoRA at rank 16 to the query and value projections only — which is usually enough.

full fine-tuningLoRA
trainable parameters3,623,878,6566,291,456
share of the model100%0.174%
optimiser + gradient memory54.0 GiB96 MiB
what you ship afterwards6.75 GiB12 MiB

The memory line is what changes who can do this. 54 GiB of optimiser state needs several datacentre GPUs; 96 MiB needs none of them, and the frozen weights can sit in a compressed form because nothing is writing to them.

The shipping line changes what you can distribute. Ten fine-tunes of the same base model are ten 12 MiB files, not ten copies of a 6.75 GiB one.

Merging and swapping

After training you have a choice.

Fold the adapter in — Wmerged=W+αrBAW_{\text{merged}} = W + \frac{\alpha}{r}BA — and you get an ordinary weight matrix of the usual shape. Inference costs exactly what it did before, because there is no extra path left to run. This is one-way, though: once merged, that adapter cannot be peeled back off.

Or leave it separate, and the base model can serve many tasks at once. Load one set of weights, keep several adapters beside it, and switch which one is active per request. That only works while they are unmerged, which is why serving systems usually keep them that way and accept the small cost of the second path.

The short version

  • Full fine-tuning is expensive because of optimiser state, not the weights — roughly 16 bytes per trainable parameter.
  • LoRA freezes the original weights and learns ΔW=BA\Delta W = BA, with AA and BB thin.
  • BB starts at zero, so training begins from the pretrained model's exact behaviour.
  • The forward pass adds the two paths, scaled by α/r\alpha/r so rank changes stay neutral.
  • At rank 16 on a 3.6B model, 0.174% of the parameters are trained and optimiser memory drops from 54 GiB to 96 MiB.
  • Merge the adapter for free inference, or keep it separate to swap tasks on one loaded model.