Model Quantization
/AI/4 min read
Fewer bits per number is the easy part. What decides whether it works is how many numbers share a scale — and in one measured case the average error looked fine at 5% while the ordinary features were off by 42%.
Quantization stores each number in fewer bits. A 7-billion-parameter model:
| bytes per weight | total | |
|---|---|---|
| fp32 | 4 | 28.0 GB |
| fp16 | 2 | 14.0 GB |
| int8 | 1 | 7.0 GB |
| int4 | 0.5 | 3.5 GB |
That table is the entire motivation and none of the difficulty. The difficulty is that an 8-bit integer has 256 possible values, and you have to decide which 256 real numbers they stand for.
Scale, and who shares one
The mapping is one multiplication. Pick a scale s, store round(v / s), recover q × s. For a signed 8-bit integer the representable range is −127 to 127, so s must be large enough that the biggest value in the group fits.
Which means everything in the group is measured in units of the largest member. A group containing one enormous value forces a coarse scale on everything else.
So the real question is never "how many bits". It is how many numbers share a scale.
Weights: the grouping axis matters
A weight matrix has structure. Different output channels can differ enormously in magnitude — in a matrix built with a realistic spread, the largest channel was 227× the smallest.
Quantizing all of it with one scale, versus one scale per output channel:
| error, relative to the weights themselves | |
|---|---|
| one scale for the whole tensor | 4.49% |
| one scale per output channel | 0.69% |
A 6.5× improvement, for the cost of storing one extra number per channel — negligible against the channel's own hundreds of weights.
This is why per-channel quantization is the default for weights and the per-tensor variant is essentially never used. It is not a refinement; it is the difference between usable and not.
Activations: the same idea, much worse stakes
Weights are known in advance and can be grouped carefully offline. Activations are computed at run time and have a property that makes them far harder: in a transformer, a handful of feature dimensions carry values vastly larger than the rest.
Taking 256 features where four of them run about 44× larger than the others:
| all features | ordinary features only | |
|---|---|---|
| one scale for the tensor | 5.31% | 41.93% |
| one scale per feature | 0.83% | 0.69% |
Look at what the aggregate number does. Measured across everything, per-tensor quantization looks acceptable at 5.31% — because the four outlier features are large, they dominate the total magnitude, and their own error is proportionally small.
Measured on the 252 features that carry the actual signal, the error is 42%. The representation has been destroyed and the summary statistic says it is fine.
This is the central trap of quantization work. An average error weighted by magnitude hides exactly the failure you care about, because the outliers that caused it are also the values that dominate the average.
The consequences shape the whole field. Weight-only quantization is common and safe, because weights have no run-time outliers to surprise you. Activation quantization needs either per-feature scales computed on the fly, or the outlier dimensions held at higher precision, or a transformation that moves the problem from activations into weights before quantizing.
Symmetric or not
A symmetric scheme centres on zero: the representable values run from −127 to 127 and zero maps exactly to zero.
That is wasteful when the data is one-sided. Values after a rectifying activation are all non-negative, so half the integer range represents numbers that cannot occur.
Measured on non-negative data at 8 bits:
| error | |
|---|---|
| symmetric | 0.843% |
| asymmetric | 0.421% |
Exactly 2.00× — one whole bit, recovered by adding an offset so the range starts where the data starts. Asymmetric costs one extra stored integer per group and an addition per value, which is why it is normal for activations and usually skipped for weights, whose distributions are roughly centred anyway.
After training, or during
Post-training quantization takes a finished model and converts it, using a small calibration sample to choose scales. Cheap, fast, and the right first attempt.
Quantization-aware training simulates the rounding during training, so the weights learn to be robust to it. Far more expensive, and the option worth reaching for at very low bit widths where post-training conversion visibly degrades.
At 8 bits, post-training is usually indistinguishable from the original. At 4 bits it depends heavily on the method — the better ones reduce error layer by layer, or spend their precision budget on the weights that matter most rather than spreading it evenly.
What to take away
Bit width gets the headline and grouping does the work.
Before trusting any quantized model, ask two questions: which numbers share a scale, and — more importantly — what the error looks like on the values that are not the largest ones. An aggregate error figure computed over data containing outliers is not evidence of anything.