All posts

GGUF

/AI/4 min read

GGUF is one file holding the weights, the tokenizer and everything needed to run a model — laid out so the runtime can map it into memory instead of parsing it, and quantized in small blocks so 4 bits per weight is actually usable.

GGUF is a file format for a trained model. One file contains the weights, the tokenizer, and the metadata describing how to run the thing — architecture, layer counts, context length, special token ids.

That "one file" is the whole design goal, and the two decisions that follow from it are the interesting part.

What running a model locally needs

A model is a large collection of numbers arranged into named tensors. To run it you also need to know how those tensors connect, how text becomes token ids, and which ids mean stop.

Split that across a weights file, a tokenizer file and a couple of config files and every one of them can drift out of step with the others. A tokenizer from a slightly different revision produces ids the weights were never trained on — and nothing crashes. The model just answers slightly wrong, forever.

GGUF keeps them together because they are not separable in practice. The file starts with the four bytes GGUF, then a table of key–value metadata, then the tensor data. A reader can learn everything about the model from the header before touching a single weight.

Why 4-bit is really 4.5-bit

Weights are trained at 16 bits each. A 7-billion-weight model is therefore 14 GB, which is more memory than most laptops will give a single process.

Quantization stores each weight in fewer bits. The naive version — pick one scale for the whole tensor, divide, round to the nearest of 16 levels — does not work, and it is worth seeing how badly.

Taking 32,768 weights drawn to look like a real tensor, mostly small with a scattering of large values:

error, as a share of the weights' own size
one scale for the whole tensor55.3%
one scale per block of 329.9%

A single outlier stretches the scale to cover its own magnitude, and every ordinary weight then rounds to the same couple of levels. Almost all the information is gone.

The fix is to quantize in small blocks. Thirty-two weights share one scale, so an outlier only ruins its own block. The error drops 5.6×.

That scale has to be stored. A block of 32 costs 16 bytes of 4-bit values plus a 2-byte scale — 18 bytes, or 4.5 bits per weight:

bits/weight7B model
fp1616.0014.00 GB
Q8_08.507.44 GB
Q4_04.503.94 GB

So a "4-bit" model is 3.56× smaller than fp16, not 4×. The missing half-bit is the scale, and it is the reason the format works at all.

ONE BLOCK: 32 SMALL NUMBERS AND THE SCALE THEY SHARE scale 32 × 4 bits = 16 bytes + 2 bytes = 18 bytes per 32 weights = 4.50 bits each rounding error, relative to the weights themselves 55.3% — one scale for the whole tensor 9.9% — one scale per block of 32 the extra half-bit is what buys the other 45 points
An outlier can only wreck the block it lives in.

Reading the names

A file marked Q4_K_M decodes as four separate choices:

  • Q — quantized, as opposed to full precision
  • 4 — the nominal bits per weight
  • K — the k-quant family, which stores scales in a smarter two-level arrangement rather than one flat scale per block
  • M — the medium variant, meaning some tensors are kept at higher precision than others

That last letter matters more than it looks. Not all tensors tolerate quantization equally — attention projections and embeddings are more sensitive than the bulk feed-forward weights. The S/M/L variants are different answers to which tensors get spent on.

Why loading is instant

A GGUF file stores tensor data in exactly the layout the runtime uses, aligned, with no per-tensor decoding step.

So the runtime does not read the file. It calls mmap and asks the operating system to make the file appear at an address. Nothing is copied. The header is parsed, the tensor pointers are computed as offsets into that mapping, and the model is ready.

The pages arrive from disk when they are first touched, which is why a 4 GB model can start answering before 4 GB has been read. It also means the memory is file-backed and clean: under pressure the OS can drop those pages and re-read them later, rather than swapping. And two processes running the same model share one copy of the pages, because they are mapping the same file.

None of this is possible if the weights need to be decompressed or rearranged on load. The layout is the feature.

What to take away

GGUF is unglamorous on purpose. A header, some key–value metadata, and tensor bytes arranged the way the runtime wants them.

What it gets right is that a model is not just weights — and that a "4-bit" model only works because it is not really 4 bits.