Model Compression
Making a trained model fit and run on the hardware you actually have, without discovering later which part of the distribution you broke.
Model compression is the set of techniques that reduce a trained model's memory footprint, bandwidth demand and compute cost while keeping its behavior close enough to the original to be useful. It matters because the binding constraint in deployment is almost never raw arithmetic throughput — it is how many bytes have to move, how much accelerator memory the weights and activations occupy, and whether the model fits on the hardware you are actually allowed to buy.
Why size, not FLOPs, is usually the constraint
During single-stream inference a model reads all of its weights to produce each output step. That makes generation memory-bandwidth-bound: the accelerator spends its time waiting on memory, not saturating its compute units. Halving the bytes per parameter therefore roughly halves the data that must be read per step, which is why numerical precision reduction gives a speedup that a pure FLOP count does not predict. The same arithmetic decides whether a model fits at all: storage is parameter count multiplied by bytes per parameter, so a model held in 32-bit floats needs four bytes per parameter, 16-bit formats two, 8-bit integers one, and 4-bit formats half a byte. Those ratios are definitional, and they are the first thing to compute when sizing hardware.
Quantization
Quantization stores and often computes with lower-precision numbers. A group of floating-point values is mapped onto a smaller integer range using a scale (and sometimes a zero point), and dequantized on the way back. The design space has a few axes that determine whether it works:
- Post-training quantization (PTQ) converts an already-trained model, optionally using a small calibration set to choose scales. Cheap, fast, and the right starting point.
- Quantization-aware training (QAT) simulates the rounding during training or fine-tuning so the weights adapt to it. More expensive, more robust at aggressive bit widths.
- Weight-only versus weight-and-activation. Weight-only quantization shrinks the model and relieves bandwidth while computing in higher precision; quantizing activations too can unlock faster integer kernels but is harder because activations vary per input.
- Granularity. A single scale per tensor is simplest but must cover the tensor's full range. Per-channel or per-group scales cost a little metadata and tolerate much more variation in magnitude, which usually matters more than the bit width itself.
The recurring difficulty is outliers: a few activation channels with magnitudes far above the rest force a coarse scale on everything else. Techniques that isolate or rescale those channels exist precisely because of this. The other recurring difficulty is the calibration set — scales derived from data unlike production traffic produce a model that is quietly worse on the inputs you actually receive.
Pruning and sparsity
Pruning removes parameters judged unimportant, typically by magnitude or by an estimate of the loss increase from removing them, usually followed by fine-tuning to recover quality.
The critical distinction is structural. Unstructured pruning zeroes individual weights and can reach high sparsity with little quality loss — but dense kernels still read and multiply the zeros, so on ordinary hardware nothing gets faster and the model only shrinks if you store it in a sparse format. Structured pruning removes whole units — channels, attention heads, entire layers — leaving a smaller dense model that is genuinely faster everywhere, at the cost of a steeper quality drop per parameter removed. Some hardware supports specific semi-structured sparsity patterns; outside those patterns, treat unstructured sparsity as a storage technique, not a latency technique.
Distillation
Knowledge distillation trains a smaller student model to reproduce a larger teacher's outputs. Training against the teacher's full probability distribution rather than hard labels conveys more information — the relative ranking of the wrong answers carries signal about how the teacher generalizes. Distillation can also transfer intermediate representations, not just final outputs.
It is the only family here that can change architecture freely, and it produces a clean dense model with no special kernel requirements. Its cost is a real training run and a corpus of inputs to distil over; its risk is that the student inherits the teacher's errors and biases as if they were ground truth.
Factorization and architecture choices
Low-rank factorization replaces a large weight matrix with the product of two smaller ones, trading exactness for parameters. Weight sharing reuses the same parameters across layers. Neither is usually the largest win on its own, but both compose with the techniques above. The largest win of all is often upstream: choosing a smaller model that was trained well rather than compressing a larger one aggressively.
Tradeoffs and failure modes
- Averaged metrics hide the damage. Compression rarely degrades everything uniformly; it degrades the tail. Rare classes, long inputs, minority languages and edge cases go first while the headline metric barely moves. Evaluate per-segment, not just overall.
- Proxy metrics mislead. For language models, a small change in perplexity can accompany a large change in a downstream task. Measure the task you actually ship.
- "It fits" is not "it is fast". A quantized model without a kernel that operates natively on the quantized format may dequantize on the fly and run no faster — sometimes slower.
- Compounding approximations. Quantization plus pruning plus a reduced-precision KV cache plus speculative decoding can each look acceptable alone and be unacceptable together. Introduce one at a time with an evaluation between.
- Calibration and eval drift. Both the calibration set and the evaluation set need to track production traffic, or the compressed model is tuned for a world that no longer exists.
- Reproducibility. Record the recipe — method, bit width, granularity, calibration data, kernel and runtime version — or the artifact cannot be rebuilt.
When to compress, and when not to
Compress when the model does not fit, when memory bandwidth is demonstrably your bottleneck, when you are paying for idle accelerator memory, or when the target is an edge device with a fixed budget. Start with the cheapest reversible step — weight-only post-training quantization at a conservative bit width — measure, and only escalate if it is not enough.
Do not compress before you have a trustworthy evaluation harness, because you will have no way to see what you broke. Do not compress to fix a quality problem; a smaller model will not be more accurate. And do not compress a model that is already cheap relative to the rest of the request path — if retrieval or network time dominates, the saving is invisible. Related: deep learning for the architectures being compressed, MLOps for versioning the resulting artifacts, and AI model quantization for a deeper treatment.