# Quantization

URL: https://softwaredictionary.org/terms/quantization
Category: AI & Machine Learning
Last updated: 2026-09-30

In short: Quantization is a technique that shrinks an AI model by storing its parameters in fewer bits, such as 8 or 4 instead of 16, so inference is faster and cheaper.

## What is quantization in AI?

Quantization reduces the precision of the numbers inside a machine learning model. Models are usually trained with parameters stored as 16- or 32-bit floating-point numbers, and quantization converts them to smaller formats such as 8-bit or 4-bit integers. The model keeps the same architecture and behaves almost the same, but it takes a fraction of the memory and often runs faster.

The basic method maps a range of values onto a small set of levels. For 8-bit quantization, each group of weights gets a scale factor, and every weight is divided by that scale and rounded to one of 256 whole numbers; at run time the numbers are multiplied back by the scale. Post-training quantization applies this to a finished model, while quantization-aware training simulates the rounding during training so the model learns to tolerate it. The payoff is large: a 70-billion-parameter model needs about 140 GB at 16 bits but roughly 35 GB at 4 bits.

An analogy is saving a photo with fewer colors or rounding prices to the nearest dollar: you lose a little detail, but the result is much smaller and still does the job. Quantization is what makes it practical to run capable models on a single GPU, a laptop, or a phone, and it lets servers handle more requests with the same hardware. The trade-off is some loss of accuracy, which is usually small at 8 bits and more noticeable at 4 bits and below.

Quantization is sometimes confused with compressing a file, but a quantized model is not unzipped before use; it runs directly on the lower-precision numbers, and the lost precision is gone for good. It is also different from distillation, which trains a new, smaller model to imitate a larger one, and from pruning, which removes weights entirely. These techniques are often combined, and despite both turning things into numbers, quantization has nothing to do with tokenization.

## Key takeaways

- Quantization stores model parameters with fewer bits, such as 8 or 4.
- It cuts memory use sharply and often speeds up inference.
- A scale factor maps the original values to a small range of integers and back.
- Accuracy loss is usually small at 8 bits and larger at very low bit widths.
- Distillation and pruning are different ways to make models smaller.

## Example: Quantizing weights to 8-bit integers and back

```python
# Quantize a few floating-point weights to 8-bit integers and back
weights = [0.4213, -1.27, 0.0318, 0.8871, -0.5096]

scale = max(abs(w) for w in weights) / 127       # map the largest value to 127
quantized = [round(w / scale) for w in weights]  # small integers: 1 byte each
restored = [q * scale for q in quantized]        # what the model computes with

print(quantized)                        # [42, -127, 3, 89, -51]
print([round(r, 3) for r in restored])  # [0.42, -1.27, 0.03, 0.89, -0.51]
# Close to the originals, but the small rounding errors are permanent
```

## Frequently asked questions

**Does quantization reduce model accuracy?**

Usually a little. At 8 bits the difference is often hard to notice, while 4-bit and lower can cause measurable drops, especially on complex reasoning, so quantized models should be tested on your own tasks.

**What does 4-bit quantization mean?**

It means each parameter is stored in 4 bits, which allows only 16 distinct values per group of weights. It uses about a quarter of the memory of 16-bit weights, which lets much larger models fit on consumer hardware.

**What is the difference between quantization and distillation?**

Quantization keeps the same model but stores its numbers with less precision. Distillation trains a separate, smaller student model to reproduce the behavior of a larger teacher model.

---

Software Dictionary: https://softwaredictionary.org/ · https://softwaredictionary.org/llms.txt
