Key takeaways
- Quantization stores weights in fewer bits, so a 4-bit file is about a quarter to a third the size of the 16-bit original.
- llama.cpp says quantization may introduce some accuracy loss, usually measured with perplexity and KL divergence; test a quantized model on your own tasks.
- Prefer quantized files from the publisher or a source that documents how they were made, and files made from the 16- or 32-bit original, not requantized.
On this page
A model's weights are millions or billions of numbers, and each is stored using a set number of bits; many releases use 16 bits per weight. Quantization stores them in fewer bits, often 8 or 4, so the weight files shrink roughly in proportion: a 4-bit copy is about a quarter to a third the size of the 16-bit original. That is how a model too large for a laptop's memory at 16 bits can fit once quantized. The cost is accuracy, because each stored number becomes an approximation; llama.cpp's documentation says quantization "may introduce some accuracy loss," which careful methods try to keep small. Quantized files come from publishers and from third parties, so prefer the publisher's or a source that documents how they were made.
Bits and precision
A bit holds a 0 or a 1, so a number stored in 4 bits can take only 16 different values, while 16 bits allow 65,536. Fewer bits mean coarser steps between the values a weight can take. Hugging Face's Transformers documentation (external site: huggingface.co) says weights are typically stored in 32-bit floating point, with 16-bit formats increasingly common. Formats use their bits in two main ways:
- Floating point splits the bits between an exponent, which sets how large or small a value can be, and a mantissa, which sets the fine detail. NVIDIA's Transformer Engine guide (external site: docs.nvidia.com) describes two 8-bit versions, called FP8: E4M3 (1 sign, 4 exponent, and 3 mantissa bits) stores values up to ±448, while E5M2 reaches ±57,344 at lower precision.
- Integer formats store evenly spaced whole numbers plus a scale, shared by a small group of weights, that converts them back. In the "affine" mode of Apple's MLX (external site: ml-explore.github.io), each group of 32, 64, or 128 weights stores a scale and a bias, at 2, 3, 4, 5, 6, or 8 bits per weight.
The arithmetic of weight files
This section is arithmetic about weight files, not a hardware requirement. A file's size is roughly the number of weights times the bits per weight, divided by 8 bits per byte. For an imaginary model with 8 billion weights:
- 32 bits per weight: 32 GB
- 16 bits: 16 GB
- 8 bits: 8 GB
- 4 bits: 4 GB
Shared scales add a little. Hugging Face's GGUF documentation (external site: huggingface.co) gives llama.cpp's Q4_K type 4.5 bits per weight, which would mean 4.5 GB here. In the OCP Microscaling (MX) specification (external site: opencompute.org), 32 values share one 8-bit scale, so MXFP4 uses 32 × 4 + 8 = 136 bits per 32 values, or 4.25 bits each.
llama.cpp's quantize README (external site: github.com) lists sizes for Meta's Llama 3.1 8B that follow this pattern: 14.96 GiB at 16 bits (F16), 7.95 GiB for Q8_0, and 4.58 GiB for Q4_K_M at about 4.89 bits per weight. A GiB is 2 to the 30th power bytes, about 7 percent more than a GB.
Running a model takes more memory than its file. Ollama's context length page (external site: docs.ollama.com) says a longer context increases the memory required; see understanding inference hardware.
Common formats: 16, 8, and 4 bits
- 16-bit. The GGUF documentation describes F16 as the standard half-precision format and BF16 as a shortened version of 32-bit floating point.
- 8-bit. The Transformer Engine guide says NVIDIA's H100 GPU introduced FP8 support. For integers, the bitsandbytes (external site: huggingface.co) library's LLM.int8() uses 8 bits for most of its matrix math and handles outlier values separately in 16 bits.
- 4-bit. MXFP4, from the MX specification, uses 4-bit floating-point values (E2M1) in blocks of 32 that share a scale. OpenAI's gpt-oss model card (external site: arxiv.org) says the mixture-of-experts weights, over 90% of all parameters, were post-trained with MXFP4 quantization at 4.25 bits per parameter, which lets gpt-oss-20b run on systems with as little as 16 GB of memory.
Weight-only methods: AWQ and GPTQ
Weight-only methods shrink the stored weights but leave activations, the intermediate values computed while the model runs, at higher precision. The AWQ paper gives W4A16 as an example of this setting, in which only the weights become low-bit integers. Both methods below are applied to an already trained model.
- GPTQ. The GPTQ paper (external site: arxiv.org) describes a one-shot method that uses approximate second-order information to reduce weights to 3 or 4 bits, and reports quantizing 175-billion-parameter models in about four GPU hours.
- AWQ. MIT HAN Lab's AWQ paper (external site: arxiv.org) starts from the finding that weights are not equally important: protecting about 1% of salient weights greatly reduces quantization error. It finds them from activation statistics collected on a small calibration set, sample data run through the model, then scales them up before quantizing. Its repository (external site: github.com) targets INT3 and INT4 weights.
The Transformers documentation adds that other methods work on the fly, without calibration; Ai2's Olmo 3 7B Instruct card (external site: huggingface.co) loads that model in 8-bit through bitsandbytes.
llama.cpp types and GGUF files
GGUF is a binary file format that holds a model's tensors (its arrays of weights) with a standard set of metadata, made for the GGML library and tools such as llama.cpp. The quantize README describes two phases: convert the original model to GGUF, typically in a high-precision format such as F32 or BF16, then pick a type:
./build/bin/llama-quantize gemma-4-E2B-it-bf16.gguf gemma-4-E2B-it-Q4_K_M.gguf Q4_K_M
The number in a type's name is its base bit width. "K" types group weights into super-blocks with their own scales, and variants differ in size: for Llama 3.1 8B, Q4_K_S is about 4.67 bits per weight and Q4_K_M about 4.89. "IQ" types use an importance matrix, which llama.cpp computes from the model and a sample text file and says can improve the quality of quantized models. On Hugging Face, a GGUF viewer shows each tensor's name, shape, and precision.
What you give up, in the tools' own words
- llama.cpp says the accuracy loss is usually measured with perplexity and KL divergence. Its perplexity README (external site: github.com) says perplexity measures how well a model predicts the next token, lower being better, and is not directly comparable between models; KL divergence measures how similar the quantized and 16-bit models' output distributions are, with 0 meaning identical.
- The quantize README warns that requantizing an already quantized file can severely reduce quality compared with starting from 16 or 32 bits.
- The bitsandbytes documentation describes LLM.int8() as halving the memory needed for inference without performance degradation.
- OpenAI's gpt-oss-20b card on Hugging Face (external site: huggingface.co) says all its evaluations were performed with the same MXFP4 quantization.
These are the publishers' descriptions. USASI does not publish scores; test a quantized model on your own tasks.
Choosing a quantized file
- Look to the publisher first. OpenAI's gpt-oss model card lists a 12.8 GiB checkpoint for gpt-oss-20b, with the mixture-of-experts weights in MXFP4. Nous Research publishes Hermes 4 70B in BF16 and FP8 itself, and its card points to GGUF files made by the LM Studio team.
- Otherwise, pick a documented source. Its card should name the original repository, the tool, and the type. Hugging Face's model card documentation (external site: huggingface.co) lists "quantized" among the relationships the Hub infers from the
base_modelfield. - Prefer files made from the 16- or 32-bit original, not requantized ones.
- Read the original license. A quantized copy is made from the original weights.
Worked example: picking a file for a laptop
Sam is a fictional reader invented for this page. Sam has a laptop with 16 GB of memory and wants to run an 8-billion-parameter model with llama.cpp.
- Do the arithmetic. At 16 bits, the weights alone would take about all 16 GB. At about 4.9 bits per weight, they come to under 5 GB.
- Leave room. The context needs memory too, so a file that nearly fills memory is a poor fit.
- Check the source. Sam looks for an official GGUF first, then for a repository that documents how its files were made.
- Pick a type. Using llama.cpp's Llama 3.1 8B figures, Q4_K_M (4.58 GiB) leaves more headroom than Q8_0 (7.95 GiB).
- Test. Sam runs the same prompts on both and compares the answers.
Sam's outcome: start with Q4_K_M and move to Q8_0 only if answers are worse on Sam's tasks and memory allows.
What you can do next
- Browse llama.cpp, AWQ, MLX, and gpt-oss-20b in the catalog.
- Look up quantization, GGUF, weights, and self-hosting in the glossary.
- Read understanding inference hardware, hosted or local, and pretraining, fine-tuning, and post-training.
- Visit the local AI hub.
Sources
All read on October 8, 2026.
- llama.cpp (ggml-org): quantize README (external site: github.com), imatrix README (external site: github.com), perplexity README (external site: github.com)
- Hugging Face: GGUF (external site: huggingface.co), Transformers quantization overview (external site: huggingface.co), bitsandbytes documentation (external site: huggingface.co), Model cards: specifying a base model (external site: huggingface.co)
- NVIDIA: Using FP8 and FP4 with Transformer Engine (external site: docs.nvidia.com)
- Open Compute Project: OCP Microscaling Formats (MX) Specification, version 1.0 (external site: opencompute.org)
- OpenAI: gpt-oss-120b and gpt-oss-20b Model Card (arXiv 2508.10925) (external site: arxiv.org), gpt-oss-20b on Hugging Face (external site: huggingface.co)
- MIT HAN Lab: AWQ paper (arXiv 2306.00978) (external site: arxiv.org), llm-awq repository (external site: github.com)
- IST Austria, ETH Zurich, and Neural Magic: GPTQ paper (arXiv 2210.17323) (external site: arxiv.org)
- Apple: MLX quantize documentation (external site: ml-explore.github.io), MLX README (external site: github.com)
- Nous Research: Hermes-4-70B card (external site: huggingface.co)
- Ai2: Olmo-3-7B-Instruct card (external site: huggingface.co)
- Ollama: Context length (external site: docs.ollama.com)
Check your understanding
Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.
0 of 3 answered.