Understanding GGUF Quantization
Read GGUF quantization labels, estimate the VRAM they save, and choose the right precision for your hardware.
TL;DR
GGUFpacks weights, tokenizer, and metadata into one file.Q4_K_Mbalances speed and quality for most models.- 4-bit quantization cuts VRAM use by roughly
70%versus fp16.
What GGUF Is
Single FileGGUF stores weights and metadata in one file.
ls *.gguf # one self-contained model fileBundled MetadataTokenizer and config travel inside the GGUF file.
# No separate config.json needed for GGUFRun From GGUFImport a GGUF directly with a short Modelfile.
echo 'FROM ./model.gguf' > ModelfileReading Quant Labels
Q4_K_M4-bit K-quant, medium size: the common default.
# Q4 = 4-bit, K = K-quant, M = mediumQ8_08-bit, near-full quality but a larger file.
ollama pull llama3.1:8b-instruct-q8_0Q4_0Older 4-bit scheme, smaller but lower quality.
# Prefer Q4_K_M over legacy Q4_0Bits Per WeightThe leading number is bits stored per weight.
# Q4 ~ 4 bits, Q8 ~ 8 bits per weightVRAM Savings
fp16 BaselineFull precision uses about 2 bytes per weight.
# 7B fp16 ~ 14 GB of weights4-bit SizeFour-bit weights are roughly a quarter the size.
# 7B Q4_K_M ~ 4-5 GB of weightsRough SavingsDropping fp16 to 4-bit cuts size about 70%.
# ~70-75% smaller than fp16Add ContextThe K/V cache adds memory on top of weights.
# A bigger num_ctx needs more VRAMChoosing A Quant
BalancedQ4_K_M suits most machines and general use.
ollama pull llama3.1:8b-instruct-q4_K_MMax QualityQ8_0 when accuracy matters and VRAM allows.
ollama pull llama3.1:8b-instruct-q8_0Tight MemoryQ3_K_M fits bigger models into small VRAM.
# Q3_K_M trades quality for a smaller sizeTips
- Start with
Q4_K_Mfor a strong balance of size, speed, and quality, then move toQ8_0only if you need more accuracy. - Match the quantization to your VRAM; a smaller quant like
Q3_K_Mlets a bigger model fit when memory is tight.
Warnings
- Very low quants like
Q2_Ksave memory but noticeably hurt quality; avoid them unless memory leaves you no other option. - A quant level is not free quality;
Q4_0is smaller but usually worse than the K-quantQ4_K_Mat a similar size.
In Practice
Pull the same model at 8-bit and 4-bit, compare their sizes, then watch the smaller one's speed.
- The q8_0 build is near-full quality but the largest download.
- The q4_K_M build is roughly half the size with little quality loss.
ollama listshows both sizes side by side on disk.--verboseprints tokens per second so you can feel the speed gain.
# High quality, larger download
ollama pull llama3.1:8b-instruct-q8_0
# Balanced default, ~half the size
ollama pull llama3.1:8b-instruct-q4_K_M
# Compare sizes on disk
ollama list
# Watch tokens/sec on the smaller one
ollama run llama3.1:8b-instruct-q4_K_M --verboseFAQ
GGUF is a single-file format that bundles a model's quantized weights, tokenizer, and metadata together. Because everything travels in one file, Ollama can import it with a one-line Modelfile and no separate config.
The 4 means 4 bits per weight. K marks a K-quant scheme that holds quality better at a given size, and M is the medium variant. Q4_K_M is the common, balanced default.
Full precision (fp16) uses about 2 bytes per weight, so a 7B model is roughly 14 GB. At 4-bit it drops to about 4-5 GB, a cut of roughly 70-75%. You still add memory for the context cache.
Use Q4_K_M on most machines. Choose Q8_0 when accuracy matters and you have the VRAM, and drop to Q3_K_M only to squeeze a larger model into limited memory.