Understanding GGUF Quantization

Read GGUF quantization labels, estimate the VRAM they save, and choose the right precision for your hardware.

TL;DR

  1. GGUF packs weights, tokenizer, and metadata into one file.
  2. Q4_K_M balances speed and quality for most models.
  3. 4-bit quantization cuts VRAM use by roughly 70% versus fp16.

What GGUF Is

    Single File

    GGUF stores weights and metadata in one file.

    ls *.gguf   # one self-contained model file
    Bundled Metadata

    Tokenizer and config travel inside the GGUF file.

    # No separate config.json needed for GGUF
    Run From GGUF

    Import a GGUF directly with a short Modelfile.

    echo 'FROM ./model.gguf' > Modelfile

Reading Quant Labels

    Q4_K_M

    4-bit K-quant, medium size: the common default.

    # Q4 = 4-bit, K = K-quant, M = medium
    Q8_0

    8-bit, near-full quality but a larger file.

    ollama pull llama3.1:8b-instruct-q8_0
    Q4_0

    Older 4-bit scheme, smaller but lower quality.

    # Prefer Q4_K_M over legacy Q4_0
    Bits Per Weight

    The leading number is bits stored per weight.

    # Q4 ~ 4 bits, Q8 ~ 8 bits per weight

VRAM Savings

    fp16 Baseline

    Full precision uses about 2 bytes per weight.

    # 7B fp16 ~ 14 GB of weights
    4-bit Size

    Four-bit weights are roughly a quarter the size.

    # 7B Q4_K_M ~ 4-5 GB of weights
    Rough Savings

    Dropping fp16 to 4-bit cuts size about 70%.

    # ~70-75% smaller than fp16
    Add Context

    The K/V cache adds memory on top of weights.

    # A bigger num_ctx needs more VRAM

Choosing A Quant

    Balanced

    Q4_K_M suits most machines and general use.

    ollama pull llama3.1:8b-instruct-q4_K_M
    Max Quality

    Q8_0 when accuracy matters and VRAM allows.

    ollama pull llama3.1:8b-instruct-q8_0
    Tight Memory

    Q3_K_M fits bigger models into small VRAM.

    # Q3_K_M trades quality for a smaller size

Tips

  1. Start with Q4_K_M for a strong balance of size, speed, and quality, then move to Q8_0 only if you need more accuracy.
  2. Match the quantization to your VRAM; a smaller quant like Q3_K_M lets a bigger model fit when memory is tight.

Warnings

  1. Very low quants like Q2_K save memory but noticeably hurt quality; avoid them unless memory leaves you no other option.
  2. A quant level is not free quality; Q4_0 is smaller but usually worse than the K-quant Q4_K_M at a similar size.

In Practice

FAQ