CPU And GPU Offloading

Understand how Ollama splits model layers between GPU and CPU, tune the split, and use Apple unified memory.

TL;DR

  1. Ollama offloads model layers to the GPU for speed.
  2. It spills extra layers to the CPU when VRAM fills.
  3. Cap GPU layers with the num_gpu parameter per session.

How Offloading Works

    Layers On GPU

    Ollama loads as many layers as VRAM allows.

    # More GPU layers means faster inference
    ollama ps
    Spill To CPU

    Remaining layers run on the CPU when VRAM is full.

    # CPU layers are much slower than GPU
    Hybrid Inference

    A model can split across GPU and CPU at once.

    # e.g. 24/33 layers on GPU, rest on CPU

Control Layer Offload

    num_gpu (session)

    Set how many layers load on the GPU this session.

    >>> /set parameter num_gpu 24
    num_gpu (Modelfile)

    Bake the GPU layer count into a custom model.

    PARAMETER num_gpu 24
    Force CPU Only

    Set zero GPU layers to run entirely on the CPU.

    >>> /set parameter num_gpu 0
    Check Result

    Confirm the new split took effect after changing it.

    ollama ps   # confirm the new split

Apple Silicon Unified Memory

    Unified Memory

    Mac CPU and GPU share one memory pool.

    # System RAM doubles as VRAM on a Mac
    Metal Backend

    Ollama uses Apple's Metal API automatically.

    # Metal runs the GPU layers on macOS
    Bigger Models

    Large shared memory lets Macs load bigger models.

    # A 32 GB Mac can run 27B-class models
    ollama run gemma3:27b

Diagnose With Logs

    ollama ps Split

    Read the CPU/GPU split for each loaded model.

    ollama ps   # see the PROCESSOR column
    Offload In Logs

    The log reports how many layers reached the GPU.

    journalctl -u ollama | grep -i offload
    Enable Debug

    Verbose logging shows detailed placement decisions.

    OLLAMA_DEBUG=1 ollama serve

Tips

  1. Watch the PROCESSOR column in ollama ps to see the CPU and GPU split for each loaded model, so you know where inference runs.
  2. Lower num_gpu when a model overflows VRAM, so fewer layers load on the GPU and the server stops crashing on start.

Warnings

  1. When VRAM runs out, layers fall back to the CPU and inference slows sharply; expect far fewer tokens per second.
  2. A high num_gpu on a small GPU forces an overflow; set it below the model's total layer count to keep loading stable.

In Practice

FAQ