CPU And GPU Offloading
Understand how Ollama splits model layers between GPU and CPU, tune the split, and use Apple unified memory.
TL;DR
- Ollama offloads model layers to the
GPUfor speed. - It spills extra layers to the
CPUwhen VRAM fills. - Cap GPU layers with the
num_gpuparameter per session.
How Offloading Works
Layers On GPUOllama loads as many layers as VRAM allows.
# More GPU layers means faster inference
ollama psSpill To CPURemaining layers run on the CPU when VRAM is full.
# CPU layers are much slower than GPUHybrid InferenceA model can split across GPU and CPU at once.
# e.g. 24/33 layers on GPU, rest on CPUControl Layer Offload
num_gpu (session)Set how many layers load on the GPU this session.
>>> /set parameter num_gpu 24num_gpu (Modelfile)Bake the GPU layer count into a custom model.
PARAMETER num_gpu 24Force CPU OnlySet zero GPU layers to run entirely on the CPU.
>>> /set parameter num_gpu 0Check ResultConfirm the new split took effect after changing it.
ollama ps # confirm the new splitApple Silicon Unified Memory
Unified MemoryMac CPU and GPU share one memory pool.
# System RAM doubles as VRAM on a MacMetal BackendOllama uses Apple's Metal API automatically.
# Metal runs the GPU layers on macOSBigger ModelsLarge shared memory lets Macs load bigger models.
# A 32 GB Mac can run 27B-class models
ollama run gemma3:27bDiagnose With Logs
ollama ps SplitRead the CPU/GPU split for each loaded model.
ollama ps # see the PROCESSOR columnOffload In LogsThe log reports how many layers reached the GPU.
journalctl -u ollama | grep -i offloadEnable DebugVerbose logging shows detailed placement decisions.
OLLAMA_DEBUG=1 ollama serveTips
- Watch the
PROCESSORcolumn inollama psto see the CPU and GPU split for each loaded model, so you know where inference runs. - Lower
num_gpuwhen a model overflows VRAM, so fewer layers load on the GPU and the server stops crashing on start.
Warnings
- When VRAM runs out, layers fall back to the
CPUand inference slows sharply; expect far fewer tokens per second. - A high
num_gpuon a small GPU forces an overflow; set it below the model's total layer count to keep loading stable.
In Practice
Check the current CPU/GPU split, cap GPU layers to avoid an overflow, then confirm the change in the logs.
ollama psreveals whether a model already spilled onto the CPU.- Lowering
num_gpuleaves headroom so the GPU load stops crashing. - The model still answers, now with a stable, smaller GPU footprint.
- The server log's offload count confirms how many layers reached the GPU.
# See the CPU/GPU split for loaded models
ollama ps
# Limit layers on the GPU to avoid overflow
ollama run llama3.1:8b
>>> /set parameter num_gpu 24
>>> Write a haiku about servers
# Check the server log for the offload count
journalctl -u ollama | grep -i offloadFAQ
The model probably does not fit in VRAM, so Ollama spilled some layers to the CPU. Check ollama ps for the CPU/GPU split, then use a smaller model, a lower quant, or a smaller num_gpu.
On Apple Silicon, the CPU and GPU share one memory pool, so system RAM doubles as VRAM. A 32 GB Mac can load models that would need a large dedicated GPU on a PC. That is a real advantage for local inference.
Set the num_gpu parameter to the number of layers to place on the GPU. Use /set parameter num_gpu 24 in a session, or add PARAMETER num_gpu 24 to a Modelfile. Set it to 0 to force CPU-only.
Run ollama ps and read the PROCESSOR column, which shows the CPU/GPU percentage split. For detail, check the server log for the offloaded layer count, or start with OLLAMA_DEBUG=1.