Multi-GPU Model Splitting
Split large models across multiple GPUs in Ollama to fit models that no single card can hold.
TL;DR
- Pick which GPUs Ollama uses with
CUDA_VISIBLE_DEVICESfirst. - Check the per-GPU memory split in
ollama psoutput. - Force a spread across all GPUs with
OLLAMA_SCHED_SPREAD.
How Splitting Works
Automatic SplitOllama spreads a model across GPUs when it won't fit one.
# Fits on one GPU if it can, else splits
ollama run llama3.1:70bPipeline ParallelismLayers run in sequence across GPUs, not in parallel.
# One GPU active at a time per requestVRAM, Not SpeedExtra GPUs add capacity for bigger models, not throughput.
# 2 GPUs = more VRAM, similar tok/sSelect GPUs
CUDA_VISIBLE_DEVICESList the GPU indices Ollama is allowed to use.
export CUDA_VISIBLE_DEVICES=0,1Single GPURestrict to one card to keep a model together.
export CUDA_VISIBLE_DEVICES=0Reserve A CardHide a GPU you use for display from Ollama.
export CUDA_VISIBLE_DEVICES=1Force A Spread
OLLAMA_SCHED_SPREADSpread a model across all GPUs even if it fits one.
export OLLAMA_SCHED_SPREAD=1When To UseHandy for benchmarking or freeing VRAM on one card.
# Frees VRAM on the primary GPUDefault BehaviorWithout it, Ollama fills one GPU before the next.
# Default packs onto the fewest GPUsVerify The Split
ollama psShow how a loaded model is split across GPUs.
ollama ps # PROCESSOR shows the GPUsCheck LogsThe server log lists layers assigned per GPU.
journalctl -u ollama | grep -i gpunvidia-smiWatch live VRAM use on each GPU.
nvidia-smi # per-GPU memory in useTips
- Keep a model on a single GPU whenever it fits, since layer splitting adds VRAM capacity but not real single-request speed.
- Use
CUDA_VISIBLE_DEVICESto reserve a card for your display, so Ollama loads the model only on the GPUs you choose.
Warnings
- Ollama does pipeline parallelism, not tensor parallelism; a second GPU lets bigger models fit but rarely doubles your tokens per second.
- Setting
OLLAMA_SCHED_SPREAD=1forces a split even when a model fits one GPU, which can slow a single request needlessly.
In Practice
Expose both GPUs, run a 70B model that splits automatically, then confirm the per-GPU memory split.
- Listing both indices makes both cards visible to Ollama.
- A 70B model exceeds one card, so Ollama splits its layers across both.
ollama psshows the model spread over both GPUs.nvidia-smiconfirms real VRAM use on each card.
# Make both GPUs visible to Ollama
export CUDA_VISIBLE_DEVICES=0,1
# Too big for one card, so it splits
ollama run llama3.1:70b "Draft a release note"
# Confirm the split and per-GPU memory
ollama ps
nvidia-smiFAQ
Not for a single request. Ollama splits models by layer, so only one GPU is active at a time. A second GPU adds VRAM so bigger models fit, but it does not meaningfully raise tokens per second for one prompt.
If a model fits on one GPU, Ollama loads it there. If it is too big, Ollama automatically splits its layers across the visible GPUs. This is pipeline parallelism, where layers run in sequence across cards.
Set CUDA_VISIBLE_DEVICES to a comma-separated list of GPU indices, such as 0,1. Ollama only sees those cards, which lets you reserve a GPU for your display or a different workload.
By default, no. Ollama uses layer-based pipeline parallelism, which favors fitting a model over raw speed. An experimental tensor split mode exists, but layer splitting remains the default for stability.