Multi-GPU Model Splitting

Split large models across multiple GPUs in Ollama to fit models that no single card can hold.

TL;DR

  1. Pick which GPUs Ollama uses with CUDA_VISIBLE_DEVICES first.
  2. Check the per-GPU memory split in ollama ps output.
  3. Force a spread across all GPUs with OLLAMA_SCHED_SPREAD.

How Splitting Works

    Automatic Split

    Ollama spreads a model across GPUs when it won't fit one.

    # Fits on one GPU if it can, else splits
    ollama run llama3.1:70b
    Pipeline Parallelism

    Layers run in sequence across GPUs, not in parallel.

    # One GPU active at a time per request
    VRAM, Not Speed

    Extra GPUs add capacity for bigger models, not throughput.

    # 2 GPUs = more VRAM, similar tok/s

Select GPUs

    CUDA_VISIBLE_DEVICES

    List the GPU indices Ollama is allowed to use.

    export CUDA_VISIBLE_DEVICES=0,1
    Single GPU

    Restrict to one card to keep a model together.

    export CUDA_VISIBLE_DEVICES=0
    Reserve A Card

    Hide a GPU you use for display from Ollama.

    export CUDA_VISIBLE_DEVICES=1

Force A Spread

    OLLAMA_SCHED_SPREAD

    Spread a model across all GPUs even if it fits one.

    export OLLAMA_SCHED_SPREAD=1
    When To Use

    Handy for benchmarking or freeing VRAM on one card.

    # Frees VRAM on the primary GPU
    Default Behavior

    Without it, Ollama fills one GPU before the next.

    # Default packs onto the fewest GPUs

Verify The Split

    ollama ps

    Show how a loaded model is split across GPUs.

    ollama ps   # PROCESSOR shows the GPUs
    Check Logs

    The server log lists layers assigned per GPU.

    journalctl -u ollama | grep -i gpu
    nvidia-smi

    Watch live VRAM use on each GPU.

    nvidia-smi   # per-GPU memory in use

Tips

  1. Keep a model on a single GPU whenever it fits, since layer splitting adds VRAM capacity but not real single-request speed.
  2. Use CUDA_VISIBLE_DEVICES to reserve a card for your display, so Ollama loads the model only on the GPUs you choose.

Warnings

  1. Ollama does pipeline parallelism, not tensor parallelism; a second GPU lets bigger models fit but rarely doubles your tokens per second.
  2. Setting OLLAMA_SCHED_SPREAD=1 forces a split even when a model fits one GPU, which can slow a single request needlessly.

In Practice

FAQ