Ollama Environment Variables

Configure Ollama storage, memory, GPU use, and networking with the environment variables the server reads.

TL;DR

  1. Point OLLAMA_MODELS at a bigger drive for weights.
  2. Pin a model in memory with OLLAMA_KEEP_ALIVE=-1 to skip reloads.
  3. Reserve GPU memory for other apps with OLLAMA_GPU_OVERHEAD.

Storage And Networking

    OLLAMA_MODELS

    Set the directory where model weights are stored.

    export OLLAMA_MODELS=/mnt/ssd/ollama
    OLLAMA_HOST

    Set the address and port the server binds to.

    export OLLAMA_HOST=0.0.0.0:11434
    OLLAMA_ORIGINS

    Allow specific web origins to call the local API.

    export OLLAMA_ORIGINS=https://myapp.example

Memory And Model Loading

    OLLAMA_KEEP_ALIVE

    Set how long a model stays loaded when idle.

    export OLLAMA_KEEP_ALIVE=-1   # never unload
    OLLAMA_MAX_LOADED_MODELS

    Cap how many models can be resident at once.

    export OLLAMA_MAX_LOADED_MODELS=2
    OLLAMA_NUM_PARALLEL

    Set how many requests one model handles concurrently.

    export OLLAMA_NUM_PARALLEL=4

GPU And Performance

    OLLAMA_GPU_OVERHEAD

    Reserve VRAM per GPU for other applications, in bytes.

    export OLLAMA_GPU_OVERHEAD=1073741824
    # reserves 1 GiB of VRAM
    OLLAMA_FLASH_ATTENTION

    Enable flash attention to cut memory use on long context.

    export OLLAMA_FLASH_ATTENTION=1
    OLLAMA_KV_CACHE_TYPE

    Quantize the K/V cache to save memory on context.

    export OLLAMA_KV_CACHE_TYPE=q8_0
    CUDA_VISIBLE_DEVICES

    Restrict Ollama to specific NVIDIA GPUs by index.

    export CUDA_VISIBLE_DEVICES=0

Context And Debugging

    OLLAMA_CONTEXT_LENGTH

    Set the default context window size for models.

    export OLLAMA_CONTEXT_LENGTH=8192
    OLLAMA_MAX_QUEUE

    Cap how many requests wait in line before rejection.

    export OLLAMA_MAX_QUEUE=512
    OLLAMA_DEBUG

    Turn on verbose logging for troubleshooting the server.

    export OLLAMA_DEBUG=1

Applying Variables

    Linux (systemd)

    Add Environment lines to the service override file.

    sudo systemctl edit ollama.service
    # Environment="OLLAMA_KEEP_ALIVE=-1"
    macOS

    Set the variable for the app, then restart Ollama.

    launchctl setenv OLLAMA_KEEP_ALIVE -1
    Windows

    Set account environment variables, then restart Ollama.

    # Settings > Edit environment variables
    Verify (Linux)

    Confirm the running service actually sees the variable.

    cat /proc/$(pgrep ollama)/environ |
      tr '\0' '\n' | grep OLLAMA

Tips

  1. Set OLLAMA_KEEP_ALIVE=-1 to keep a busy model resident between requests, which avoids the reload delay on every call.
  2. Raise OLLAMA_NUM_PARALLEL to handle several requests at once, but only if your GPU has enough free memory for the extra load.

Warnings

  1. OLLAMA_MAX_VRAM was removed from Ollama; use OLLAMA_GPU_OVERHEAD to reserve VRAM per GPU instead of that old variable.
  2. Setting OLLAMA_KEEP_ALIVE=-1 pins a model in memory forever, which can block other models from loading on limited hardware.

In Practice

FAQ