Ollama Environment Variables
Configure Ollama storage, memory, GPU use, and networking with the environment variables the server reads.
TL;DR
- Point
OLLAMA_MODELSat a bigger drive for weights. - Pin a model in memory with
OLLAMA_KEEP_ALIVE=-1to skip reloads. - Reserve GPU memory for other apps with
OLLAMA_GPU_OVERHEAD.
Storage And Networking
OLLAMA_MODELSSet the directory where model weights are stored.
export OLLAMA_MODELS=/mnt/ssd/ollamaOLLAMA_HOSTSet the address and port the server binds to.
export OLLAMA_HOST=0.0.0.0:11434OLLAMA_ORIGINSAllow specific web origins to call the local API.
export OLLAMA_ORIGINS=https://myapp.exampleMemory And Model Loading
OLLAMA_KEEP_ALIVESet how long a model stays loaded when idle.
export OLLAMA_KEEP_ALIVE=-1 # never unloadOLLAMA_MAX_LOADED_MODELSCap how many models can be resident at once.
export OLLAMA_MAX_LOADED_MODELS=2OLLAMA_NUM_PARALLELSet how many requests one model handles concurrently.
export OLLAMA_NUM_PARALLEL=4GPU And Performance
OLLAMA_GPU_OVERHEADReserve VRAM per GPU for other applications, in bytes.
export OLLAMA_GPU_OVERHEAD=1073741824
# reserves 1 GiB of VRAMOLLAMA_FLASH_ATTENTIONEnable flash attention to cut memory use on long context.
export OLLAMA_FLASH_ATTENTION=1OLLAMA_KV_CACHE_TYPEQuantize the K/V cache to save memory on context.
export OLLAMA_KV_CACHE_TYPE=q8_0CUDA_VISIBLE_DEVICESRestrict Ollama to specific NVIDIA GPUs by index.
export CUDA_VISIBLE_DEVICES=0Context And Debugging
OLLAMA_CONTEXT_LENGTHSet the default context window size for models.
export OLLAMA_CONTEXT_LENGTH=8192OLLAMA_MAX_QUEUECap how many requests wait in line before rejection.
export OLLAMA_MAX_QUEUE=512OLLAMA_DEBUGTurn on verbose logging for troubleshooting the server.
export OLLAMA_DEBUG=1Applying Variables
Linux (systemd)Add Environment lines to the service override file.
sudo systemctl edit ollama.service
# Environment="OLLAMA_KEEP_ALIVE=-1"macOSSet the variable for the app, then restart Ollama.
launchctl setenv OLLAMA_KEEP_ALIVE -1WindowsSet account environment variables, then restart Ollama.
# Settings > Edit environment variablesVerify (Linux)Confirm the running service actually sees the variable.
cat /proc/$(pgrep ollama)/environ |
tr '\0' '\n' | grep OLLAMATips
- Set
OLLAMA_KEEP_ALIVE=-1to keep a busy model resident between requests, which avoids the reload delay on every call. - Raise
OLLAMA_NUM_PARALLELto handle several requests at once, but only if your GPU has enough free memory for the extra load.
Warnings
OLLAMA_MAX_VRAMwas removed from Ollama; useOLLAMA_GPU_OVERHEADto reserve VRAM per GPU instead of that old variable.- Setting
OLLAMA_KEEP_ALIVE=-1pins a model in memory forever, which can block other models from loading on limited hardware.
In Practice
Relocate model storage to a large drive and pin a model in memory using systemd environment lines.
systemctl editadds the variables in an override that survives updates.OLLAMA_MODELSmoves large weights off the small system disk.OLLAMA_KEEP_ALIVE=-1keeps the model resident so requests skip reloading.- Reading
/proc/.../environproves the service picked up both variables.
# Edit the service to add two variables
sudo systemctl edit ollama.service
# Under [Service], add:
# Environment="OLLAMA_MODELS=/mnt/ssd/ollama"
# Environment="OLLAMA_KEEP_ALIVE=-1"
# Apply the changes
sudo systemctl daemon-reload
sudo systemctl restart ollama
# Verify the running service sees them
cat /proc/$(pgrep ollama)/environ |
tr '\0' '\n' | grep OLLAMAFAQ
Set OLLAMA_MODELS to a directory on the drive you want, then restart the server. This is useful for moving large weights off a small system disk onto an external SSD.
Set OLLAMA_KEEP_ALIVE to a duration like 30m, or to -1 to keep the model loaded indefinitely. The default is 5 minutes of idle time before Ollama unloads it.
The old OLLAMA_MAX_VRAM variable was removed. Use OLLAMA_GPU_OVERHEAD to reserve VRAM in bytes for other apps, and OLLAMA_MAX_LOADED_MODELS to cap how many models load at once.
Set OLLAMA_HOST to 0.0.0.0:11434 so the server binds to all network interfaces. On a shared network, also review OLLAMA_ORIGINS to control which web origins may call the API.