Scaling With A Proxy
Serve many users by raising per-instance concurrency and load balancing Ollama replicas with Nginx.
TL;DR
- Let one server handle several requests with
OLLAMA_NUM_PARALLEL. - Run multiple Ollama replicas behind an
Nginxproxy. - Spread load across replicas with an Nginx
upstreamblock.
In-Process Concurrency
OLLAMA_NUM_PARALLELSet how many requests one model serves at once.
export OLLAMA_NUM_PARALLEL=4DefaultDefaults to 1 or 4 based on free memory.
# 1 serializes; raise for concurrencyMemory CostEach slot needs its own KV cache.
# More parallel = more VRAM usedScale Out
Multiple ReplicasRun several Ollama containers in parallel.
# docker compose up --scale ollama=3Same ModelsEach replica pulls the same models.
# Share a models volume across replicasFront With NginxPut a reverse proxy in front of them.
# Nginx balances across the replicasNginx Upstream
upstream BlockList the backend Ollama instances.
upstream ollama {
server a:11434;
server b:11434;
}proxy_passForward requests to the upstream group.
location / {
proxy_pass http://ollama;
}Round RobinNginx rotates requests by default.
# Default balances requests evenlyProduction Tips
Health ChecksProbe each instance before routing to it.
curl http://a:11434/api/tagsKeep Models WarmPin models so replicas respond instantly.
export OLLAMA_KEEP_ALIVE=-1One Per GPURun one instance per available GPU.
# 1 replica per GPU, not oversubscribedTips
- Set
OLLAMA_NUM_PARALLELabove 1 so a single instance serves several requests at once, as long as memory allows the extra KV cache. - Scale horizontally by running several Ollama containers and balancing them with an Nginx
upstreamblock for round-robin routing.
Warnings
- By default
OLLAMA_NUM_PARALLELmay be 1, so a single instance queues extra requests; raise it or add replicas for concurrency. - Each parallel slot needs its own KV cache memory; setting
OLLAMA_NUM_PARALLELtoo high can exhaust VRAM and crash model loads.
In Practice
Raise per-instance concurrency, then put two Ollama replicas behind an Nginx upstream for round-robin routing.
- Raising
OLLAMA_NUM_PARALLELlets each replica serve several requests. - The
upstreamblock lists both Ollama backends. - Nginx round-robins requests across the two servers.
proxy_passforwards client traffic to the upstream group.
# Each Ollama instance: raise concurrency
# export OLLAMA_NUM_PARALLEL=4
# nginx.conf: balance two replicas
upstream ollama {
server 10.0.0.1:11434;
server 10.0.0.2:11434;
}
server {
listen 80;
location / {
proxy_pass http://ollama;
}
}FAQ
Yes. Set OLLAMA_NUM_PARALLEL above 1 and a single loaded model serves that many requests at once. The default can be 1, which queues extra requests, so raise it when memory allows.
First raise OLLAMA_NUM_PARALLEL for in-process concurrency. For more load, run several Ollama instances and put an Nginx reverse proxy in front to balance requests across them.
It sets how many requests one loaded model handles simultaneously. Each slot uses its own KV cache memory, so higher concurrency costs more VRAM. Tune it to your hardware.
Define an Nginx upstream block listing each instance, then proxy_pass to it. Nginx round-robins requests across the backends by default, spreading load evenly.