Scaling With A Proxy

Serve many users by raising per-instance concurrency and load balancing Ollama replicas with Nginx.

TL;DR

  1. Let one server handle several requests with OLLAMA_NUM_PARALLEL.
  2. Run multiple Ollama replicas behind an Nginx proxy.
  3. Spread load across replicas with an Nginx upstream block.

In-Process Concurrency

    OLLAMA_NUM_PARALLEL

    Set how many requests one model serves at once.

    export OLLAMA_NUM_PARALLEL=4
    Default

    Defaults to 1 or 4 based on free memory.

    # 1 serializes; raise for concurrency
    Memory Cost

    Each slot needs its own KV cache.

    # More parallel = more VRAM used

Scale Out

    Multiple Replicas

    Run several Ollama containers in parallel.

    # docker compose up --scale ollama=3
    Same Models

    Each replica pulls the same models.

    # Share a models volume across replicas
    Front With Nginx

    Put a reverse proxy in front of them.

    # Nginx balances across the replicas

Nginx Upstream

    upstream Block

    List the backend Ollama instances.

    upstream ollama {
      server a:11434;
      server b:11434;
    }
    proxy_pass

    Forward requests to the upstream group.

    location / {
      proxy_pass http://ollama;
    }
    Round Robin

    Nginx rotates requests by default.

    # Default balances requests evenly

Production Tips

    Health Checks

    Probe each instance before routing to it.

    curl http://a:11434/api/tags
    Keep Models Warm

    Pin models so replicas respond instantly.

    export OLLAMA_KEEP_ALIVE=-1
    One Per GPU

    Run one instance per available GPU.

    # 1 replica per GPU, not oversubscribed

Tips

  1. Set OLLAMA_NUM_PARALLEL above 1 so a single instance serves several requests at once, as long as memory allows the extra KV cache.
  2. Scale horizontally by running several Ollama containers and balancing them with an Nginx upstream block for round-robin routing.

Warnings

  1. By default OLLAMA_NUM_PARALLEL may be 1, so a single instance queues extra requests; raise it or add replicas for concurrency.
  2. Each parallel slot needs its own KV cache memory; setting OLLAMA_NUM_PARALLEL too high can exhaust VRAM and crash model loads.

In Practice

FAQ