Ollama With Docker
Run Ollama in Docker with a persistent volume for models and GPU passthrough for fast inference.
TL;DR
- Run Ollama in Docker with the
ollama/ollamaimage. - Mount a volume at
/root/.ollamaso weights persist. - Add
--gpus=allfor GPU access via the NVIDIA toolkit.
Run The Container
Pull ImageGet the official Ollama Docker image.
docker pull ollama/ollamaRun DetachedStart the server mapped to port 11434.
docker run -d -p 11434:11434 ollama/ollamaExec A ModelRun a model inside the container.
docker exec -it ollama ollama run llama3.2Persist Weights
Named VolumeMount a volume so models persist.
-v ollama:/root/.ollamaHost PathOr bind a host directory for weights.
-v /data/ollama:/root/.ollamaWhy It MattersWithout it, restarts lose all models.
# No volume = re-pull every restartGPU Access
NVIDIA ToolkitInstall the toolkit on the host first.
# Install nvidia-container-toolkit--gpus=allExpose all GPUs to the container.
docker run --gpus=all ollama/ollamaVerify GPUCheck the container sees the GPU.
docker exec ollama nvidia-smiFull Command
Everything TogetherVolume, port, and GPU in one command.
docker run -d --gpus=all \
-v ollama:/root/.ollama \
-p 11434:11434 ollama/ollamaName ItGive the container a stable name.
--name ollamaKubernetesMount a PVC and request GPU in the pod.
# PVC at /root/.ollama + GPU limitsTips
- Mount a named volume at
/root/.ollama, so pulled models survive container restarts instead of downloading again every time. - Install the NVIDIA Container Toolkit first, then pass
--gpus=all, so the container can actually see and use your GPU.
Warnings
- Without a volume at
/root/.ollama, every container restart loses downloaded models and has to re-pull them from scratch. - Omitting
--gpus=allmakes the container run on CPU only; the NVIDIA Container Toolkit must also be installed on the host.
In Practice
Start the official image with a persistent volume and full GPU access, then run a model inside it.
- The NVIDIA Container Toolkit is a one-time host prerequisite.
- The volume at
/root/.ollamakeeps models across restarts. --gpus=allgives the container access to the host GPUs.nvidia-smiinside the container confirms the GPU is visible.
# 1. Install the NVIDIA Container Toolkit
# (one-time host setup)
# 2. Start Ollama with volume + GPU
docker run -d --gpus=all \
-v ollama:/root/.ollama \
-p 11434:11434 \
--name ollama ollama/ollama
# 3. Pull and run a model in the container
docker exec -it ollama ollama run llama3.2
# 4. Confirm the GPU is visible
docker exec ollama nvidia-smiFAQ
Run the official image: docker run -d -p 11434:11434 ollama/ollama. Then run a model inside it with docker exec -it ollama ollama run llama3.2. Add a volume and GPU flags for real use.
Mount a persistent volume at /root/.ollama, where Ollama stores weights. With -v ollama:/root/.ollama, models survive restarts instead of being pulled again each time the container starts.
Install the NVIDIA Container Toolkit on the host, then add --gpus=all to the run command. Verify it worked by running nvidia-smi inside the container.
Use the same image in a Deployment, mount a PersistentVolumeClaim at /root/.ollama, and request GPU resources in the pod spec. The container behaves the same as under Docker.