Ollama Fundamentals
Understand how Ollama runs open language models locally on your own hardware with no network calls.
TL;DR
- Run open models locally using the
ollamacommand line. - Send prompts to the local API at
localhost:11434. - Store weights under
~/.ollamaand run them offline.
Core Architecture
ollama --versionConfirm the engine is installed and print its version.
ollama --versionllama.cpp EngineOllama wraps the llama.cpp runtime and loads GGUF weights.
# GGUF models run on the bundled engine
ollama run llama3.2Model StorageDownloaded weights live under your home directory by default.
ls ~/.ollama/modelsLocal REST APIA background server exposes an HTTP API on port 11434.
curl http://localhost:11434/api/tagsRun Your First Model
ollama runDownload if needed, then open an interactive chat session.
ollama run llama3.2
>>> Explain DNS in one sentence.One-Shot PromptPass a prompt inline to get one answer and exit.
ollama run llama3.2 "Summarize TCP in 20 words"/byeLeave the interactive prompt and return to your shell.
>>> /byeWhy Run Models Locally
On-Device DataPrompts and responses stay on your machine by default.
# No API key, no outbound request
ollama run llama3.2 "Draft a private note"Offline UseOnce pulled, a model runs with no internet connection.
ollama pull mistral # download once
# then run fully offlineZero Token FeesLocal inference has no per-token billing after download.
# Unlimited local prompts, no usage charge
ollama run gemma3The Local API
/api/generateSend a single prompt and get back a completion.
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "Why is the sky blue?",
"stream": false
}'/api/chatSend a message list for multi-turn conversations.
curl http://localhost:11434/api/chat -d '{
"model": "llama3.2",
"messages": [
{"role": "user", "content": "Hi"}
]
}'OpenAI CompatibilityPoint OpenAI SDKs at the compatible base URL.
http://localhost:11434/v1/chat/completionsTips
- Start with a small model like
llama3.2to test your hardware before pulling larger models that need much more RAM. - Use
ollama psto see which models are loaded in memory, so you can free resources before running a bigger one.
Warnings
- Large models can exhaust your RAM or VRAM; check a model's size with
ollama listbefore running it on limited hardware. - The server binds to
127.0.0.1by default, so other machines cannot reach it until you changeOLLAMA_HOST.
In Practice
Pull a small model, ask a one-shot question, then query the same model through the local REST API.
- Pulling first caches the weights so later runs start fast and work offline.
- A one-shot
ollama runprints a single answer and returns to the shell. ollama psconfirms the model is resident, so you know memory is in use.- The same model answers HTTP requests, letting your apps reuse it.
# 1. Download a small, fast model
ollama pull llama3.2
# 2. Ask a one-shot question from the shell
ollama run llama3.2 "Name 3 uses for local LLMs"
# 3. Check what is loaded in memory
ollama ps
# 4. Call the same model over the local API
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "Say hello",
"stream": false
}'FAQ
Yes. Ollama is open source and needs no API key. Once you run ollama pull to download a model, it runs fully offline with no per-token cost.
No. Inference happens on your own machine, so prompts and responses never leave the device by default. The only network use is downloading models with ollama pull.
A small model like llama3.2:1b runs on a modern laptop with 8 GB of RAM. Larger models need more RAM, and a GPU with enough VRAM makes responses much faster.
Send an HTTP request to the local server. Use http://localhost:11434/api/chat, or point an OpenAI SDK at http://localhost:11434/v1 for a compatible endpoint.