Ollama Fundamentals

Understand how Ollama runs open language models locally on your own hardware with no network calls.

TL;DR

  1. Run open models locally using the ollama command line.
  2. Send prompts to the local API at localhost:11434.
  3. Store weights under ~/.ollama and run them offline.

Core Architecture

    ollama --version

    Confirm the engine is installed and print its version.

    ollama --version
    llama.cpp Engine

    Ollama wraps the llama.cpp runtime and loads GGUF weights.

    # GGUF models run on the bundled engine
    ollama run llama3.2
    Model Storage

    Downloaded weights live under your home directory by default.

    ls ~/.ollama/models
    Local REST API

    A background server exposes an HTTP API on port 11434.

    curl http://localhost:11434/api/tags

Run Your First Model

    ollama run

    Download if needed, then open an interactive chat session.

    ollama run llama3.2
    >>> Explain DNS in one sentence.
    One-Shot Prompt

    Pass a prompt inline to get one answer and exit.

    ollama run llama3.2 "Summarize TCP in 20 words"
    /bye

    Leave the interactive prompt and return to your shell.

    >>> /bye

Why Run Models Locally

    On-Device Data

    Prompts and responses stay on your machine by default.

    # No API key, no outbound request
    ollama run llama3.2 "Draft a private note"
    Offline Use

    Once pulled, a model runs with no internet connection.

    ollama pull mistral   # download once
    # then run fully offline
    Zero Token Fees

    Local inference has no per-token billing after download.

    # Unlimited local prompts, no usage charge
    ollama run gemma3

The Local API

    /api/generate

    Send a single prompt and get back a completion.

    curl http://localhost:11434/api/generate -d '{
      "model": "llama3.2",
      "prompt": "Why is the sky blue?",
      "stream": false
    }'
    /api/chat

    Send a message list for multi-turn conversations.

    curl http://localhost:11434/api/chat -d '{
      "model": "llama3.2",
      "messages": [
        {"role": "user", "content": "Hi"}
      ]
    }'
    OpenAI Compatibility

    Point OpenAI SDKs at the compatible base URL.

    http://localhost:11434/v1/chat/completions

Tips

  1. Start with a small model like llama3.2 to test your hardware before pulling larger models that need much more RAM.
  2. Use ollama ps to see which models are loaded in memory, so you can free resources before running a bigger one.

Warnings

  1. Large models can exhaust your RAM or VRAM; check a model's size with ollama list before running it on limited hardware.
  2. The server binds to 127.0.0.1 by default, so other machines cannot reach it until you change OLLAMA_HOST.

In Practice

FAQ