Ollama CLI Commands

Master the core Ollama command line for running models, managing local weights, and chatting interactively.

TL;DR

  1. Download and chat with a model using ollama run.
  2. Audit local weights on disk with ollama list.
  3. Enter multi-line prompts by wrapping them in """ quotes.

Run And Chat

    ollama run

    Start an interactive chat, downloading the model first if needed.

    ollama run llama3.2
    One-Shot Prompt

    Pass a prompt inline to print one answer and exit.

    ollama run llama3.2 "Explain HTTP briefly"
    --verbose

    Show timing and tokens-per-second stats after each reply.

    ollama run llama3.2 --verbose
    /bye

    End the interactive session and return to the shell.

    >>> /bye

Manage Models

    ollama pull

    Download a model's weights without starting a chat.

    ollama pull mistral
    ollama list

    List local models with size and last-modified date.

    ollama list
    ollama ps

    Show models currently loaded in memory right now.

    ollama ps
    ollama show

    Print a model's parameters, template, and license.

    ollama show llama3.2
    ollama cp

    Copy a model under a new name for customization.

    ollama cp llama3.2 my-llama
    ollama rm

    Delete a model's weights from disk permanently.

    ollama rm mistral

Interactive Commands

    """ multi-line

    Wrap several lines in triple quotes to send one block.

    >>> """
    ... first line of the prompt
    ... second line of the prompt
    ... """
    /set parameter

    Change a runtime setting like temperature for this session.

    >>> /set parameter temperature 0.2
    /show info

    Display details about the model in the current session.

    >>> /show info
    /clear

    Clear the conversation context without leaving the session.

    >>> /clear

Tags And Custom Models

    model:tag

    Choose a specific size or variant with a tag.

    ollama run llama3.2:1b
    :latest default

    A bare name resolves to the latest tag automatically.

    ollama run mistral   # same as mistral:latest
    ollama create

    Build a custom model from a Modelfile recipe.

    ollama create my-bot -f Modelfile

Tips

  1. Run ollama pull ahead of time to cache weights, so the first ollama run starts chatting without a download wait.
  2. Use ollama ps to see loaded models and their memory use before starting another, which avoids an out-of-memory stall.

Warnings

  1. ollama rm deletes a model's weights from disk permanently; you must run ollama pull again to get the model back.
  2. A bare model name defaults to the :latest tag, which can change over time; pin a specific tag for reproducible results.

In Practice

FAQ