Choosing Ollama Models

Compare base and instruct models, match parameter counts to your RAM, and set the right context window.

TL;DR

  1. Pick instruct models for chat, base models for completion.
  2. Budget about 8GB of RAM for a 7B model.
  3. Raise the context window with /set parameter num_ctx per session.

Base vs Instruct

    Instruct Models

    Tuned to follow chat instructions and answer questions.

    ollama run llama3.1:8b   # instruct default
    Base Models

    Only continue raw text; not tuned for chat.

    ollama run llama3.1:8b-text-q4_K_M   # base
    Check The Tag

    Registry tags mark base variants, often with -text.

    ollama show llama3.1:8b-text-q4_K_M

Parameters And RAM

    Model Size

    More parameters mean better quality but more memory.

    # 3B is light; 8B mid; 70B needs a big GPU
    7B Rule Of Thumb

    A 7-8B model at 4-bit needs about 8 GB RAM.

    # ~5-6 GB weights + context overhead
    Check Disk Size

    The list command shows each model's size on disk.

    ollama list   # see the SIZE column

Context Window

    num_ctx

    Set the context window for the current session.

    >>> /set parameter num_ctx 8192
    Default Window

    Many setups default to a small 2048-token window.

    # Raise it when long inputs get truncated
    Global Default

    Set a server-wide default context length.

    export OLLAMA_CONTEXT_LENGTH=8192

Picking A Model

    Start Small

    A 3B instruct model is a fast first choice.

    ollama pull llama3.2:3b
    Specialized Models

    Use a coding model for programming tasks.

    ollama run qwen2.5-coder
    Inspect Details

    Show a model's parameters and context length.

    ollama show llama3.2:3b

Tips

  1. Choose an instruct model for assistants and chatbots, since base models only continue text and ignore chat-style instructions.
  2. Check the SIZE column in ollama list against your free RAM before pulling, so a model actually fits in memory.

Warnings

  1. A large context window raises memory use sharply; a big num_ctx can push a model off the GPU onto the slower CPU.
  2. Base models labeled -text are not chat-tuned; they ramble or repeat when given instructions meant for instruct models.

In Practice

FAQ