Choosing Ollama Models
Compare base and instruct models, match parameter counts to your RAM, and set the right context window.
TL;DR
- Pick
instructmodels for chat, base models for completion. - Budget about
8GBof RAM for a 7B model. - Raise the context window with
/set parameter num_ctxper session.
Base vs Instruct
Instruct ModelsTuned to follow chat instructions and answer questions.
ollama run llama3.1:8b # instruct defaultBase ModelsOnly continue raw text; not tuned for chat.
ollama run llama3.1:8b-text-q4_K_M # baseCheck The TagRegistry tags mark base variants, often with -text.
ollama show llama3.1:8b-text-q4_K_MParameters And RAM
Model SizeMore parameters mean better quality but more memory.
# 3B is light; 8B mid; 70B needs a big GPU7B Rule Of ThumbA 7-8B model at 4-bit needs about 8 GB RAM.
# ~5-6 GB weights + context overheadCheck Disk SizeThe list command shows each model's size on disk.
ollama list # see the SIZE columnContext Window
num_ctxSet the context window for the current session.
>>> /set parameter num_ctx 8192Default WindowMany setups default to a small 2048-token window.
# Raise it when long inputs get truncatedGlobal DefaultSet a server-wide default context length.
export OLLAMA_CONTEXT_LENGTH=8192Picking A Model
Start SmallA 3B instruct model is a fast first choice.
ollama pull llama3.2:3bSpecialized ModelsUse a coding model for programming tasks.
ollama run qwen2.5-coderInspect DetailsShow a model's parameters and context length.
ollama show llama3.2:3bTips
- Choose an
instructmodel for assistants and chatbots, since base models only continue text and ignore chat-style instructions. - Check the
SIZEcolumn inollama listagainst your free RAM before pulling, so a model actually fits in memory.
Warnings
- A large context window raises memory use sharply; a big
num_ctxcan push a model off the GPU onto the slower CPU. - Base models labeled
-textare not chat-tuned; they ramble or repeat when given instructions meant forinstructmodels.
In Practice
Audit local sizes, pull a small instruct model that fits 8 GB, then give it a larger context window.
ollama listshows sizes so you can compare them against free RAM.- A 3B instruct model fits comfortably on an 8 GB machine.
/set parameter num_ctxwidens the window for longer inputs.- The model answers a long prompt without truncating the earlier text.
# See how much memory each model needs
ollama list
# Pull a small instruct model for 8 GB RAM
ollama pull llama3.2:3b
# Give it a bigger context window
ollama run llama3.2:3b
>>> /set parameter num_ctx 8192
>>> Summarize this meeting transcript
>>> /byeFAQ
A base model only predicts the next token, so it continues text but ignores commands. An instruct model is fine-tuned to follow chat instructions, which is what you want for assistants and Q&A.
A 7-8B model at 4-bit needs roughly 5-6 GB just for weights. Plan for about 8 GB of RAM once you add the context cache and system overhead. Larger models need proportionally more.
Inside a session, run /set parameter num_ctx 8192. To change the server-wide default, set OLLAMA_CONTEXT_LENGTH. A bigger window remembers more conversation but uses more memory.
Start with a small instruct model like llama3.2:3b. It runs on modest hardware and responds quickly, so you can test your setup before pulling a larger model.