Modelfile Prompt Templates
Customize the Modelfile TEMPLATE directive and control tokens so a model's chat format works correctly.
TL;DR
TEMPLATEcontrols the exact prompt sent to the model.- Insert content with
.System,.Prompt, and.Responsevariables. - Match control tokens like
<|im_end|>to the model's format.
The TEMPLATE Directive
TEMPLATEDefine the full prompt string sent to the model.
TEMPLATE """{{ .Prompt }}""".SystemInsert the system prompt into the template.
{{ if .System }}{{ .System }}{{ end }}.PromptInsert the user's message into the template.
{{ .Prompt }}.ResponseMarks where the model's reply begins.
{{ .Response }}Control Tokens
Special TokensModels use tokens to mark turn boundaries.
<|im_start|>user
{{ .Prompt }}<|im_end|>Match TrainingUse the exact tokens the model was trained on.
# Wrong tokens break the chat formatStop SequencesSet a stop token so generation ends cleanly.
PARAMETER stop "<|im_end|>"A Full Template
System BlockRender the system prompt only when present.
{{ if .System }}<|system|>
{{ .System }}
{{ end }}User BlockWrap the user prompt in the model's tags.
<|user|>
{{ .Prompt }}Assistant TagEnd with the assistant tag to cue the reply.
<|assistant|>View And Edit
Show TemplatePrint the model's current Modelfile and template.
ollama show llama3.2 --modelfileEdit And RebuildChange the template, then rebuild the model.
ollama create my-model -f ModelfileTest ItRun the model and check the formatting.
ollama run my-model "Say hi"Tips
- Start from the model's official template with
ollama show --modelfile, then edit that rather than writing aTEMPLATEfrom scratch. - Keep any special control tokens exactly as the model was trained, since even a small change can break the chat format.
Warnings
- A broken
TEMPLATEcan make a model ramble, loop, or ignore the system prompt; verify the format before relying on the model. - Text placed after
.Responsein aTEMPLATEis dropped during generation, so keep the assistant tag before that point.
In Practice
Export a model's template, correct the control tokens, rebuild it, and confirm the chat format works.
- Exporting the real Modelfile gives you a correct starting point.
- Editing the tokens in place avoids typos that break the format.
ollama createbakes the corrected template into a new model.- A quick test run confirms the reply ends cleanly without looping.
# 1. Save the current Modelfile
ollama show llama3.2 --modelfile > Modelfile
# 2. Edit the TEMPLATE block to match tokens,
# e.g. <|start|> ... <|end|> per the model
# 3. Rebuild with the corrected template
ollama create llama3-fixed -f Modelfile
# 4. Verify the reply is clean, no loops
ollama run llama3-fixed "Hello"FAQ
TEMPLATE defines the full prompt string Ollama passes to the model. It uses Go template syntax to place the system prompt, the user message, and control tokens in the exact layout the model expects.
The core variables are {{ .System }} for the system prompt, {{ .Prompt }} for the user message, and {{ .Response }} for where the reply begins. You wrap them in {{ if }} blocks to include parts only when present.
The TEMPLATE usually has the wrong control tokens or a missing stop sequence. Export the model's original template, compare the tokens, and add a matching PARAMETER stop so generation ends cleanly.
Run ollama show to print the full Modelfile, including the TEMPLATE block. Copy that as your starting point instead of guessing the format.