Guardrails For Small Models

Constrain small local models with rigid prompts, stop tokens, and examples to stop rambling and filler.

TL;DR

  1. Give small models rigid, explicit rules in the SYSTEM prompt.
  2. Cut rambling early with a PARAMETER stop sequence.
  3. Lower randomness with PARAMETER temperature 0 for steady output.

Write Rigid System Prompts

    Be Explicit

    State the exact task and format precisely.

    SYSTEM """Answer in one sentence only."""
    Ban Filler

    Forbid preambles and sign-offs directly.

    SYSTEM """No preamble. Output only code."""
    One Job

    Give the model a single, narrow task.

    SYSTEM """Return valid JSON. Nothing else."""

Constrain Output

    temperature 0

    Remove randomness for consistent formatting.

    PARAMETER temperature 0
    stop

    End generation at a known boundary token.

    PARAMETER stop "```"
    num_predict

    Cap tokens to prevent runaway output.

    PARAMETER num_predict 200

Prime With Examples

    MESSAGE pair

    Show one ideal question and answer.

    MESSAGE user "2+2"
    MESSAGE assistant "4"
    Match The Format

    Make the example output the exact target shape.

    MESSAGE assistant """{"ok": true}"""
    Keep It Short

    One or two pairs is enough for small models.

    # 1-2 pairs, not ten

Structure The Prompt

    Linear Steps

    Ask for one action at a time.

    ollama run llama3.1:8b "List 3 fruits"
    Explicit Format

    Tell the model exactly how to format.

    "Reply as a numbered list, no intro"
    Test And Tighten

    Iterate on wording until output is stable.

    # Rerun, then tighten the SYSTEM rule

Tips

  1. Spell out the exact output format for small models, since an 8B model follows rigid, concrete rules far better than vague ones.
  2. Show one or two MESSAGE examples of the ideal answer, which anchors the format better than description alone for small models.

Warnings

  1. Small models drift with long, multi-part prompts; split complex tasks into single, linear steps to keep them on track.
  2. Without a stop sequence, a small model may add filler like "Sure, here is" or keep talking well past the answer.

In Practice

FAQ