Few-Shot With MESSAGE

Use Modelfile MESSAGE directives to preload few-shot examples and lock in a strict output format.

TL;DR

  1. MESSAGE lines preload example turns into the model.
  2. Pair MESSAGE user with MESSAGE assistant to show format.
  3. Prime strict output without training a LoRA adapter.

The MESSAGE Directive

    MESSAGE user

    Add an example user turn to the history.

    MESSAGE user "Convert 5 to roman"
    MESSAGE assistant

    Add the ideal assistant reply to copy.

    MESSAGE assistant "V"
    MESSAGE system

    Seed a system role message when needed.

    MESSAGE system "Reply in roman numerals."
    Order Matters

    List pairs in the order the model should read.

    # user, then assistant, repeated

Prime A Format

    Show The Pattern

    Give one clean example of the exact output.

    MESSAGE user "3"
    MESSAGE assistant "III"
    Repeat For Strength

    Two or three pairs reinforce the format.

    MESSAGE user "9"
    MESSAGE assistant "IX"
    Strict Syntax

    Examples teach the exact symbols to emit.

    MESSAGE assistant "VII"

Few-Shot vs LoRA

    No Training

    MESSAGE primes behavior with zero training runs.

    # No dataset, no GPU training needed
    ADAPTER

    A LoRA adapter changes weights via a trained file.

    ADAPTER ./lora-adapter.gguf
    When To Train

    Use LoRA for deep style or knowledge shifts.

    # LoRA for big, permanent behavior change

Build And Test

    ollama create

    Bake the examples into a named model.

    ollama create roman-bot -f Modelfile
    Run It

    The model follows the primed format.

    ollama run roman-bot "7"
    Adjust Examples

    Edit the pairs and rebuild to refine.

    ollama create roman-bot -f Modelfile

Tips

  1. Add two or three MESSAGE example pairs to lock in a strict output format, which is faster than fine-tuning for many tasks.
  2. Keep few-shot examples short and consistent, since the model copies their style, spacing, and structure in its own replies.

Warnings

  1. Too many MESSAGE examples eat the context window and slow each request; a few strong examples usually beat a long list.
  2. Baked-in MESSAGE examples apply to every request; send runtime messages instead when the examples should change per call.

In Practice

FAQ