Vision Models With Ollama

Send images to local vision models like llava for captioning, UI analysis, and reading text.

TL;DR

  1. Send images as Base64 strings to a vision model.
  2. Use a vision model like llava or llama3.2-vision.
  3. Pass an images array alongside your text prompt.

Vision Basics

    Pull A Vision Model

    Download a model with a vision architecture.

    ollama pull llava
    images Field

    Attach Base64 images to the request.

    # body: { model, prompt, images: [b64] }
    Ask About It

    Pair the image with a specific question.

    "prompt": "Describe this screenshot"

Encode Images

    Base64 Encode

    Read a file and encode it to Base64.

    import base64
    b64 = base64.b64encode(data).decode()
    Python Client

    The ollama client takes file paths directly.

    ollama.chat(model="llava",
      messages=[{"role": "user", ...}])
    images In Message

    Add an images list to the message.

    "images": ["screenshot.png"]

cURL Example

    generate With Image

    Send a Base64 image to /api/generate.

    curl http://localhost:11434/api/generate -d '{
      "model": "llava",
      "prompt": "What is in this image?",
      "images": ["<base64>"],
      "stream": false
    }'
    Base64 String

    The images array holds encoded strings.

    "images": ["iVBORw0KGgo..."]
    stream false

    Get one response for scripting.

    "stream": false

Use Cases

    UI Analysis

    Describe or audit a UI screenshot automatically.

    "prompt": "List the buttons in this UI"
    Captioning

    Generate alt text for images in bulk.

    "prompt": "Write concise alt text"
    OCR-Like Reading

    Read visible text from an image.

    "prompt": "Transcribe the text shown"

Tips

  1. Use the ollama Python client, which accepts image file paths in images and Base64-encodes them for you automatically.
  2. Ask a specific question about the image, such as extracting text or listing UI elements, to get a focused, useful answer.

Warnings

  1. A text-only model ignores the images field; the model must have a vision architecture like llava to see pictures.
  2. Large images inflate token use and slow inference; resize or compress them before Base64-encoding for faster responses.

In Practice

FAQ