Vision Models With Ollama
Send images to local vision models like llava for captioning, UI analysis, and reading text.
TL;DR
- Send
imagesas Base64 strings to a vision model. - Use a vision model like
llavaorllama3.2-vision. - Pass an
imagesarray alongside your textprompt.
Vision Basics
Pull A Vision ModelDownload a model with a vision architecture.
ollama pull llavaimages FieldAttach Base64 images to the request.
# body: { model, prompt, images: [b64] }Ask About ItPair the image with a specific question.
"prompt": "Describe this screenshot"Encode Images
Base64 EncodeRead a file and encode it to Base64.
import base64
b64 = base64.b64encode(data).decode()Python ClientThe ollama client takes file paths directly.
ollama.chat(model="llava",
messages=[{"role": "user", ...}])images In MessageAdd an images list to the message.
"images": ["screenshot.png"]cURL Example
generate With ImageSend a Base64 image to /api/generate.
curl http://localhost:11434/api/generate -d '{
"model": "llava",
"prompt": "What is in this image?",
"images": ["<base64>"],
"stream": false
}'Base64 StringThe images array holds encoded strings.
"images": ["iVBORw0KGgo..."]stream falseGet one response for scripting.
"stream": falseUse Cases
UI AnalysisDescribe or audit a UI screenshot automatically.
"prompt": "List the buttons in this UI"CaptioningGenerate alt text for images in bulk.
"prompt": "Write concise alt text"OCR-Like ReadingRead visible text from an image.
"prompt": "Transcribe the text shown"Tips
- Use the
ollamaPython client, which accepts image file paths inimagesand Base64-encodes them for you automatically. - Ask a specific question about the image, such as extracting text or listing UI elements, to get a focused, useful answer.
Warnings
- A text-only model ignores the
imagesfield; the model must have a vision architecture likellavato see pictures. - Large images inflate token use and slow inference; resize or compress them before Base64-encoding for faster responses.
In Practice
Use the ollama Python client to send a screenshot to a vision model and get a caption back.
- The ollama client reads the image path and Base64-encodes it.
llavahas a vision architecture, so it can see the image.- The content question focuses the model on what you need.
- The reply is a text caption describing the screenshot.
import ollama
resp = ollama.chat(
model="llava",
messages=[{
"role": "user",
"content": "Describe this screenshot",
"images": ["screenshot.png"],
}],
)
print(resp.message.content)FAQ
Include an images array of Base64-encoded strings in the request, alongside your prompt. With the ollama Python client you can pass file paths directly and it encodes them for you.
Vision models such as llava, llama3.2-vision, and moondream. A text-only model has no vision architecture, so it cannot interpret an image even if you send one.
Yes. Send a screenshot and ask the model to list elements, describe layout, or spot issues. It works well for local, automated UI and accessibility checks without any cloud service.
Most likely you are using a text-only model. Switch to a vision model like llava. Also confirm the image is valid Base64 and sits in the images field, not the prompt text.