The Ollama Python Library
Use the official ollama Python library to chat, stream tokens, manage models, and call a remote host.
TL;DR
- Install the official client with
pip install ollama. - Call
ollama.chat()with a messages list to get replies. - Stream tokens as they arrive by passing
stream=True.
Install And Chat
InstallAdd the official Python client.
pip install ollamachatSend messages and read the reply.
import ollama
r = ollama.chat(model="llama3.2",
messages=[{"role":"user","content":"Hi"}])Read ContentThe reply text is on message.content.
print(r["message"]["content"])generateSend a single prompt without roles.
ollama.generate(model="llama3.2",
prompt="Explain DNS")Streaming
stream=TrueReturn a generator of partial chunks.
stream = ollama.chat(model="llama3.2",
messages=msgs, stream=True)Iterate ChunksPrint each token as it arrives.
for c in stream:
print(c["message"]["content"], end="")generate StreamChunks expose response for generate.
# c["response"] for generate streamsManage Models
listList locally installed models.
ollama.list()pullDownload a model from Python.
ollama.pull("llama3.2")embedCreate embeddings from text.
ollama.embed(model="nomic-embed-text",
input="hello")showInspect a model's details.
ollama.show("llama3.2")Clients
ClientPoint the client at a specific host.
from ollama import Client
c = Client(host="http://box:11434")AsyncClientUse async for concurrent requests.
from ollama import AsyncClient
r = await AsyncClient().chat(...)toolsPass Python functions as tools.
ollama.chat(model="llama3.1",
messages=msgs, tools=[get_weather])Tips
- Iterate the generator from
stream=Trueand printchunk['message']['content'], so output appears token by token like a live chat. - Use
AsyncClientfor concurrent requests in async apps, orClient(host=...)to target a remote Ollama server by URL.
Warnings
- The client needs a running server; start
ollama serveor the desktop app, or calls fail to reachlocalhost:11434. ollama.chatis stateless like the API; resend the fullmessageslist on each call to keep conversation context.
In Practice
Send a prompt with the Python client and print the reply token by token as it streams in.
pip install ollamagives the official client.stream=Truemakeschatreturn a generator of chunks.- Each chunk holds a slice of text on
message.content. - Printing with
end=""streams the reply smoothly.
import ollama
stream = ollama.chat(
model="llama3.2",
messages=[{"role": "user",
"content": "Explain RAG in 3 lines"}],
stream=True,
)
for chunk in stream:
print(chunk["message"]["content"],
end="", flush=True)FAQ
Run pip install ollama with Python 3.8 or newer. Then import ollama and call ollama.chat(...). A local Ollama server must be running for the calls to succeed.
Pass stream=True to ollama.chat. It returns a generator; iterate it and print chunk['message']['content'] with end='' so tokens appear as they arrive instead of all at once.
ollama.chat takes a messages list with roles for multi-turn conversation. ollama.generate takes a single prompt for one-off completions. Chat streams expose message.content; generate streams expose response.
Create a client with a host: Client(host="http://box:11434"), then call .chat(...) on it. Use AsyncClient the same way inside async code for concurrent requests.