The Ollama Python Library

Use the official ollama Python library to chat, stream tokens, manage models, and call a remote host.

TL;DR

  1. Install the official client with pip install ollama.
  2. Call ollama.chat() with a messages list to get replies.
  3. Stream tokens as they arrive by passing stream=True.

Install And Chat

    Install

    Add the official Python client.

    pip install ollama
    chat

    Send messages and read the reply.

    import ollama
    r = ollama.chat(model="llama3.2",
      messages=[{"role":"user","content":"Hi"}])
    Read Content

    The reply text is on message.content.

    print(r["message"]["content"])
    generate

    Send a single prompt without roles.

    ollama.generate(model="llama3.2",
      prompt="Explain DNS")

Streaming

    stream=True

    Return a generator of partial chunks.

    stream = ollama.chat(model="llama3.2",
      messages=msgs, stream=True)
    Iterate Chunks

    Print each token as it arrives.

    for c in stream:
      print(c["message"]["content"], end="")
    generate Stream

    Chunks expose response for generate.

    # c["response"] for generate streams

Manage Models

    list

    List locally installed models.

    ollama.list()
    pull

    Download a model from Python.

    ollama.pull("llama3.2")
    embed

    Create embeddings from text.

    ollama.embed(model="nomic-embed-text",
      input="hello")
    show

    Inspect a model's details.

    ollama.show("llama3.2")

Clients

    Client

    Point the client at a specific host.

    from ollama import Client
    c = Client(host="http://box:11434")
    AsyncClient

    Use async for concurrent requests.

    from ollama import AsyncClient
    r = await AsyncClient().chat(...)
    tools

    Pass Python functions as tools.

    ollama.chat(model="llama3.1",
      messages=msgs, tools=[get_weather])

Tips

  1. Iterate the generator from stream=True and print chunk['message']['content'], so output appears token by token like a live chat.
  2. Use AsyncClient for concurrent requests in async apps, or Client(host=...) to target a remote Ollama server by URL.

Warnings

  1. The client needs a running server; start ollama serve or the desktop app, or calls fail to reach localhost:11434.
  2. ollama.chat is stateless like the API; resend the full messages list on each call to keep conversation context.

In Practice

FAQ