Ollama Native API

Call the native /api/generate and /api/chat endpoints, and handle Ollama's stateless, streaming design.

TL;DR

  1. Send a single prompt to the /api/generate endpoint.
  2. Resend the whole message history to /api/chat each call.
  3. Streaming is on by default; set stream: false to disable.

Generate Endpoint

    /api/generate

    Send one prompt and receive a completion.

    curl http://localhost:11434/api/generate -d '{
      "model": "llama3.2",
      "prompt": "Why is grass green?",
      "stream": false
    }'
    Single Turn

    Best for one-off prompts with no chat history.

    # No messages array, just a prompt
    Model Field

    Name the local model to run in the body.

    { "model": "llama3.2", "prompt": "Hi" }

Chat Endpoint

    /api/chat

    Send a messages array for multi-turn chat.

    curl http://localhost:11434/api/chat -d '{
      "model": "llama3.2",
      "messages": [
        {"role":"user","content":"Hello"}
      ]
    }'
    Roles

    Each message has a role and content field.

    { "role": "user", "content": "Hello" }
    Resend History

    Include all prior turns to keep context.

    # messages: [user, assistant, user, ...]

Streaming

    Default Stream

    Responses arrive as newline-delimited JSON chunks.

    # Each line is a partial JSON object
    Disable Stream

    Set stream false for one complete response.

    "stream": false
    Done Flag

    The final chunk carries a done field of true.

    { "done": true }

Statelessness

    No Server Memory

    The server keeps nothing between requests.

    # Every call is fully independent
    Client Holds State

    Your app stores and resends the conversation.

    # Append each reply to your messages
    generate Context

    generate can return a context array to reuse.

    # Pass returned context on the next call

Tips

  1. Track the conversation on the client and resend it in messages, since the server remembers nothing at all between requests.
  2. Set "stream": false when you want one complete JSON response instead of parsing a series of newline-delimited chunks.

Warnings

  1. The server is stateless; if you forget to resend prior messages, the model loses all context from the earlier turns.
  2. Responses stream as newline-delimited JSON by default, so code expecting one JSON object must pass "stream": false.

In Practice

FAQ