OpenAI-Compatible Endpoint

Drop local models into existing OpenAI apps by pointing the SDK at Ollama's /v1 compatible endpoint.

TL;DR

  1. Ollama serves an OpenAI-compatible API under the /v1 path.
  2. Point the OpenAI SDK base_url at the local /v1 path.
  3. Pass any non-empty string as the api_key value.

The /v1 Endpoint

    Base URL

    The OpenAI-compatible path lives under /v1.

    http://localhost:11434/v1
    Chat Completions

    The standard chat completions route works locally.

    POST /v1/chat/completions
    Other Routes

    Completions, embeddings, and models are available.

    /v1/completions  /v1/embeddings  /v1/models

Python SDK

    Point The Client

    Set base_url and a placeholder api_key.

    from openai import OpenAI
    
    client = OpenAI(
        base_url="http://localhost:11434/v1",
        api_key="ollama",
    )
    Call A Local Model

    Use a local model name in the request.

    res = client.chat.completions.create(
        model="llama3.2",
        messages=[
            {"role": "user", "content": "Hi"}
        ])
    Read The Reply

    Parse the response like any OpenAI call.

    print(res.choices[0].message.content)

cURL Example

    Chat Request

    Call the endpoint directly with curl.

    curl http://localhost:11434/v1/chat/completions \
      -H 'Content-Type: application/json' \
      -d '{"model":"llama3.2","messages":
      [{"role":"user","content":"Hi"}]}'
    Authorization

    Send any bearer token in the header.

    -H "Authorization: Bearer ollama"
    Model Name

    Use your local model, not a cloud name.

    "model": "llama3.2"

Drop-In Migration

    Change base_url

    Repoint an existing client to localhost.

    base_url="http://localhost:11434/v1"
    Dummy Key

    Any non-empty api_key satisfies the SDK.

    api_key="ollama"
    Keep The Rest

    Messages and parsing code stay the same.

    # No other code changes needed

Tips

  1. Reuse existing OpenAI SDK code by changing only base_url and api_key, so a local model drops into your app unchanged.
  2. Set api_key to any placeholder like "ollama", since the local server ignores the value but the SDK still requires one.

Warnings

  1. Some OpenAI-only parameters are ignored or unsupported on /v1; test features like function calling per individual model.
  2. The model must name a local Ollama model, not a cloud name like gpt-4o, or the request comes back with an error.

In Practice

FAQ