OpenAI-Compatible Endpoint
Drop local models into existing OpenAI apps by pointing the SDK at Ollama's /v1 compatible endpoint.
TL;DR
- Ollama serves an OpenAI-compatible API under the
/v1path. - Point the OpenAI SDK
base_urlat the local/v1path. - Pass any non-empty string as the
api_keyvalue.
The /v1 Endpoint
Base URLThe OpenAI-compatible path lives under /v1.
http://localhost:11434/v1Chat CompletionsThe standard chat completions route works locally.
POST /v1/chat/completionsOther RoutesCompletions, embeddings, and models are available.
/v1/completions /v1/embeddings /v1/modelsPython SDK
Point The ClientSet base_url and a placeholder api_key.
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama",
)Call A Local ModelUse a local model name in the request.
res = client.chat.completions.create(
model="llama3.2",
messages=[
{"role": "user", "content": "Hi"}
])Read The ReplyParse the response like any OpenAI call.
print(res.choices[0].message.content)cURL Example
Chat RequestCall the endpoint directly with curl.
curl http://localhost:11434/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"llama3.2","messages":
[{"role":"user","content":"Hi"}]}'AuthorizationSend any bearer token in the header.
-H "Authorization: Bearer ollama"Model NameUse your local model, not a cloud name.
"model": "llama3.2"Drop-In Migration
Change base_urlRepoint an existing client to localhost.
base_url="http://localhost:11434/v1"Dummy KeyAny non-empty api_key satisfies the SDK.
api_key="ollama"Keep The RestMessages and parsing code stay the same.
# No other code changes neededTips
- Reuse existing OpenAI SDK code by changing only
base_urlandapi_key, so a local model drops into your app unchanged. - Set
api_keyto any placeholder like"ollama", since the local server ignores the value but the SDK still requires one.
Warnings
- Some OpenAI-only parameters are ignored or unsupported on
/v1; test features like function calling per individual model. - The
modelmust name a local Ollama model, not a cloud name likegpt-4o, or the request comes back with an error.
In Practice
Repoint an existing OpenAI Python client at Ollama by changing only the base URL and key.
- Only the
base_urlandapi_keylines differ from cloud code. - The local server ignores the key, so a placeholder is fine.
- A local
modelname routes the call to Ollama. - Response parsing uses the same
choicesshape as OpenAI.
from openai import OpenAI
# Only these two lines change
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama",
)
res = client.chat.completions.create(
model="llama3.2",
messages=[
{"role": "user", "content": "Hello"}
])
print(res.choices[0].message.content)FAQ
Yes. Ollama exposes an OpenAI-compatible API at /v1. Point the SDK's base_url at http://localhost:11434/v1 and pass any api_key, then call it like the real OpenAI API.
Use http://localhost:11434/v1 as the base_url. The api_key can be any non-empty string, such as "ollama", because the local server does not check it.
The common ones: /v1/chat/completions, /v1/completions, /v1/embeddings, and /v1/models. Advanced or provider-specific features may not be supported, so verify the ones your app relies on.
Barely. Change base_url, set a placeholder api_key, and use a local model name. Your message construction and response parsing stay exactly the same.