Ollama Native API
Call the native /api/generate and /api/chat endpoints, and handle Ollama's stateless, streaming design.
TL;DR
- Send a single prompt to the
/api/generateendpoint. - Resend the whole message history to
/api/chateach call. - Streaming is on by default; set
stream: falseto disable.
Generate Endpoint
/api/generateSend one prompt and receive a completion.
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "Why is grass green?",
"stream": false
}'Single TurnBest for one-off prompts with no chat history.
# No messages array, just a promptModel FieldName the local model to run in the body.
{ "model": "llama3.2", "prompt": "Hi" }Chat Endpoint
/api/chatSend a messages array for multi-turn chat.
curl http://localhost:11434/api/chat -d '{
"model": "llama3.2",
"messages": [
{"role":"user","content":"Hello"}
]
}'RolesEach message has a role and content field.
{ "role": "user", "content": "Hello" }Resend HistoryInclude all prior turns to keep context.
# messages: [user, assistant, user, ...]Streaming
Default StreamResponses arrive as newline-delimited JSON chunks.
# Each line is a partial JSON objectDisable StreamSet stream false for one complete response.
"stream": falseDone FlagThe final chunk carries a done field of true.
{ "done": true }Statelessness
No Server MemoryThe server keeps nothing between requests.
# Every call is fully independentClient Holds StateYour app stores and resends the conversation.
# Append each reply to your messagesgenerate Contextgenerate can return a context array to reuse.
# Pass returned context on the next callTips
- Track the conversation on the client and resend it in
messages, since the server remembers nothing at all between requests. - Set
"stream": falsewhen you want one complete JSON response instead of parsing a series of newline-delimited chunks.
Warnings
- The server is stateless; if you forget to resend prior
messages, the model loses all context from the earlier turns. - Responses stream as newline-delimited JSON by default, so code expecting one JSON object must pass
"stream": false.
In Practice
Call /api/chat, then send a second request that includes the prior turns so context carries over.
- The first call sends only the opening user message.
- The server replies but stores nothing about the exchange.
- The second call resends both prior turns plus the new question.
- Because the history is included, the model answers with the name.
# First turn
curl http://localhost:11434/api/chat -d '{
"model": "llama3.2",
"messages": [
{"role":"user","content":"My name is Sam"}
]
}'
# Second turn: resend prior turns + the new one
curl http://localhost:11434/api/chat -d '{
"model": "llama3.2",
"messages": [
{"role":"user","content":"My name is Sam"},
{"role":"assistant","content":"Hi Sam!"},
{"role":"user","content":"What is my name?"}
]
}'FAQ
No. The API is stateless, so it keeps no memory between requests. To continue a conversation, your client must include the full messages history on every call to /api/chat.
/api/generate takes a single prompt for one-off completions. /api/chat takes a messages array with roles, which is the right choice for multi-turn conversations where context matters.
Add "stream": false to the request body. The server then returns one complete JSON object instead of a stream of newline-delimited partial objects that you assemble yourself.
Append each assistant reply to your local messages list, then send the whole list on the next /api/chat request. The server has no memory, so the history you send is the only context.