Ollama In Your IDE
Connect editor tools like Continue.dev to a local Ollama server for private, free code completion.
TL;DR
- Point IDE extensions at the local
11434port. - Pick a coder model like
qwen2.5-coderfor autocomplete. - Set
apiBasetolocalhost:11434in the extension config.
Prepare Ollama
Pull A CoderDownload a coding model for the editor.
ollama pull qwen2.5-coderKeep It LoadedPin the model so autocomplete stays fast.
export OLLAMA_KEEP_ALIVE=-1Confirm It RunsCheck the server responds on its port.
curl http://localhost:11434Continue.dev Config
Model EntryAdd a model block to ~/.continue/config.yaml.
models:
- name: Ollama Coder
provider: ollama
model: qwen2.5-coder
apiBase: http://localhost:11434providerTell Continue to use the Ollama provider.
provider: ollamaapiBasePoint the block at the local server URL.
apiBase: http://localhost:11434Cursor And Others
OpenAI Base URLCursor needs an OpenAI-compatible endpoint.
# Set base URL to the /v1 path in settingsPublic TunnelExpose localhost, since Cursor calls from its cloud.
# A tunnel maps a public URL to :11434Prefer Local ToolsContinue.dev runs fully local, no tunnel.
# No cloud round-trip with Continue.devPick A Model
AutocompleteA small coder model gives instant suggestions.
ollama pull qwen2.5-coder:1.5bChat And EditsA larger model handles refactors and Q&A.
ollama pull qwen2.5-coder:7bGeneral Codercodellama is a solid all-round choice.
ollama pull codellamaTips
- Use a fast small coder model like
qwen2.5-coder:1.5bfor tab autocomplete, and a larger one for chat and refactors. - Keep
ollama serverunning and pin the model withOLLAMA_KEEP_ALIVE=-1, so suggestions stay instant while you type.
Warnings
- Continue.dev talks to Ollama over its
HTTPAPI, not the Language Server Protocol; there is no LSP bridge to configure. - Cursor proxies requests through its cloud, so a
localhostOllama needs a public tunnel or it cannot be reached.
In Practice
Pull a coder model, keep it loaded, and point Continue.dev at the local server in config.yaml.
- A small coder model keeps tab suggestions near-instant.
- Pinning it in memory avoids a reload pause mid-typing.
- The config block routes Continue to the local server.
- A restart loads the config; no LSP or cloud key is needed.
# 1. Pull a fast coder model
ollama pull qwen2.5-coder:1.5b
# 2. Keep it loaded for instant suggestions
export OLLAMA_KEEP_ALIVE=-1
# 3. Add this to ~/.continue/config.yaml:
# models:
# - name: Ollama Coder
# provider: ollama
# model: qwen2.5-coder:1.5b
# apiBase: http://localhost:11434
# 4. Restart the editor, then verify
curl http://localhost:11434FAQ
Install the Continue.dev extension, then add a model with provider: ollama and apiBase: http://localhost:11434 to ~/.continue/config.yaml. Pull a coder model first and restart the editor.
A small, fast coder model like qwen2.5-coder:1.5b gives near-instant tab suggestions. Use a larger model such as qwen2.5-coder:7b for chat, explanations, and refactors where quality matters more.
Partly. Cursor needs an OpenAI-compatible URL, but it routes requests through its own cloud, so a local server must be exposed with a public tunnel. Continue.dev is the better fully-local option.
No. Extensions like Continue.dev call Ollama's HTTP API directly at http://localhost:11434. There is no LSP involved, so you only configure a base URL and a model name.