Grounding With Web Search
Give local models live information with a web search tool, then ground answers in the results.
TL;DR
- Give the model a
search_webtool to fetch facts. - Inject live results into the prompt, then call
/api/chat. - Or use Ollama's built-in
web_searchAPI afterollama signin.
The Grounding Loop
Expose search_webOffer a web search function as a tool.
# tools: [ search_web ]Model Requests ItThe model asks to search when unsure.
# tool_calls -> search_web(query)Run The SearchYour app calls a real search API.
# results = search_web("...")Inject ResultsAdd snippets back as a tool message.
{ "role": "tool", "content": snippets }Built-in Web Search
ollama signinAuthenticate to enable hosted web search.
ollama signinweb_searchQuery the hosted search API directly.
ollama.web_search(query="latest news")web_fetchFetch a page's content by URL.
ollama.web_fetch(url="https://...")DIY Search Tool
Define The ToolWrap any search API as a function.
def search_web(query: str) -> str: ...Bind ItAttach the tool to a chat model.
llm.bind_tools([search_web])Return SnippetsFeed the results back for the answer.
# Append results, then call againKeep It Grounded
Cite SourcesAsk the model to use retrieved text only.
"prompt": "Answer using only the results"Trim NoisePass only the top relevant snippets.
# Keep the best 3 results, not allFresh Over MemoryPrefer live data over model recall.
# Instruct: prefer search over memoryTips
- Treat web search as a tool the model calls with
bind_toolsor atoolsarray, then feed the results back for grounding. - Insert only the most relevant snippets into the prompt, since dumping entire pages fills the context window and buries the answer.
Warnings
- Local models have a training cutoff; without a
search_webtool, they answer current-events questions with stale or made-up facts. - Ollama's hosted
web_searchproxies through ollama.com and needsollama signin; asearch_webfunction you write stays fully local.
In Practice
Expose a search_web tool, let the model call it, then feed results back so the answer uses live data.
- The model cannot know current weather from training alone.
search_webgives it a tool to fetch live data.- The model returns a
tool_callwith the search query. - Your app runs it and feeds results back for a grounded reply.
import ollama
def search_web(query: str) -> str:
# call your search API here
return "Oslo: 7C, light rain today"
msgs = [{"role": "user",
"content": "Weather in Oslo now?"}]
r = ollama.chat(model="llama3.1",
messages=msgs, tools=[search_web])
# Model returns a tool_call; run it,
# append the result, then chat again
print(r.message.tool_calls)FAQ
Expose a web search function as a tool. The model requests a search, your app runs it against a real search API, and you feed the results back into the prompt so the model answers from live data.
Yes. The hosted web_search and web_fetch APIs are available after ollama signin. They proxy through ollama.com, so they need an account and network access, unlike a local search tool you build yourself.
By putting real retrieved text in the prompt and instructing the model to answer only from it, you replace guesses from memory with current facts. The model summarizes evidence instead of inventing it.
Only the top few snippets. Too much fills the context window, raises latency, and dilutes the signal. Retrieve broadly, then keep the most relevant results before injecting.