Local RAG Pipeline
Build a local RAG pipeline that chunks, embeds, retrieves, and injects private text into a model prompt.
TL;DR
- Chunk text, embed each piece, and store it in a
vector DB. - Embed the query with
/api/embed, then search the DB. - Inject retrieved text into the prompt for
/api/chat.
The RAG Flow
ChunkSplit documents into small passages.
# Split docs into ~200-word chunksEmbedTurn each chunk into a vector locally.
# Embed chunks with nomic-embed-textStoreSave vectors and their text in a DB.
# Upsert (vector, text) into the DBRetrieveSearch the DB with the query vector.
# Top-k nearest vectors to the queryGenerateInject matches and call the chat model.
# Add context to the prompt, then chatIndex With ChromaDB
Create CollectionMake a collection to hold your vectors.
col = client.create_collection("docs")Add DocumentsStore text and its embedding together.
col.add(ids=ids, embeddings=vecs,
documents=texts)Chroma + OllamaUse Ollama embeddings as the vectors.
# Embed chunks via ollama.embed firstQuery And Retrieve
Embed QueryTurn the question into a vector.
q = ollama.embed(
model="nomic-embed-text", input=question)SearchFind the nearest stored chunks.
hits = col.query(query_embeddings=q,
n_results=3)Get TextPull the original passages from results.
context = hits["documents"][0]Augment The Prompt
Build ContextJoin retrieved passages into one block.
ctx = "\n".join(context)InjectAdd the context to the chat message.
msg = f"Context:\n{ctx}\n\nQ: {query}"AnswerCall the chat model with the context.
ollama.chat(model="llama3.1",
messages=[{"role":"user","content":msg}])Tips
- Store the original text alongside each vector, so retrieval returns the passage you inject, not just a numeric array.
- Chunk documents into small, overlapping passages, which keeps each embedding focused and improves retrieval accuracy for specific questions.
Warnings
- Use the same embedding model for indexing and querying; otherwise the query vector cannot be compared to stored vectors.
- Inject only the top retrieved chunks into the prompt; stuffing too much text overflows the
num_ctxwindow and slows replies.
In Practice
Embed two chunks, retrieve the closest to a question by similarity, and inject it into the model prompt.
- One batch call embeds both document chunks locally.
- The query is embedded with the same model for comparison.
- A dot product ranks chunks; the top match becomes context.
- That passage is injected so the model answers from it.
import ollama, numpy as np
# 1. Embed and store two chunks
docs = ["Refund window is 30 days.",
"Support hours are 9 to 5."]
E = ollama.embed(model="nomic-embed-text",
input=docs)["embeddings"]
# 2. Retrieve best match for the query
q = ollama.embed(model="nomic-embed-text",
input="How long for refunds?")
qv = np.array(q["embeddings"][0])
sims = [qv @ np.array(e) for e in E]
best = docs[int(np.argmax(sims))]
# 3. Inject and answer
msg = f"Context: {best}\nQ: Refund time?"
r = ollama.chat(model="llama3.1",
messages=[{"role":"user","content":msg}])
print(r.message.content)FAQ
RAG, or retrieval-augmented generation, finds relevant text and injects it into the prompt so the model answers from your data. It gives a model knowledge it was never trained on, without fine-tuning.
Any of them. ChromaDB and FAISS are common local choices. You generate vectors with Ollama's /api/embed, store them in the database, and query by similarity. Ollama handles embedding, not storage.
Your documents are embedded and stored locally. At query time, the closest passages are retrieved and placed in the prompt, so the model reasons over private text that stays entirely on your machine.
Only the top few chunks. Injecting too much fills the context window, raises latency, and can bury the answer. Retrieve more than you need, then keep the best three or so.