Local RAG Pipeline

Build a local RAG pipeline that chunks, embeds, retrieves, and injects private text into a model prompt.

TL;DR

  1. Chunk text, embed each piece, and store it in a vector DB.
  2. Embed the query with /api/embed, then search the DB.
  3. Inject retrieved text into the prompt for /api/chat.

The RAG Flow

    Chunk

    Split documents into small passages.

    # Split docs into ~200-word chunks
    Embed

    Turn each chunk into a vector locally.

    # Embed chunks with nomic-embed-text
    Store

    Save vectors and their text in a DB.

    # Upsert (vector, text) into the DB
    Retrieve

    Search the DB with the query vector.

    # Top-k nearest vectors to the query
    Generate

    Inject matches and call the chat model.

    # Add context to the prompt, then chat

Index With ChromaDB

    Create Collection

    Make a collection to hold your vectors.

    col = client.create_collection("docs")
    Add Documents

    Store text and its embedding together.

    col.add(ids=ids, embeddings=vecs,
            documents=texts)
    Chroma + Ollama

    Use Ollama embeddings as the vectors.

    # Embed chunks via ollama.embed first

Query And Retrieve

    Embed Query

    Turn the question into a vector.

    q = ollama.embed(
        model="nomic-embed-text", input=question)
    Search

    Find the nearest stored chunks.

    hits = col.query(query_embeddings=q,
                     n_results=3)
    Get Text

    Pull the original passages from results.

    context = hits["documents"][0]

Augment The Prompt

    Build Context

    Join retrieved passages into one block.

    ctx = "\n".join(context)
    Inject

    Add the context to the chat message.

    msg = f"Context:\n{ctx}\n\nQ: {query}"
    Answer

    Call the chat model with the context.

    ollama.chat(model="llama3.1",
      messages=[{"role":"user","content":msg}])

Tips

  1. Store the original text alongside each vector, so retrieval returns the passage you inject, not just a numeric array.
  2. Chunk documents into small, overlapping passages, which keeps each embedding focused and improves retrieval accuracy for specific questions.

Warnings

  1. Use the same embedding model for indexing and querying; otherwise the query vector cannot be compared to stored vectors.
  2. Inject only the top retrieved chunks into the prompt; stuffing too much text overflows the num_ctx window and slows replies.

In Practice

FAQ