Local Vector Embeddings

Generate vector embeddings locally with the /api/embed endpoint and models like nomic-embed-text.

TL;DR

  1. Turn text into vectors with the /api/embed endpoint.
  2. Use a small model like nomic-embed-text for embeddings.
  3. Batch a string or array of texts to /api/embed.

Embedding Basics

    Pull A Model

    Download a dedicated embedding model.

    ollama pull nomic-embed-text
    /api/embed

    Send text and get back a vector array.

    curl http://localhost:11434/api/embed -d '{
      "model": "nomic-embed-text",
      "input": "The sky is blue"
    }'
    Vectors

    The response holds a numeric array per input.

    # { "embeddings": [[0.01, -0.2, ...]] }

Generate Vectors

    Single Input

    Pass one string to embed it.

    "input": "one sentence to embed"
    Batch Inputs

    Pass an array to embed many at once.

    "input": ["first text", "second text"]
    Response Shape

    Each input maps to one vector in order.

    # embeddings[0] matches input[0]

Python

    ollama.embed

    Call the embed helper with text.

    r = ollama.embed(
        model="nomic-embed-text", input="hi")
    Read Vectors

    The result exposes the embeddings list.

    vec = r["embeddings"][0]
    Batch

    Pass a list to embed several texts.

    ollama.embed(model="nomic-embed-text",
        input=["a", "b"])

Why Local

    Privacy

    Text never leaves your machine to be embedded.

    # No cloud call, no data egress
    No Per-Call Cost

    Local embedding has no usage billing.

    # Unlimited embeds after download
    Offline

    Works with no internet once pulled.

    # Index private data fully offline

Tips

  1. Pull a dedicated embedding model like nomic-embed-text; it is small, fast, and purpose-built for turning text into vectors.
  2. Embed many texts at once by passing an array to input, which is far faster than one request per chunk in a loop.

Warnings

  1. Use the same embedding model for documents and queries; mixing models produces vectors you cannot compare meaningfully.
  2. Legacy /api/embeddings uses a prompt field and returns embedding; current /api/embed uses input and returns embeddings.

In Practice

FAQ