Local Vector Embeddings
Generate vector embeddings locally with the /api/embed endpoint and models like nomic-embed-text.
TL;DR
- Turn text into vectors with the
/api/embedendpoint. - Use a small model like
nomic-embed-textfor embeddings. - Batch a string or array of texts to
/api/embed.
Embedding Basics
Pull A ModelDownload a dedicated embedding model.
ollama pull nomic-embed-text/api/embedSend text and get back a vector array.
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": "The sky is blue"
}'VectorsThe response holds a numeric array per input.
# { "embeddings": [[0.01, -0.2, ...]] }Generate Vectors
Single InputPass one string to embed it.
"input": "one sentence to embed"Batch InputsPass an array to embed many at once.
"input": ["first text", "second text"]Response ShapeEach input maps to one vector in order.
# embeddings[0] matches input[0]Python
ollama.embedCall the embed helper with text.
r = ollama.embed(
model="nomic-embed-text", input="hi")Read VectorsThe result exposes the embeddings list.
vec = r["embeddings"][0]BatchPass a list to embed several texts.
ollama.embed(model="nomic-embed-text",
input=["a", "b"])Why Local
PrivacyText never leaves your machine to be embedded.
# No cloud call, no data egressNo Per-Call CostLocal embedding has no usage billing.
# Unlimited embeds after downloadOfflineWorks with no internet once pulled.
# Index private data fully offlineTips
- Pull a dedicated embedding model like
nomic-embed-text; it is small, fast, and purpose-built for turning text into vectors. - Embed many texts at once by passing an array to
input, which is far faster than one request per chunk in a loop.
Warnings
- Use the same embedding model for documents and queries; mixing models produces vectors you cannot compare meaningfully.
- Legacy
/api/embeddingsuses apromptfield and returnsembedding; current/api/embedusesinputand returnsembeddings.
In Practice
Generate embeddings for two sentences in one batch call, then compare them with cosine similarity.
- One batch call embeds both sentences in a single request.
- Each vector comes back in the same order as the inputs.
- Cosine similarity measures how close the two vectors are.
- Similar sentences score near 1.0; unrelated ones near 0.
import ollama, numpy as np
r = ollama.embed(
model="nomic-embed-text",
input=["I love cats", "I adore cats"],
)
a, b = r["embeddings"]
a, b = np.array(a), np.array(b)
cos = a @ b / (np.linalg.norm(a) *
np.linalg.norm(b))
print(round(float(cos), 3))FAQ
An embedding is a list of numbers that captures the meaning of text. Similar texts get similar vectors, so you can measure closeness mathematically, which powers search, clustering, and retrieval.
Use a dedicated embedding model such as nomic-embed-text. It is small and fast, and it is built for embeddings rather than chat, so it produces better vectors than a general chat model.
/api/embed is current: it takes an input field (a string or array) and returns embeddings, supporting batches. /api/embeddings is legacy, takes a single prompt, and returns embedding.
Local embedding keeps text on your machine, so sensitive documents or proprietary code never reach a third-party server. It also has no per-call cost and works fully offline once the model is pulled.