Local Chunking And OCRCheatsheet
Prepare messy documents for RAG with recursive semantic chunking and local OCR before embedding.
Prepare messy documents for RAG with recursive semantic chunking and local OCR before embedding.
Run Ollama in Docker with a persistent volume for models and GPU passthrough for fast inference.
Serve many users by raising per-instance concurrency and load balancing Ollama replicas with Nginx.
Use the Model Context Protocol to give local model agents standardized access to files, databases, and tools.
Open Ollama to your network safely with host binding, CORS origins, and an authenticating proxy.
Give Ollama a ChatGPT-style browser interface with Open WebUI, running locally in Docker.
Use the official ollama Python library to chat, stream tokens, manage models, and call a remote host.
Send images to local vision models like llava for captioning, UI analysis, and reading text.
Give local models live information with a web search tool, then ground answers in the results.
Compare base and instruct models, match parameter counts to your RAM, and set the right context window.
Master the core Ollama command line for running models, managing local weights, and chatting interactively.
Understand how Ollama splits model layers between GPU and CPU, tune the split, and use Apple unified memory.
Generate vector embeddings locally with the /api/embed endpoint and models like nomic-embed-text.
Configure Ollama storage, memory, GPU use, and networking with the environment variables the server reads.
Use Modelfile MESSAGE directives to preload few-shot examples and lock in a strict output format.
Let a local model request function calls with a tools schema, then run them and return the results.
Understand how Ollama runs open language models locally on your own hardware with no network calls.
Read GGUF quantization labels, estimate the VRAM they save, and choose the right precision for your hardware.