Local Chunking And OCR
Prepare messy documents for RAG with recursive semantic chunking and local OCR before embedding.
TL;DR
- Split text recursively with a set
chunk_sizeandchunk_overlap. - Keep whole sentences together with the
RecursiveCharacterTextSplitterclass. - Extract text from scanned PDFs with
pytesseractOCR first.
Chunking Strategy
chunk_sizeSet a target size for each text chunk.
chunk_size = 800 # characters per chunkchunk_overlapOverlap chunks so split sentences stay whole.
chunk_overlap = 100Recursive SplitSplit on large separators, then smaller ones.
RecursiveCharacterTextSplitter(
chunk_size=800, chunk_overlap=100)SeparatorsPrefer paragraph and sentence breaks first.
separators=["\n\n", "\n", ". ", " "]Semantic Boundaries
Never Mid-SentenceAvoid cutting a sentence across chunks.
# Split on sentence ends, not raw lengthKeep Code TogetherDo not break a code block across chunks.
# Treat a fenced code block as one unitHeadings As AnchorsStart new chunks at section headings.
# Split on markdown headings firstOCR For Scanned PDFs
pytesseractRead text from an image with Tesseract OCR.
import pytesseract
from PIL import ImageOCR A PageConvert a page image into plain text.
text = pytesseract.image_to_string(img)pdf2imageRender PDF pages to images for OCR.
from pdf2image import convert_from_pathPrepare For Embedding
Clean TextStrip headers, footers, and extra whitespace.
text = text.strip()Then ChunkSplit the cleaned text into chunks.
chunks = splitter.split_text(text)Then EmbedEmbed each clean chunk with Ollama.
ollama.embed(model="nomic-embed-text",
input=chunks)Tips
- Add a small
chunk_overlapso a sentence split across two chunks still appears whole in at least one of them. - Run OCR with
pytesseracton scanned or image-based PDFs before embedding, since those pages contain no selectable text.
Warnings
- Splitting by a fixed character count alone cuts sentences and
codeblocks apart, which produces vague, low-quality embeddings. - Chunks that are too large bury the relevant sentence among noise; keep each focused so retrieval returns a precise passage.
In Practice
OCR a scanned PDF page, recursively chunk the text with overlap, then embed the clean chunks locally.
pdf2imagerenders the page so OCR has an image to read.pytesseractextracts the page's text from that image.- Recursive splitting keeps sentences whole with a small overlap.
- Each clean chunk is embedded locally for the vector store.
from pdf2image import convert_from_path
import pytesseract, ollama
from langchain_text_splitters import \
RecursiveCharacterTextSplitter
# 1. OCR the first page to text
page = convert_from_path("doc.pdf")[0]
text = pytesseract.image_to_string(page)
# 2. Recursively chunk with overlap
splitter = RecursiveCharacterTextSplitter(
chunk_size=800, chunk_overlap=100)
chunks = splitter.split_text(text)
# 3. Embed the clean chunks locally
vecs = ollama.embed(
model="nomic-embed-text", input=chunks)FAQ
Split recursively on natural boundaries, paragraphs first, then sentences, with a target chunk_size and a small chunk_overlap. This keeps each chunk coherent so its embedding captures one clear idea.
Recursive chunking tries large separators first, like blank lines, then falls back to smaller ones, like sentence ends, until chunks fit the size limit. A RecursiveCharacterTextSplitter implements this.
Scanned PDFs are images, so render each page with pdf2image and run pytesseract OCR to get text. Only then should you clean, chunk, and embed it.
A few hundred words is a common starting point. Smaller chunks give precise retrieval but more of them; larger chunks hold context but can bury the answer. Tune to your documents.