Local Chunking And OCR

Prepare messy documents for RAG with recursive semantic chunking and local OCR before embedding.

TL;DR

  1. Split text recursively with a set chunk_size and chunk_overlap.
  2. Keep whole sentences together with the RecursiveCharacterTextSplitter class.
  3. Extract text from scanned PDFs with pytesseract OCR first.

Chunking Strategy

    chunk_size

    Set a target size for each text chunk.

    chunk_size = 800   # characters per chunk
    chunk_overlap

    Overlap chunks so split sentences stay whole.

    chunk_overlap = 100
    Recursive Split

    Split on large separators, then smaller ones.

    RecursiveCharacterTextSplitter(
      chunk_size=800, chunk_overlap=100)
    Separators

    Prefer paragraph and sentence breaks first.

    separators=["\n\n", "\n", ". ", " "]

Semantic Boundaries

    Never Mid-Sentence

    Avoid cutting a sentence across chunks.

    # Split on sentence ends, not raw length
    Keep Code Together

    Do not break a code block across chunks.

    # Treat a fenced code block as one unit
    Headings As Anchors

    Start new chunks at section headings.

    # Split on markdown headings first

OCR For Scanned PDFs

    pytesseract

    Read text from an image with Tesseract OCR.

    import pytesseract
    from PIL import Image
    OCR A Page

    Convert a page image into plain text.

    text = pytesseract.image_to_string(img)
    pdf2image

    Render PDF pages to images for OCR.

    from pdf2image import convert_from_path

Prepare For Embedding

    Clean Text

    Strip headers, footers, and extra whitespace.

    text = text.strip()
    Then Chunk

    Split the cleaned text into chunks.

    chunks = splitter.split_text(text)
    Then Embed

    Embed each clean chunk with Ollama.

    ollama.embed(model="nomic-embed-text",
      input=chunks)

Tips

  1. Add a small chunk_overlap so a sentence split across two chunks still appears whole in at least one of them.
  2. Run OCR with pytesseract on scanned or image-based PDFs before embedding, since those pages contain no selectable text.

Warnings

  1. Splitting by a fixed character count alone cuts sentences and code blocks apart, which produces vague, low-quality embeddings.
  2. Chunks that are too large bury the relevant sentence among noise; keep each focused so retrieval returns a precise passage.

In Practice

FAQ