Context Windows & Token Budget

Understand the context window as a finite budget, what counts against it, and how to spend it on the code and facts that matter.

TL;DR

  1. The context window is the total text, your prompt plus the reply, the model can consider at once.
  2. Everything counts against it: instructions, pasted code, files the agent read, and the conversation so far.
  3. Relevance beats volume; a focused prompt usually outperforms one stuffed with extra code.

What Counts

    The Whole Window

    Input and output share one budget; a long reply leaves less room for your input.

    window = system + prompt
      + pasted code + history + reply
    Tokens, Roughly

    Estimate about four tokens per word so you can gauge a big paste before you send it.

    ~4 tokens / word
    500 lines of code ~= a lot of budget
    History Adds Up

    Every earlier turn stays in context until you trim or restart the session.

    Turn 20 still 'remembers' turn 1
    unless you start fresh.

Spend It Well

    Paste The Slice

    Include the exact function, type, or error, not the entire file around it.

    // just the failing function + its types,
    // not the whole 800-line module
    Describe The Rest

    Summarize surrounding architecture in words instead of pasting it all.

    "This runs in an Express route handler;
    req.user is already populated."
    Put Instructions Last

    After a long paste, restate the task so it sits closest to the reply.

    <paste code>
    
    Now: add input validation to the above.

Avoid Context Rot

    Fresh Session

    Begin a new chat for an unrelated task so old context cannot interfere.

    New task -> new session.
    Summarize Forward

    When a thread gets long, ask for a summary and carry it into a clean start.

    "Summarize the decisions and the
    current code state in 10 lines."
    Prune Dead Ends

    Do not keep debugging on top of abandoned attempts; restate from the good version.

    Re-paste the known-good file,
    then ask for the next change.

Signals Of Overload

    Ignored Instructions

    When the model drops a rule you set earlier, it has likely fallen out of focus.

    Repeat the rule near your question.
    Wrong Target

    If it answers a side point, your real question is buried too deep.

    Move the question to the end.
    Vague Replies

    Generic answers often mean the useful context is diluted by noise.

    Cut the paste down to essentials.

Tips

  1. Paste the specific function, type, or error rather than a whole file when the rest is irrelevant.
  2. Start a fresh session for a new task so stale context does not crowd out what matters now.

Warnings

  1. Models lose track of detail as the window fills; important instructions can get buried mid-context.
  2. A giant paste can push your actual question so far back that the model answers the wrong part.

In Practice

FAQ