Context Windows & Token Budget
Understand the context window as a finite budget, what counts against it, and how to spend it on the code and facts that matter.
TL;DR
- The context window is the total text, your prompt plus the reply, the model can consider at once.
- Everything counts against it: instructions, pasted code, files the agent read, and the conversation so far.
- Relevance beats volume; a focused prompt usually outperforms one stuffed with extra code.
What Counts
The Whole WindowInput and output share one budget; a long reply leaves less room for your input.
window = system + prompt
+ pasted code + history + replyTokens, RoughlyEstimate about four tokens per word so you can gauge a big paste before you send it.
~4 tokens / word
500 lines of code ~= a lot of budgetHistory Adds UpEvery earlier turn stays in context until you trim or restart the session.
Turn 20 still 'remembers' turn 1
unless you start fresh.Spend It Well
Paste The SliceInclude the exact function, type, or error, not the entire file around it.
// just the failing function + its types,
// not the whole 800-line moduleDescribe The RestSummarize surrounding architecture in words instead of pasting it all.
"This runs in an Express route handler;
req.user is already populated."Put Instructions LastAfter a long paste, restate the task so it sits closest to the reply.
<paste code>
Now: add input validation to the above.Avoid Context Rot
Fresh SessionBegin a new chat for an unrelated task so old context cannot interfere.
New task -> new session.Summarize ForwardWhen a thread gets long, ask for a summary and carry it into a clean start.
"Summarize the decisions and the
current code state in 10 lines."Prune Dead EndsDo not keep debugging on top of abandoned attempts; restate from the good version.
Re-paste the known-good file,
then ask for the next change.Signals Of Overload
Ignored InstructionsWhen the model drops a rule you set earlier, it has likely fallen out of focus.
Repeat the rule near your question.Wrong TargetIf it answers a side point, your real question is buried too deep.
Move the question to the end.Vague RepliesGeneric answers often mean the useful context is diluted by noise.
Cut the paste down to essentials.Tips
- Paste the specific function, type, or error rather than a whole file when the rest is irrelevant.
- Start a fresh session for a new task so stale context does not crowd out what matters now.
Warnings
- Models lose track of detail as the window fills; important instructions can get buried mid-context.
- A giant paste can push your actual question so far back that the model answers the wrong part.
In Practice
Instead of pasting an 800-line file and asking 'why is this broken?', give the model the failing function, its types, the error, and a one-line description of where it runs. Same budget, far more signal.
- The error message tells the model exactly what went wrong and where.
- The failing function and its types are the only code that matters here.
- A one-line note about the environment replaces pasting the whole module.
- The question sits last, right next to where the answer will go.
Error:
TypeError: Cannot read properties of undefined
(reading 'id') at getOwnerName (orders.ts:42)
Function:
function getOwnerName(order: Order): string {
return users.find(u => u.id === order.ownerId).name;
}
Types:
type Order = { id: string; ownerId?: string };
type User = { id: string; name: string };
Context: runs inside an Express route; `users`
is loaded at startup and may not contain every id.
Why does this throw, and how do I fix it safely?FAQ
A token is a chunk of text the model processes, roughly three to four characters of English or code, so about four tokens per word. Your prompt, any pasted code, and the model's reply all consume tokens from the same window.
You can, but you usually should not. Even with large windows, models attend less reliably to detail buried in a long context, and extra code dilutes the signal. Curated context beats a data dump.
The window fills with earlier turns, old code, and dead ends. That 'context rot' pushes current instructions toward the edges and crowds out room for a good reply. Summarize and restart when it happens.
Agents pull files into context as they work, which can fill the window fast. Good agent prompts point at the specific files or symbols that matter so the agent does not read half the repo first.