The Gap Between Capacity and Usefulness
Frontier models offer 200k token context windows. Some offer a million. The marketing implies you can dump your repository in and ask questions. The engineering reality is more complicated. Effective context is the subset of the window that the model actually uses for the task. Everything else is overhead, token cost, latency cost, and noise that degrades reasoning.
Our internal benchmark across 12 codebases showed accuracy plateauing around 40k input tokens for code generation, with a sharp regression past 90k. Adding more context past the plateau hurt more than it helped. This post is the set of patterns we use to keep workflows in the productive zone of the window.
What Belongs in Context, What Does Not
Three buckets:
Belongs: The ticket, the files the AI will edit, the immediate callers of any function being changed, the type definitions and interfaces touched by the change, the relevant tests, and the recent commit messages on the involved files.
Sometimes belongs: Sibling files in the same module, configuration files, dependency type definitions, related ADRs, the most recent issue conversation.
Does not belong: The full repository, unrelated test suites, lockfiles, generated code, documentation that is not specifically about the changed area, license files, large binary blobs decoded to base64 by an over-eager indexer.
The third bucket is where most context window waste happens. Teams index the entire repo, retrieve too broadly, and assume the model will sort it out. The model will not. It will use the irrelevant context as evidence and hallucinate confidently.
The Retrieval Strategy
We use four retrieval stages, each progressively more expensive:
- Path-based. Files referenced by the ticket, files in the same directory, files matching the keyword. Fast, deterministic, cheap.
- Symbol graph. Files that define or use symbols mentioned in the ticket or in the path-based set. This is where AST-aware tooling pays off.
- Recent history. Files modified in the last 30 days that touch the same modules. Catches the "we just refactored this" context that pure retrieval misses.
- Vector search. Last resort. Cheap to query but produces the most noise. We use it only when the first three stages return less than 8 files.
Most tickets stop at stage 2. Stage 4 fires on 12% of tickets, mostly cross-cutting refactors.
The Compression Patterns That Work
Even with disciplined retrieval, you will sometimes have 40 files and 60k tokens of relevant content. Three compression patterns that preserve accuracy:
Skeletonization. Replace function bodies with signatures and docstrings for files where the AI does not need to read the implementation. A 4000-line file collapses to 400 lines that still convey the API surface. Keep the file the AI is editing in full.
Diff windowing. For files where the change is local, send a 50-line window around the relevant region instead of the full file. Combined with skeletons of the rest of the file, this is the largest token reduction in our pipeline.
Conversation summarization. For tickets with long comment threads, summarize older comments and keep the last three verbatim. Older context is usually historical reasoning, not active constraints.
The Patterns That Failed
Three patterns that sounded good and did not work:
LLM-as-compressor. Running a cheap model to summarize files before the expensive model reads them. The compressor confidently dropped the wrong details. We tried this for six weeks and rolled it back.
Pure vector retrieval. Embedding the entire codebase and retrieving top-K for every ticket. Top-K embeddings are biased toward keyword overlap and miss the type-level context that is critical for correctness. Symbol-graph retrieval beat vector retrieval on every benchmark we ran.
Repo-wide context dumping with long-context models. Even with a million-token window, dumping the repo regressed accuracy. The "lost in the middle" effect is real, and the cost was prohibitive.
The Window Budget
A budget we hold ourselves to, per stage:
- Ticket parsing: under 4k tokens.
- Plan generation: under 12k tokens.
- File selection: under 20k tokens (mostly file headers and signatures).
- Code generation: under 40k tokens, ideally under 25k.
- Reviewer summary: under 30k tokens.
Anything above the budget triggers a routing decision: split the ticket, escalate to a human, or split the change across multiple PRs. The budget enforces the discipline of decomposing big changes instead of throwing them at a model.
How To Diagnose A Context Problem
Three diagnostics for context-related failures:
The hallucinated reference. If the AI references a function or file that does not exist, it is filling in missing context with invention. Retrieval is missing something it needs.
The over-broad change. If the AI changes 14 files for a 1-file ticket, retrieval gave it too much. The model treated unrelated files as evidence that they needed updating.
The under-changed fix. If the AI fixes the symptom but misses the caller that also has the bug, retrieval missed the call graph. Symbol-graph retrieval would have caught this.
What This Adds Up To
Context window engineering is the unglamorous half of working with LLMs. The flashy part is the prompt; the durable part is the context. The teams that get reliable AI coding outcomes are not the ones with the cleverest prompts. They are the ones with the most disciplined retrieval and the tightest compression.
For more on how this fits into a full pipeline, see the multi-agent architecture. For how this interacts with cost, see the token economy.
Frequently asked questions
Does a bigger context window make AI code generation better?
Not on its own. Effective context is only the subset of the window the model actually uses; in one benchmark across 12 codebases, code-generation accuracy plateaued around 40k input tokens and regressed sharply past 90k. Beyond that plateau, extra context becomes token cost, latency, and noise that degrades reasoning rather than help.
How much context should you give an AI coding agent?
Only what the task needs: the ticket, the files being edited, immediate callers, the touched type definitions and interfaces, relevant tests, and recent commit messages. Keep repo-wide dumps, unrelated suites, lockfiles, and generated code out entirely. A practical per-stage budget keeps code generation under 40k tokens and ideally under 25k.
Is vector search the best retrieval method for AI coding?
No, treat it as a last resort. Pure vector retrieval biases toward keyword overlap and misses the type-level context that correctness depends on, and symbol-graph retrieval beat it on every benchmark tested. Use path-based and symbol-graph retrieval first, and reach for vector search only when the earlier stages return too few files.
How do you fit a large codebase into an LLM context window?
Retrieve narrowly, then compress. Skeletonization replaces function bodies with signatures and docstrings, diff windowing sends a small window around the changed region instead of the whole file, and conversation summarization condenses old comment threads while keeping the last few verbatim. This interacts directly with cost, see the token economy of per-stage model routing.
Why does an AI agent hallucinate code or change too many files?
Both are retrieval symptoms. A hallucinated reference to a nonexistent function means retrieval is missing something the model then invents; an over-broad change across many files for a one-file ticket means retrieval gave it too much and it treated unrelated files as evidence they needed updating. The fix is disciplined, symbol-graph-aware retrieval, not a cleverer prompt. For how this sits in the wider system, see the multi-agent architecture for code generation.
EnsureFix Research Team
The EnsureFix research team studies model behavior, confidence calibration, and evaluation methodology for autonomous coding agents, translating findings into the production pipeline.