Guides10 min read

Onboarding AI to a 10-Year-Old Codebase: A Field Report

Greenfield benchmarks lie about how AI handles legacy code. A field report from deploying an AI coding agent into a 10-year-old monolith, with the surprises, the failures, and what worked.

EnsureFix Customer Success · Customer Success, EnsureFix
Onboarding AI to a 10-Year-Old Codebase: A Field Report, EnsureFix

The Honest Premise

AI coding benchmarks are mostly run on small, modern, well-documented repositories. Most production code is not that. It is decade-old monoliths with outdated conventions, undocumented assumptions, and 12-year-old comments that contradict the code beneath them.

This is a report from one of those environments. A 10-year-old monolith. Mixed Python 2 and Python 3 (Python 2 not yet removed). Three generations of ORM patterns coexisting. Tests of varying quality. Documentation last updated 2019.

The team rolled out an AI coding agent over six months. This is what we learned.

Month One: The AI Could Not Find Anything

The first month was about retrieval, not generation. The AI's default retrieval (symbol graph, recent history, path-based), assumed conventions the codebase did not follow. Symbols were duplicated across modules with subtly different semantics. The same function name meant different things in two namespaces. Recent history was noisy because of an ongoing refactoring effort.

Findings:

  • Acceptance rate: 14%.
  • Most common rejection reason: "the AI edited the wrong file."

The fix was not better prompting. It was retrieval tuning specific to this codebase. We added:

  • An exclusion list of legacy modules the AI should not touch without elevated review.
  • A symbol disambiguation layer that asked the AI to confirm which of multiple same-named functions it meant.
  • A demoted weight on files older than 5 years that had not been touched in 12 months.

After tuning, retrieval accuracy roughly tripled. Generation followed.

Month Two: The AI Did Not Speak The Local Dialect

By month two, retrieval was reliable. The AI was finding the right files. The PRs it produced were technically correct but did not look like the team's code.

  • It used for loops where the team used list comprehensions.
  • It imported with absolute paths where the team used relative.
  • It used type hints where the team had not adopted them.
  • It used dataclasses where the team used plain classes with __init__ methods.

Each of these is correct Python. None of them was the team's code. Reviewers rejected for "style," which sounded petty but was actually a real issue, when code does not match the surrounding code, future readers waste time on the mismatch.

The fix was per-repo style profile, generated from 200 recently-merged PRs. The AI learned the team's dialect after a week of operation, and reviewer style complaints dropped to near zero.

Month Three: The Untestable Half

By month three, generation was working for the well-tested modules. The other half of the codebase had no tests at all. The AI's behavior on those modules was an extra dimension of risk: a change that the AI was confident in had no test guard.

We decided to slow down. The policy:

  • High-risk modules (legacy core, payments, auth): AI in suggest-only mode, human implements.
  • Medium-risk modules (recent code with partial tests): AI assists, human reviews.
  • Low-risk modules (new code, well-tested utilities): AI auto-applies low-risk changes, human reviews higher-risk.

The classification took a week of work with the team's tech leads. It was the most useful week of the rollout.

Month Four: The Hidden Coupling

Month four surfaced the deepest problem with legacy codebases: hidden coupling. Function A's behavior depended on Function B in a way no comment, type, or test recorded. The AI changed A correctly per the ticket. B silently broke under the new contract. The breakage shipped to production and triggered an incident.

This is not solvable by better prompts. It is the structural property of the codebase. The mitigation:

  • Per-incident, capture the hidden coupling as a learned rule. The rule says "if you change A, also check B." Next time, the AI catches it.
  • Over time, replace hidden couplings with explicit contracts. This is regular tech debt work the AI assists with rather than the AI's special responsibility.
  • For the riskiest hidden couplings, the team did a structured refactoring sprint. The AI helped, but the work was human-led.

The escape that triggered this was instructive, not catastrophic. It reset the team's calibration on what the AI could and could not do safely in this codebase. See postmortem-to-prompt pipeline for the rule-capture pattern.

Month Five: The Throughput Inflection

By month five, with retrieval tuned, style profile learned, risk classification in place, and hidden-coupling rules accumulating, the AI's effective contribution surpassed an entry-level engineer's typical throughput on this codebase. Specifically:

  • Acceptance rate: 71%.
  • Tickets handled per week: 38 (across the team's backlog of bug fixes, refactors, and small features).
  • Escape rate: 0.8%.
  • Per-ticket cost: $1.90.

For comparison, in month one those numbers were 14%, 6 tickets, 4.2% escape, and (effectively) infinite cost per useful PR.

Month Six: The Team Asks For More

Month six was when the team started asking the AI to do things we had explicitly excluded. "Can it touch the auth module now?" "Can we have it do small feature work?" The data was good enough to support cautious expansion.

The pattern that worked: incremental scope expansion, with one new category per month, always starting in suggest-only mode. By month nine, the AI was handling small features in the auth module with a human reviewer always in the loop.

What Generalizes

Three lessons that we have seen repeat across other legacy deployments:

Retrieval tuning is the long pole. Most teams underestimate this. They expect to spend most of the rollout time on prompting. They actually spend it on retrieval. Plan accordingly.

Style profile matters more in legacy than in greenfield. A greenfield codebase has no dialect yet, so reviewer style friction is low. A legacy codebase has a strong dialect, and AI deviations sting.

Risk classification is a one-time investment that pays back forever. Two weeks of senior engineer time to classify modules buys six months of safer rollout.

What Does Not Generalize

Three things that were specific to this team:

  • The decision to keep Python 2 around. This was a constraint we worked with, not one we tried to fix in scope.
  • The size of the team (28 engineers). Smaller teams can be more aggressive; larger teams are more conservative.
  • The team's tolerance for incremental scope expansion. Some teams want all-or-nothing.

Your rollout will have its own equivalents.

The Takeaway

A 10-year-old monolith is not a hostile environment for AI coding. It is a slow environment. The first 90 days produce numbers that look worse than the benchmark headlines. By month six, the numbers are competitive with greenfield deployments, and the institutional value of having the AI inside the legacy code is larger because the team had no other affordable path to backlog throughput.

The teams that give up at month two never see this. The teams that invest in retrieval, style, and risk classification do.

For the manager-side of this rollout, see engineering manager playbook. For the modernization angle, see AI code generation for legacy modernization.

Frequently asked questions

Can AI coding agents work on legacy codebases?

Yes, a 10-year-old monolith is a slow environment, not a hostile one. In one field deployment across a mixed Python 2/3 monolith with three ORM generations, the agent went from 14% acceptance in month one to 71% by month six, handling 38 tickets a week at a 0.8% escape rate and $1.90 per ticket. The teams that succeed invest in retrieval, style, and risk classification instead of giving up at month two. For the modernization angle, see AI code generation for legacy modernization.

How long does it take to onboard an AI coding agent to an old codebase?

Plan for roughly 90 days before the numbers become competitive. In practice the first month goes to retrieval tuning, the second to learning the team's code style, the third to risk classification and slowing down on untested modules, and later months to handling hidden coupling and expanding scope. By month six a legacy deployment can match greenfield metrics; the early numbers deliberately look worse than benchmark headlines.

Why does AI-generated code get rejected on style even when it's correct?

Because correct code that doesn't match the surrounding code makes future readers waste time on the mismatch. An agent might use for-loops where the team uses comprehensions, absolute imports where the team uses relative, or dataclasses where the team uses plain classes, all valid Python, none of it the team's dialect. The fix is a per-repo style profile generated from recently merged PRs, which the AI can learn in about a week. See the reviewer's toolkit for how style shows up in review.

What is hidden coupling and why does it break AI changes?

Hidden coupling is when one function's behavior depends on another in a way no comment, type, or test records. The AI can change the first function correctly per the ticket while silently breaking the second under the new contract, a real cause of a shipped production incident. It isn't solvable by better prompts because it's a structural property of the codebase; the mitigation is to capture each instance as a learned rule and, over time, replace hidden couplings with explicit contracts. See the postmortem-to-prompt pipeline for the rule-capture pattern.

Should an AI coding agent touch authentication and payment code?

Not at first, classify modules by risk and gate accordingly. High-risk modules like legacy core, payments, and auth start in suggest-only mode with a human implementing; medium-risk code gets AI assistance with human review; low-risk, well-tested utilities can auto-apply low-risk changes. Expand incrementally, one new category per month, always starting in suggest-only mode, until the data supports letting the AI do small feature work in sensitive modules with a reviewer always in the loop.

EnsureFix Customer Success

Customer Success, EnsureFix

The EnsureFix customer success team works with engineering teams post-deployment, capturing the playbooks and metrics that separate successful rollouts from stalled pilots.

legacy codemonolithsonboardingAI deploymentcase study

From reading to running

Ready to automate your tickets?

Watch EnsureFix take a real item from your backlog all the way to a pull request.