AI & LLMs9 min read

The Token Economy: How Per-Stage Model Routing Cut Our Ticket Cost 73%

Sending every step of an AI coding pipeline to the most expensive model is a tax on engineering budgets. A breakdown of the routing rules that dropped average ticket cost from $6.10 to $1.64.

EnsureFix Platform Team · Platform Engineers, EnsureFix
The Token Economy: How Per-Stage Model Routing Cut Our Ticket Cost 73%, EnsureFix

The Lesson Hidden in the Bill

Every team that runs an AI coding pipeline at scale hits the same wall: the bill is real, and it grows faster than ticket throughput. Our internal data set covered 14,200 production tickets across 38 repositories over six months. The first quarter ran every stage on a frontier model. The second quarter ran a routed pipeline. Average ticket cost dropped from $6.10 to $1.64, a 73% reduction, with no measurable change in acceptance rate.

This post is the routing table.

The Pipeline Stages Worth Routing

Our pipeline has nine stages: ticket parsing, plan generation, file selection, code generation, test synthesis, security scan, root-cause classification, reviewer summary, and audit packaging. Each stage has a different reasoning profile. Most do not need frontier-tier reasoning, and the ones that do only need it for a fraction of the work.

The rule that emerged: route by entropy, not by importance. Stages that produce a small, structured output from a large input (file selection, classification, summary) are low-entropy and route to small models. Stages that produce novel multi-line output from sparse input (code generation, root-cause reasoning) are high-entropy and route to large models.

The Routing Table

StageModelWhy
Ticket parsingHaikuPure extraction. JSON in, JSON out.
Plan generationHaikuSelecting from a fixed set of templates and constraints.
File selectionHaikuRanking from an enumerated candidate list.
Code generationSonnetHigh entropy, multi-file output, novel reasoning.
Test synthesisSonnetNeeds to reason about edge cases the change introduced.
Security scanHaikuPattern matching against known classes. Escalates to Sonnet on hit.
Root-cause classificationSonnetMulti-hypothesis reasoning.
Reviewer summaryHaikuSummarizing a known diff into a known format.
Audit packagingHaikuSerialization. No reasoning.

Two stages (code generation and root-cause classification), account for 78% of remaining spend. We do not try to cheap-route those. The savings come from refusing to run the other seven on Sonnet.

The Escalation Rule

A fixed-routing table is naive. Some tickets that look simple need Sonnet reasoning at a stage that normally runs on Haiku. The pattern that works: start at the cheap model, escalate on signal.

The signals we escalate on:

  • The cheap model output fails its own schema validation twice.
  • The cheap model reports low confidence in its own structured output.
  • The diff size or file count crosses a per-stage threshold.
  • The security scan flags any class above informational severity.

Escalation runs on 9% of stage invocations. The cost overhead of running cheap-then-expensive on those 9% is dwarfed by the savings on the 91% that stay cheap.

Where We Tried To Save And Failed

Three places where we tried to push the cheap model and rolled back:

Code generation in unfamiliar repos. A Haiku-generated first draft followed by a Sonnet polish looked promising. In practice the second pass had to discard most of the first pass for repos the agent had not seen before. Net cost was worse than just running Sonnet end-to-end.

Reviewer summaries for high-risk changes. Haiku is fine on summaries of 50-line diffs but produced unhelpful summaries on 400-line diffs. The fix was not Sonnet, it was a diff-size limit on summary jobs. Anything larger gets split first.

Root-cause classification across unrelated stack traces. Haiku produced a confident classification, often wrong. Escalating only on low confidence missed the cases where the model was wrong but confident. We now run Sonnet for any classification that touches more than one service.

The Counterintuitive Win: Latency

Routing did not just lower cost. It cut median ticket latency from 4m12s to 2m48s. Haiku is materially faster than Sonnet, and seven of nine stages now run on Haiku. The pipeline parallelizes some stages, but the sequential stages (ticket parsing, plan generation, file selection), all became faster.

For interactive workflows where a human is waiting for the plan before approving, this matters. A plan in 12 seconds keeps the reviewer engaged. A plan in 90 seconds means they context-switch and come back later.

What This Means For Teams Building Their Own Pipelines

Three questions worth asking before you finalize your stack:

  • Which of your stages are low-entropy structured output? Those are the cost-saving opportunities. Audit your pipeline and find the stages where you are paying for reasoning you do not need.
  • What is your escalation signal? "Schema validation failed" and "model self-reported low confidence" are cheap, reliable, and easy to wire up. Use them before reaching for more exotic signals.
  • Are you measuring the right thing? Cost per ticket is the headline number, but cost per accepted PR is what actually matters. If your routing hurts acceptance, the cost win evaporates.

Per-stage routing is the highest-ROI optimization most teams have not yet done. It is invisible to users, takes a few days of plumbing, and pays back permanently. For the deeper architecture rationale, see why single-agent LLMs fail in enterprise code and build vs. buy on multi-agent pipelines.

Frequently asked questions

How do you reduce the cost of an AI coding pipeline?

Stop sending every pipeline stage to a frontier model. Audit which stages are low-entropy structured output and route those to a small, cheap model, reserving the large model for genuinely novel reasoning like code generation. In one dataset this dropped average ticket cost 73% with no change in acceptance rate. For the architecture rationale behind splitting the pipeline this way, see why single-agent LLMs fail at enterprise code.

What is per-stage model routing in an AI coding pipeline?

It's assigning each pipeline stage the cheapest model that can do its job. Ticket parsing, plan generation, file selection, security scanning, reviewer summaries, and audit packaging run on a small model; code generation, test synthesis, and root-cause classification run on a larger one. The rule of thumb is to route by entropy (the novelty of the output), rather than by how 'important' the stage feels.

Should you use the same LLM for every step of a coding pipeline?

No. Running every stage on a frontier model is a tax on your budget: most stages produce a small, structured output from a large input and need no frontier reasoning at all. Using a mixed stack is one of the highest-ROI optimizations available. If you're weighing that build effort, see building vs buying a multi-agent code pipeline.

When should a cheap model escalate to a more expensive one?

Use cheap, reliable signals: the small model's output fails its own schema validation twice, the model self-reports low confidence, the diff size or file count crosses a per-stage threshold, or the security scan flags anything above informational severity. In practice escalation fires on about 9% of stage invocations, so the cost of running cheap-then-expensive on those is negligible.

Does model routing hurt code quality or acceptance rate?

It doesn't have to, the 73% cost reduction came with no measurable change in acceptance rate. The key is to measure cost per accepted PR, not just cost per ticket; if routing starts hurting acceptance, the cost win evaporates. Some things should never be cheap-routed, like code generation in unfamiliar repos, where a cheap first draft ends up discarded.

EnsureFix Platform Team

Platform Engineers, EnsureFix

The EnsureFix platform team builds the integrations, token economy, and context-engineering layer that let the agents scale across large repositories.

model routingAI cost optimizationClaude HaikuClaude Sonnettoken economics

From reading to running

Ready to automate your tickets?

Watch EnsureFix take a real item from your backlog all the way to a pull request.