The Metric That Predicts the Other Metrics
Every team rolling out an AI coding agent watches the same headline numbers: acceptance rate, escape rate, cycle time, cost per ticket. These are lagging indicators. By the time they move, the rollout is already going well or already in trouble.
The cheapest leading indicator we have found, across hundreds of deployments, is PR description quality. When the AI writes descriptions that read like a senior engineer wrote them, every downstream metric trends positive. When the descriptions read like AI slop, every downstream metric trends negative. The correlation is strong enough that we treat description quality as a circuit-breaker on rollout expansion.
Why Descriptions Predict the Rest
A PR description is the artifact closest to the reasoning behind the change. If the AI cannot articulate why it made a change, it almost certainly did not have the context to make a good change.
Three failure modes show up in descriptions first:
Hallucinated context. A description that references a function that does not exist, a ticket number that was not assigned, or a previous PR that was never merged. These references are easy to write and easy to verify, and they are the canary for an agent that is confabulating context.
Missing intent. A description that says "fix the bug" without saying which bug. The bug class, the trigger condition, and the chosen approach are exactly what a reviewer needs. Their absence is not a description problem, it is a thinking problem that leaked into the description.
Mismatch with the diff. A description that promises a one-line fix when the diff changes 14 files. Reviewers stop reading on mismatch, so the AI has lost its strongest reviewer aid.
How to Score Descriptions
We score on five attributes, each 0 or 1:
- Does the description name the specific bug class or feature being added?
- Does it identify the root cause (for fixes) or the design choice (for features)?
- Does it reference the ticket and any blocking dependencies?
- Does it list the files changed and why each one was needed?
- Does it include a verification plan?
A description scoring 5/5 is "shipping ready." A description below 3/5 is a signal to investigate the underlying change. Below 3/5 changes have 2.4x the escape rate of 5/5 changes in our dataset.
What Breaks Description Quality
Three patterns that degrade description quality over time:
Stale prompts. The description prompt was tuned to the team's PR conventions in week 1. By week 12, the team's conventions have drifted, and the AI is producing descriptions for the old conventions. Refresh the prompt quarterly using actual recent merged PRs as examples.
Token starvation on summaries. Some pipelines truncate the diff before sending it to the summary stage. If the truncation cuts off the file most relevant to the change, the description is going to miss the point. The fix is diff-aware sampling, not blind truncation.
Reviewer-summary-from-plan. A description generated from the plan, not from the actual diff, drifts whenever the implementation diverged from the plan. Generate descriptions from the diff plus the ticket. Not from the plan.
The Reviewer-Side Signal
Watch for one specific reviewer behavior: are they reading the description, or skipping straight to the diff? When descriptions are good, reviewers read them and the review takes minutes. When descriptions are bad, reviewers skip them and the review takes longer because the reviewer rebuilds context from scratch.
You can measure this if you have telemetry on review duration: per-PR review time correlates with description score. Below 3/5 descriptions add an average 4-6 minutes to review time. Multiplied across hundreds of PRs per week, the cost of bad descriptions is on the order of an FTE.
What To Do When Descriptions Degrade
When the score trend turns negative, the response order matters:
- Look at the recent rejected PRs first. A degradation in description quality usually shows up in rejected PRs before merged ones, because reviewers reject PRs they cannot understand.
- Sample 20 recent descriptions and read them as a human. Automated scoring is a proxy. Read the actual text. The failure mode is usually obvious, hallucinated context, missing intent, or mismatch.
- Check whether the underlying agent changed. Did the model version update? Did a prompt change? Did a new file-selection rule ship? Tie the degradation to a specific change.
- Roll back or repair. If a change shipped, roll back. If the model updated, retune the description prompt.
The Broader Point
Lagging indicators tell you what happened. Leading indicators tell you what is about to happen. For AI coding agents, the PR description is a window into the agent's reasoning that no other artifact gives you for free. It is generated anyway, and reading it costs nothing.
For the full set of metrics worth tracking, see 12 metrics for AI coding agent success. For how the description feeds back into the learning loop, see self-improving AI from code reviews.
Frequently asked questions
What is a leading indicator for AI coding agent success?
PR description quality. It's generated for free on every PR and it's the artifact closest to the agent's reasoning, so it predicts downstream metrics before they move. When descriptions read like a senior engineer wrote them, acceptance and escape rate trend positive; when they read like AI slop, everything trends negative. For the full lagging-indicator set to pair it with, see 12 metrics for AI coding agent success.
How do you measure PR description quality?
Score each description on five yes/no attributes: does it name the specific bug class or feature, identify the root cause or design choice, reference the ticket and blocking dependencies, list the files changed and why each was needed, and include a verification plan. A 5/5 is shipping-ready; anything below 3/5 is a signal to investigate the underlying change, since those changes carry 2.4x the escape rate.
Why do AI-generated PR descriptions get worse over time?
Three causes: stale prompts tuned to conventions the team has since drifted away from, token starvation where the diff is blindly truncated before the summary stage, and descriptions generated from the plan instead of the actual diff. Fix them with quarterly prompt refreshes on recent merged PRs, diff-aware sampling instead of blind truncation, and always summarizing from the diff plus ticket.
Do PR descriptions affect code review time?
Yes, measurably. When descriptions are good, reviewers read them and the review takes minutes; when they're bad, reviewers skip straight to the diff and rebuild context from scratch. Below-3/5 descriptions add an average of 4 to 6 minutes per PR. For more ways to compress review time, see how to reduce pull request review time with AI.
What should an AI-generated PR description include?
It should name the specific bug class or feature, state the root cause for a fix or the design choice for a feature, reference the ticket and any blocking dependencies, list each changed file and why it was touched, and provide a verification plan. Watch for the three failure modes that surface first, hallucinated context, missing intent, and a description that mismatches the diff.
How does PR description quality feed back into agent improvement?
The description is a free window into the agent's reasoning, so tracking its score trend tells you when the underlying agent has degraded, often visible in rejected PRs before merged ones. Tie any degradation to a specific change (a model update, a prompt change, a new file-selection rule) and roll back or retune. This closes into a broader learning loop covered in how self-improving AI learns from code reviews.
EnsureFix Engineering Team
The EnsureFix engineering team designs and operates the multi-agent pipeline that turns tickets into production-ready pull requests. They write about architecture, model routing, safety validation, and what actually ships in enterprise codebases.