The Refusal Premise
There is a deep cultural pull toward "the AI should try." Teams want their agent to take every ticket because trying feels like progress. In practice, the agents with the best outcomes refuse roughly 30-45% of tickets at intake. The agents with the worst outcomes accept everything.
The reasoning is symmetric to a senior engineer's instinct. A senior engineer who never says "this ticket is not actionable" is not being helpful. They are creating bad work. The AI is no different.
This post is the taxonomy of refusals that we have found valuable.
The Six Refusal Categories
1. Underspecified intent. The ticket says "fix the dashboard" with no description of what is broken, no error message, no steps to reproduce. The AI should not guess. The refusal is "please specify which dashboard, what the expected vs. actual behavior is, and how to reproduce."
2. Ambiguous scope. The ticket says "improve the API" without saying which endpoint, which dimension (performance, correctness, ergonomics), or what success looks like. Any improvement would be a guess about intent. Refuse and ask for the scope.
3. Requires product judgment. The ticket asks the AI to choose between two reasonable approaches that have business implications. Choosing one without product input is not the AI's call. Refuse and route to a human PM.
4. Cross-team coordination. The ticket touches code owned by two teams and requires negotiation about API contracts. The AI does not have the social context to broker. Refuse and route to a tech lead.
5. Migration or one-shot. A one-shot data migration, a one-time cleanup, or a manual operations task. These are not the AI's wheelhouse, they require exact human verification, and the cost of automation does not pay back on a single execution. Refuse and route to a human.
6. Out of confidence range. A ticket that touches a part of the codebase where the AI has high historical rejection rates or low confidence in its plan. Refuse with a confidence note. A human can take it.
Why Refusal Is Better Than Bad Work
A bad AI PR has three costs:
- Reviewer time to evaluate and reject.
- Reviewer trust erosion. Each bad PR makes the next AI PR more suspect.
- Opportunity cost. The reviewer could have been doing real work.
A refusal has one cost:
- A few seconds of intake compute, and a Jira comment.
The math is decisive. A refusal that prevents one bad PR pays back instantly.
Building The Refusal Layer
Three components:
Intake classifier. A cheap model classifies the ticket against the six categories above plus an "accept" bucket. Runs in under 2 seconds per ticket. We use Haiku.
Confidence floor. If the planner's confidence on the ticket is below a threshold (we use 0.55 on a 0-1 scale), the ticket is refused with a "low confidence, needs human pickup" reason.
Refusal templates. Each refusal category has a templated comment that includes the reason and what the human needs to do to make the ticket actionable. The AI is not telling the human "no", it is telling the human "yes, after these changes."
The Refusal Tone
This matters more than it sounds. A refusal that reads as criticism of the ticket author is corrosive. A refusal that reads as helpful clarification is the opposite. We tune the templates with the team's PM and tech lead to land in the right register. The goal is for refusals to feel like a senior engineer's question, not a robot's complaint.
What To Measure
Three metrics specific to refusals:
Refusal rate. Should land between 20-50% in a healthy deployment. Below 20% suggests the agent is taking work it should not. Above 50% suggests the ticket intake quality is poor and the team should fix it upstream.
Refusal acceptance rate. When the AI refuses, does the human agree? We measure this by tagging refused tickets and seeing if the human edits the ticket and resubmits (the AI was right) or overrides the refusal (the AI was wrong). Healthy: 80%+ refusals stand.
Time-to-resolve on refused tickets. Refused tickets should not languish. If they do, the refusal templates are not clear enough about what the human needs to do.
A Specific Example
A ticket landed: "users complaining checkout is slow, fix it."
The AI's analysis: which checkout flow (mobile, web, embed)? Which step (cart, payment, confirmation)? What "slow" means (p50, p99, hangs)? Which environment (production, staging, specific region)?
The refusal: "I can take this once we know which flow and which step, and what p95 latency target counts as 'fixed.' If you can update the ticket with those details, I'll pick it back up."
This was the right outcome. The human PM came back with "web checkout, payment step, p95 above 4s, target 1.5s." Now the AI has a clean problem and produced a working fix in 12 minutes.
What This Adds Up To
Refusal is a feature, not a bug. The agents that refuse well are the ones that earn the team's trust. The agents that accept everything burn that trust quickly.
For the broader safety architecture, see enterprise safety layers. For how confidence scores feed into refusal decisions, see calibration of AI confidence scores.
Frequently asked questions
Should an AI coding agent accept every ticket?
No. The strong cultural pull toward 'the AI should try' produces the worst outcomes, because trying a poorly-specified ticket creates bad work that erodes reviewer trust. The best-performing agents refuse roughly 30 to 45% of tickets at intake, exactly as a senior engineer would push back on a ticket that isn't actionable yet.
What is a good ticket acceptance rate for an AI coding agent?
A healthy deployment refuses 20 to 50% of tickets. Below 20% suggests the agent is taking work it shouldn't; above 50% usually means intake quality is poor and should be fixed upstream. Also track refusal-acceptance rate, when the AI refuses, 80%+ of those refusals should stand after a human looks.
Why would an AI coding agent refuse a ticket?
It refuses when the work isn't safely actionable: intent is underspecified, scope is ambiguous, the choice requires product judgment, the change needs cross-team coordination, it's a one-shot migration that doesn't pay back automation, or it's outside the agent's confidence range. In each case the AI routes to the right human and says what would make the ticket actionable.
How do you build a ticket refusal layer for an AI agent?
Combine three components: a cheap intake classifier (a small model running in under two seconds) that sorts tickets into the six refusal buckets plus 'accept', a confidence floor that refuses anything the planner scores below about 0.55, and templated refusal comments that tell the human exactly what's needed to make the ticket actionable. The confidence side of this leans on calibrated scores, see a calibration method for AI confidence scores.
How do you write a good ticket for an AI coding agent?
Specify the intent and scope the agent would otherwise have to guess. For a bug, name the flow, the step, the expected vs. actual behavior, and how to reproduce; for a performance ticket, give the metric and target, 'web checkout, payment step, p95 above 4s, target 1.5s' turns a refusal into a clean problem the AI can solve. For the upstream discipline around this, see enterprise safety for AI-generated code.
EnsureFix Solutions Team
The EnsureFix solutions team helps engineering leaders evaluate, pilot, and roll out autonomous coding agents, drawing on real deployment data across customer teams.