The Asymmetry
Human reviewers and AI reviewers are good at different things. Humans understand intent, architecture, and the social context of code changes. AI is patient and exhaustive, it can apply the same check to every line of every file in every PR without fatigue.
The best code review workflow uses both. The AI catches the patterns humans habitually miss; the human catches the things AI cannot evaluate. This post is the seven AI heuristics that produce the highest catch rate in our data.
Heuristic 1: Cross-Caller Inconsistency
The change updates one caller of a function but not the others. The function's contract is implicitly changing, and the unchanged callers will break or behave inconsistently.
Why humans miss it: the reviewer is looking at the diff, not the absent diff. Code that is not changing is invisible.
How the AI catches it: the symbol graph identifies all callers. If the diff changes the call shape at one site, the AI checks every other site against the new shape and flags mismatches.
In our data, this heuristic catches about 9% of real bugs across all categories. It is the single highest-yield AI-side check we run.
Heuristic 2: Untested Error Path
The new code adds an error case (a new branch, a new exception, a new return-with-error path), but the tests do not exercise it.
Why humans miss it: error paths read as defensive code. Reviewers approve them on intent, not coverage. They feel safe.
How the AI catches it: after the change, run a coverage delta. If new branches show up in the source and the test diff does not include a test that exercises them, flag.
This catches bugs that hide until production triggers the error path, which is exactly when you do not want the bug.
Heuristic 3: Boundary Condition Drift
The change adjusts a condition (x > 5 becomes x >= 5, or for i in range(n) becomes for i in range(n+1)). The intent might be correct, but the boundary case test is now wrong or absent.
Why humans miss it: off-by-one changes look minor. They are visually small. The reviewer's eye skips them.
How the AI catches it: it specifically flags any change that alters a boundary operator or loop bound. The reviewer is asked: is this intentional, and is the boundary case tested?
Heuristic 4: Layer Violation
Code that belongs in a service layer is added to a controller, or business logic is added to a data access layer. The change works correctly but pollutes the boundary.
Why humans miss it: the diff looks fine in isolation. The reviewer would need to mentally place the change in the architecture diagram, which they often skip.
How the AI catches it: it has a learned model of the team's architectural conventions (built from accepted PRs over time). When a change adds code that violates the learned pattern, the AI flags. See self-improving AI from code reviews.
Heuristic 5: Resource Lifetime Mismatch
A connection is opened but not closed under all paths. A lock is acquired but the release is conditional. A file is opened in one function and expected to be closed elsewhere.
Why humans miss it: resource management is verbose. Reviewers skim it.
How the AI catches it: it traces every resource acquisition through every return path and flags any path that exits without release.
Heuristic 6: Schema-Migration Compatibility
The code change references a new column or table that the corresponding migration has not added yet, or vice versa.
Why humans miss it: they look at the code in one repo and the migration in another. The mismatch is across files in a way the eye does not catch.
How the AI catches it: it has the schema (from the migration files) and the code (from the change). It cross-references column names and table references against the schema and flags mismatches.
Heuristic 7: Performance Regression Pattern
A new database query inside a loop. A new synchronous call in a hot path. A new full-table scan where an indexed query would work.
Why humans miss it: performance regressions are not usually obvious from code reading. They show up under load.
How the AI catches it: pattern-matches against a curated set of known anti-patterns (N+1 queries, sync-in-async, missing index). The catch rate on performance regressions in our data is around 40% of the issues that would otherwise have shown up in load tests.
What This Set Cannot Do
Three categories the AI does not catch reliably:
Architectural mistakes. A change that is locally correct but heads the system in a direction the team does not want to go. Only the senior reviewer sees this.
Intent mismatches. The code does X. The ticket asked for Y. They are similar but not the same. The AI usually does not catch this because the ticket itself was ambiguous.
Cultural fit. Whether the change reads like the team's code. The AI's tone-matching is decent but not perfect. Reviewers still calibrate.
How To Combine With Humans
Three patterns:
AI-first comments, human-final. The AI posts its findings as comments on the PR. The human reviewer reads them, agrees or dismisses each, and adds their own. The human is the final approver.
Human-first triage, AI-second sweep. The human does the intent review and the architecture review. After they approve on intent, the AI does the line-by-line sweep. The merge requires both.
AI-blocking, human-overriding. The AI's findings block the PR by default. The human can override individual findings with a justification that becomes part of the audit trail.
Most mature teams use a variant of the third. The override pattern matters because it forces a justification, which catches the "rubber stamp" failure mode where reviewers approve without reading.
What To Build First
If you are adding AI review to an existing workflow:
- Pick three of the seven heuristics. Cross-caller inconsistency, schema migration compatibility, and untested error path are the easiest to ship.
- Run them on the last quarter of merged PRs. Measure the catch rate and the false-positive rate.
- Promote to live review when false positives are under 5%.
- Add the next three.
The compounding effect is significant. Each heuristic shifts a class of bug from "found by users" to "found in review."
For how this fits the validation pipeline, see enterprise safety layers and AI SAST scanning inside pull requests.
Frequently asked questions
What bugs does AI code review catch that human reviewers miss?
AI reliably catches the patterns human attention skips: callers changed at one site but not others, error branches added without a test that exercises them, off-by-one boundary drift, business logic added to the wrong architectural layer, resources acquired but not released on every path, code referencing a column the migration has not added, and performance anti-patterns like queries inside loops. Humans miss these because the reviewer looks at the diff, not the absent diff, and code that is not changing is effectively invisible.
Can AI code review find off-by-one and boundary errors?
Yes. The boundary-condition-drift heuristic specifically flags any change that alters a comparison operator or loop bound, for example x > 5 becoming x >= 5, or range(n) becoming range(n+1). These changes look visually small so the reviewer's eye skips them, but the AI asks whether the change is intentional and whether the boundary case is tested. Catching it in review moves the bug out of production, where boundary errors otherwise surface.
How does AI detect N+1 queries and performance regressions in code review?
The performance heuristic pattern-matches against a curated set of known anti-patterns, a new database query inside a loop, a new synchronous call in a hot path, a new full-table scan where an indexed query would work. Because these regressions rarely show up from code reading and only appear under load, humans habitually miss them; the AI catches around 40% of the issues that would otherwise surface in load tests. It complements deeper security scanning, covered in AI SAST scanning inside pull requests.
Should AI code review block a PR or just leave comments?
Mature teams tend toward AI-blocking with human override: the AI's findings block by default, and a reviewer can override any individual finding with a justification that becomes part of the audit trail. The override pattern matters because forcing a justification catches the rubber-stamp failure mode where reviewers approve without reading. Alternatives are AI-first comments with a human final approver, or human-first intent review followed by an AI second-pass sweep.
What can't AI code review catch?
Three categories stay with the human reviewer. Architectural mistakes (a locally correct change that heads the system in a wrong direction), only a senior reviewer sees. Intent mismatches, where the code does X but the ticket asked for a similar-but-different Y, usually slip past because the ticket itself was ambiguous. And cultural fit, whether the change reads like the team's code, where the AI's tone-matching is decent but not perfect. See how the review model improves over time in self-improving AI that learns from code reviews.
How do you roll out AI code review heuristics without noise?
Start with three of the seven, cross-caller inconsistency, schema-migration compatibility, and untested error path are the easiest to ship. Run them against the last quarter of merged PRs to measure catch rate and false-positive rate, then promote to live review once false positives are under 5%, and add the next three. Each heuristic shifts a class of bug from found-by-users to found-in-review, and the compounding effect is significant. See enterprise safety for AI-generated code.
EnsureFix Engineering Team
The EnsureFix engineering team designs and operates the multi-agent pipeline that turns tickets into production-ready pull requests. They write about architecture, model routing, safety validation, and what actually ships in enterprise codebases.