Writing · AI Engineering

Backtesting AI Code Review Against Our Own Bugs

18 real bugs, two tools, three traps — turning tool adoption into a data problem


TL;DR

  • We decided whether to adopt a new AI code-review tool not with a demo, but with a backtest against 18 real bugs from our history — the verdict was run both, not "replace".
  • More important than the comparison itself: three traps. (1) Measure which model your bot is actually running. (2) Suspect selection bias in your ground-truth set. (3) Never trust a single run.
  • What we gained: tool adoption — usually a matter of taste and politics — became a reusable data problem (a backtest harness).
Detection · Python50 vs 56%tool A vs incumbent bot — a tie
Detection · Scala40 vs 67%the gap depends on language
False positives~3 vs ~0incumbent wins on precision
Higher-tier rerun56 → 61%the model changes the result

1. Why a backtest

Someone proposed adopting a new AI code-review tool. "It's better than what we have," they said. The demo was impressive. But what does "better" mean? On our codebase, against the bugs we actually hit, does it catch as much as the incumbent? Demos and benchmark scores can't answer that. So we replayed real bugs from our history through both tools and measured.

2. Designing the backtest

18 ground-truth bugs real bugs · 2 languages Replay the diffs identical inputs guaranteed Both tools review record actual model versions Pre-frozen grading criteria first, results later 4-axis scoring detection·FP·speed·severity
The backtest pipeline — the crucial step is freezing the grading criteria in writing before any review runs
  • Ground-truth set: 18 real bugs the incumbent bot had flagged as "critical" in the past, each confirmed valid by a human. Drawn from two repos with different primary languages (Python and Scala/JVM), so we could see language bias.
  • Identical conditions: we replayed the exact diff of the PR containing each bug, giving both tools the same input. Grading criteria (caught / partial / missed) were frozen in a doc before execution — so nobody could move the goalposts after seeing results.
  • Four axes: detection, false positives, speed, and severity signal. Detection alone is not enough.

The results are in the stat tiles above. The verdict was run both, not "replace" — each tool missed different classes of bugs, and tool A's "zero cost + one-click install" was real value too. But the traps we found along the way matter more than the verdict.

3. Three traps that are easy to fall into

Trap 1 — Measure which model your bot is actually running on

Our bot underperformed expectations, so we dug into the logs — the model wasn't pinned, and it had been running on the lower-tier runtime default the whole time. Rerunning on the higher tier lifted detection from 56% to 61%, with the biggest gains on cross-context bugs that only show up when you connect multiple files and rules. Don't record the tool's name — pin down the actual model and version at execution time from CI logs. Without that, your comparison is neither reproducible nor interpretable.

Trap 2 — The origin of your ground-truth set decides the outcome (selection bias)

Our ground-truth set was "bugs the incumbent bot had caught" — a set maximally biased in the incumbent's favor. So we re-measured against real hotfix bugs that had occurred independently of any bot, and both tools dropped to ~30% detection or below. That's the real skill level. This re-measurement didn't just teach humility — it exposed the classes of bugs AI review misses, which became our improvement backlog.

Trap 3 — AI review is nondeterministic. Never trust a single run

Feed the same code to the same tool multiple times and the findings differ slightly every run. A single comparison can be a snapshot of a coin flip. For borderline cases we ran multiple passes and classified them as "reliably caught / sometimes caught / reliably missed". Nondeterminism is itself a property of the tool.

4. A tool-adoption checklist

  1. Backtest against past bugs from your own codebase, not a demo
  2. Freeze grading criteria in writing before execution
  3. Verify the actual model and version the tool runs from logs — the tier changes the result
  4. Suspect bias in your ground-truth set and cross-validate with an independent set (hotfixes, etc.)
  5. Measure nondeterminism with multiple runs
  6. Score beyond detection: false positives, speed, cost, ease of setup — the decision is a vector

The biggest win wasn't which tool came out ahead. It's that we turned "tool adoption" — a question usually settled by taste and politics — into a data problem. The next time a new tool shows up (and it will), we just drop it into the same harness and hit run.