A code-review bot can leave a beautifully written comment about a bug that does not exist. It can also find a real defect that your expected-answer list missed. Before deciding whether an AI reviewer is useful, keep a record of what each finding actually proves.

On October 5, 2026, GitHub introduced ReviewBench, a research-preview benchmark for AI code review. GitHub reports a corpus of 219 public pull requests across 19 languages. Its grounded metrics compare findings with an existing reference set; augmented metrics also judge previously unmatched findings. That distinction is a useful reminder: an answer key can be incomplete. The workflow below is our proposed small-team worksheet, not the official benchmark protocol or a reproduced benchmark result.

A potato detective inspects a bug card while a robot carries a tall stack of comments.
A specific finding needs evidence; comment volume alone is not a verdict. AI-generated conceptual illustration.

Decide what you need the reviewer to catch

Write one narrow trial question before opening a leaderboard: “Can this reviewer identify behavior-changing defects in our import flow without burying the maintainer in speculative comments?” A small application does not need the same review mix as a platform library.

Choose the kinds of finding that belong in the trial. For example, count an input that produces a wrong result, a reachable crash, or a failed operation presented as successful. Track style preferences separately. Do not let a formatting suggestion outweigh a missed defect simply because both occupy one comment.

GitHub's responsible-use documentation describes code review as a supplement to human review and warns that plausible feedback can be incorrect. Keep human review and existing tests in place during the trial. Do not grant automatic merge or apply every suggested patch to make evaluation easier.

Freeze a small, varied trial set

Use code you are permitted to share with the chosen service. Public examples or sanitized fixtures are safer starting points than copying confidential production material into an unfamiliar tool. Check the tool's data handling and access scope before connecting a private repository.

Record each base and head commit, relevant instructions, available context, tool configuration and run date. Give competing reviewers equivalent inputs. If a run fails or cannot see a file, record that limitation rather than treating silence as a clean review. Start small enough that a maintainer can adjudicate every finding; the first set is a feasibility check, not a population-level ranking.

Three folders carry bug, gear and question-mark icons while a potato sorts code cards.
Use separate groups for known defects, routine changes and uncertain cases. The icons are conceptual grouping aids. AI-generated illustration.

Turn each comment into a checkable claim

Copy the finding into this three-part statement: When [trigger], [code path] causes [observable effect]. Then try to disprove it. Look for a caller-side guard, an invariant, a type constraint or an earlier validation step that makes the claimed condition unreachable.

Example, invented for this guide: a reviewer says an empty import will crash because the code reads the first item. A useful investigation checks whether empty input can reach that function. If validation rejects it earlier, the comment may be wrong for this application. If an alternate caller skips validation, the same comment may identify a real defect. The relevant result is the reachable path, not whether the warning sounds sensible.

Evidence fieldWhat to record
TriggerThe specific input, state or sequence needed
PathFile, revision, caller and relevant guard
ConsequenceExpected behavior versus the claimed result
CheckA minimal test, trace or concrete code argument
DecisionConfirmed, rejected or unresolved, with reason

A generated test is another proposal to review. Check that it exercises the claimed path and asserts the intended behavior. Run it only in an appropriate isolated development environment. A passing unrelated test neither confirms nor refutes the finding.

Trigger, Code and Effect boxes are connected left to right under Show the Path.
Trace a claim from its triggering condition through the relevant code to a consequence. AI-generated conceptual diagram.

Keep duplicates and unknowns out of the victory lap

Merge comments describing the same underlying defect into one finding before counting. Preserve the original comments as evidence. Two line-level warnings caused by one missing boundary check should not become two independent catches.

Use three explicit outcomes. Confirmed means the reviewer has supported a real, relevant issue. Rejected means you have a concrete reason the claim fails. Unresolved means evidence or a product decision is missing. Assign an owner and next check to unresolved items; do not quietly move them into the favorable column.

For your local worksheet, report confirmed, rejected and unresolved counts separately. If you calculate confirmed divided by confirmed plus rejected, label it a confirmed-finding share among resolved findings and show the unresolved count beside it. This is not ReviewBench's official precision metric. For known defects, report how many distinct reference issues were found out of the known total; that does not measure all defects in the software.

A potato examines three trays labelled Confirmed, Rejected and Unresolved.
Unresolved findings stay visible instead of being silently counted as successes or errors. AI-generated conceptual illustration.

Decide whether the extra review earns its place

Record the human minutes spent checking comments, not just the bot's response time. A reviewer that finds a useful defect but requires a long investigation of unrelated claims may fit only a narrow use case. Conversely, a quiet review is not automatically a good one: check the known defects it missed.

Before the trial, decide which misses are unacceptable and how much review effort you can afford. Afterward choose a bounded next step: keep it for a specified change type, adjust instructions and retest, or stop using it. Preserve failures and limitations. Do not describe one successful run as a reliable success rate.

Copy this review receipt

An empty Review Receipt checklist lists Snapshot, Findings, Review time and Decision.
Complete the receipt with observed evidence, not predicted performance. AI-generated conceptual illustration.

If expected behavior is itself unclear, start with the one-rule brief and boundary examples. If another person will finish the investigation, leave a restart record that names the exact revision and remaining check. The aim is a review you can explain, not a longer pile of comments.