A code-review bot can leave a beautifully written comment about a bug that does not exist. It can also find a real defect that your expected-answer list missed. Before deciding whether an AI reviewer is useful, keep a record of what each finding actually proves.
On October 5, 2026, GitHub introduced ReviewBench, a research-preview benchmark for AI code review. GitHub reports a corpus of 219 public pull requests across 19 languages. Its grounded metrics compare findings with an existing reference set; augmented metrics also judge previously unmatched findings. That distinction is a useful reminder: an answer key can be incomplete. The workflow below is our proposed small-team worksheet, not the official benchmark protocol or a reproduced benchmark result.

Decide what you need the reviewer to catch
Write one narrow trial question before opening a leaderboard: “Can this reviewer identify behavior-changing defects in our import flow without burying the maintainer in speculative comments?” A small application does not need the same review mix as a platform library.
Choose the kinds of finding that belong in the trial. For example, count an input that produces a wrong result, a reachable crash, or a failed operation presented as successful. Track style preferences separately. Do not let a formatting suggestion outweigh a missed defect simply because both occupy one comment.
GitHub's responsible-use documentation describes code review as a supplement to human review and warns that plausible feedback can be incorrect. Keep human review and existing tests in place during the trial. Do not grant automatic merge or apply every suggested patch to make evaluation easier.
Freeze a small, varied trial set
Use code you are permitted to share with the chosen service. Public examples or sanitized fixtures are safer starting points than copying confidential production material into an unfamiliar tool. Check the tool's data handling and access scope before connecting a private repository.
- Known-defect cases: retain the version before the fix and the evidence explaining the defect. Keep the answer and later fix out of the candidate's supplied context.
- Routine changes: include changes that a human has reviewed without finding a material defect. Label them “no known defect,” not “proven bug-free.”
- Uncertain cases: isolate work where the expected behavior is still disputed. Do not use unsettled product policy as a scoring answer.
Record each base and head commit, relevant instructions, available context, tool configuration and run date. Give competing reviewers equivalent inputs. If a run fails or cannot see a file, record that limitation rather than treating silence as a clean review. Start small enough that a maintainer can adjudicate every finding; the first set is a feasibility check, not a population-level ranking.

Turn each comment into a checkable claim
Copy the finding into this three-part statement: When [trigger], [code path] causes [observable effect]. Then try to disprove it. Look for a caller-side guard, an invariant, a type constraint or an earlier validation step that makes the claimed condition unreachable.
Example, invented for this guide: a reviewer says an empty import will crash because the code reads the first item. A useful investigation checks whether empty input can reach that function. If validation rejects it earlier, the comment may be wrong for this application. If an alternate caller skips validation, the same comment may identify a real defect. The relevant result is the reachable path, not whether the warning sounds sensible.
| Evidence field | What to record |
|---|---|
| Trigger | The specific input, state or sequence needed |
| Path | File, revision, caller and relevant guard |
| Consequence | Expected behavior versus the claimed result |
| Check | A minimal test, trace or concrete code argument |
| Decision | Confirmed, rejected or unresolved, with reason |
A generated test is another proposal to review. Check that it exercises the claimed path and asserts the intended behavior. Run it only in an appropriate isolated development environment. A passing unrelated test neither confirms nor refutes the finding.

Keep duplicates and unknowns out of the victory lap
Merge comments describing the same underlying defect into one finding before counting. Preserve the original comments as evidence. Two line-level warnings caused by one missing boundary check should not become two independent catches.
Use three explicit outcomes. Confirmed means the reviewer has supported a real, relevant issue. Rejected means you have a concrete reason the claim fails. Unresolved means evidence or a product decision is missing. Assign an owner and next check to unresolved items; do not quietly move them into the favorable column.
For your local worksheet, report confirmed, rejected and unresolved counts separately. If you calculate confirmed divided by confirmed plus rejected, label it a confirmed-finding share among resolved findings and show the unresolved count beside it. This is not ReviewBench's official precision metric. For known defects, report how many distinct reference issues were found out of the known total; that does not measure all defects in the software.

Decide whether the extra review earns its place
Record the human minutes spent checking comments, not just the bot's response time. A reviewer that finds a useful defect but requires a long investigation of unrelated claims may fit only a narrow use case. Conversely, a quiet review is not automatically a good one: check the known defects it missed.
Before the trial, decide which misses are unacceptable and how much review effort you can afford. Afterward choose a bounded next step: keep it for a specified change type, adjust instructions and retest, or stop using it. Preserve failures and limitations. Do not describe one successful run as a reliable success rate.
Copy this review receipt
- Trial question and included finding types:
- Authorized repository or fixture set:
- Base/head revisions and supplied context:
- Reviewer, configuration, run date and incomplete runs:
- Distinct confirmed / rejected / unresolved findings:
- Known defects found / known defects in this set:
- Important misses and disputed expectations:
- Human investigation minutes and observed usage cost:
- Decision, permitted use case, owner and next review date:

If expected behavior is itself unclear, start with the one-rule brief and boundary examples. If another person will finish the investigation, leave a restart record that names the exact revision and remaining check. The aim is a review you can explain, not a longer pile of comments.