g1t/docs/research/reviewbench.md

378 lines22,211 bytesCodeBlame
1# Measuring g1t's reviewer with ReviewBench
2
3Internal research, 2026-10-06. Prompted by GitHub's "ReviewBench: an open
4benchmark for AI code review" (github.blog). It covers:
5
6- what the benchmark is;
7- how g1t's review agent maps onto it;
8- a harness in `bench/reviewbench/`, written but not run, because running
9 it costs model credit;
10- the cost of running it;
11- where our reviewer is likely weak, and what to try;
12- how to make review quality a tracked number.
13
14Internal only. User-facing copy never names other review products or
15leaderboard positions.
16
17## Verdict
18
19- **Worth doing.** ReviewBench is open (the code and the 219-PR corpus with
20 labeled findings are MIT, in one repository). It scores the output shape
21 our reviewer already produces: file, line, message. Its judge runs
22 locally with our own key. It gives us the first quality number for
23 `@g1t-agent review` that is not anecdote.
24- **Harness status.** The code skeleton is complete in
25 `bench/reviewbench/`. The image builds from the production sandbox base
26 image. The prompt is generated from `crates/runner/src/review.rs`, and
27 `prompt.mjs --check` passes. Nothing has been run against a model.
28- **Cost, Sonnet 5.5 reviewer and Sonnet 5 judge.** These estimates come
29 from a token model and are plus or minus 50%. Calibrate them against the
30 billing ledger's real review runs.
31
32 | Run | PRs | Review | Judge | Total |
33 |---|---:|---:|---:|---:|
34 | 50-PR sample | 50 | about $62 | about $17 | **about $80** |
35 | Test set | 25 | | | about $38 |
36 | Full set, one round | 219 | about $274 | about $73 | **about $350** |
37 | Full set, three rounds (leaderboard protocol) | 657 | | | **about $1,040** |
38
39 The 50-PR sample on Opus 5.5 is about $103.
40- **Expected result as the prompt stands:** high precision and low recall.
41 The prompt asks for brevity and "only real problems". Production drops
42 comments outside changed files and anchors each comment to a single
43 line. The benchmark rewards several distinct, located findings per PR:
44 the golden set averages 12 true positives per PR, and 59% of them are
45 low severity. The first changes to try are in
46 [Improvements](#improvements-to-try-in-order).
47
48## The benchmark
49
50| | |
51|---|---|
52| Announcement | github.blog, "ReviewBench: an open benchmark for AI code review" |
53| Site, leaderboard | https://review-bench.ai (the site text is CC BY-NC 4.0) |
54| Repository | https://github.com/review-bench/ReviewBench, **MIT**; pinned at `ceb0794a` in the harness |
55| Data in the repository | `corpus/manifest.json` (219 PRs), `corpus/test/test.json` (the 25-PR test set), `golden/<pr_key>.json` (labeled findings, 7.8 MB). Each source repository is mirrored at `github.com/review-bench/<owner>_<repo>` with the base and head commits. |
56| Corpus | 219 PRs from 187 repositories, 19 languages: TypeScript 31%, Python 19%, C# 11%, Go 9%, JS 7%. Sizes: 36% over 1,000 changed lines; median 562 lines and 9 files. Types: 36% features, 27% bug fixes. |
57| Golden set | 4,632 findings, **2,623 TP** and 2,009 FP. The labeled FPs are kept so that a reviewer repeating a known false alarm is penalized. |
58| TP labels | Severity: high 181, medium 891, low **1,551**. Categories: correctness 1,068, maintainability 365, reliability 362, documentation 306, testing 223, security 98, api-architecture 81, performance 63, accessibility 57. Scope: introduced-by-pr 2,391, pre-existing 173, exacerbated 59. Context needed: diff-only 940, diff plus related files 1,529, broader project 154. 78% of TPs span several lines (median span 4 lines). |
59| Producers of the TPs | LLM reviewers 1,391; Copilot code review 1,185; human reviewers 39; deterministic tools 8. This is a source of bias, discussed below. |
60| Agent contract | One container per PR, `linux/amd64`, 15 minutes. Mounts: `/work/repo` (checkout at head, frozen, shallow, no network to GitHub), `/work/pr/diff.patch` (three-dot diff), `/work/pr/pr.json` (title and body). The agent writes `RB_OUT` as `{pr, agent, findings: [{file, start_line, end_line, message, producer}]}`. Egress is allowlisted. |
61| Judge | An LLM matcher compares candidate and golden findings per file, in chunks, semantically, many-to-many. Unmatched findings are then classified TP/FP with the same rubric used to build the golden set. Official judge: Claude Sonnet 5. Local judging uses your own provider key (`npm run judge -- --provider anthropic --model …`). |
62| Metrics | Grounded precision (matched TPs over matched) and grounded recall (golden TPs covered); the comparable pair. Augmented precision and recall, which also credit or penalize unmatched findings by the classifier. Novel TP count. F-beta with an adjustable β. Results can be split by severity and category, and reported macro and micro. Duration is reported but not scored. Leaderboard rows are the mean of 3 rounds on the full set. |
63| Leaderboard snapshot (2026-10-01) | Top rows: grounded precision 84 to 90%, **grounded recall 16 to 26%**, augmented F1 33 to 50, with 490 to 1,190 findings over 219 PRs. Precision is about the same for everyone; rank follows recall, and recall follows how many distinct, correct findings a reviewer emits. Every row is a vendor's product. |
64| Validity notes (from the docs) | The golden set is the union of what its producers found, so issues none of them found are invisible. The labels depend on the classifier (96.6% agreement with independent senior engineers). Augmented recall's denominator differs per agent; compare agents on grounded recall. |
65
66The bias matters for reading our score. Almost half the golden TPs were
67found by one commercial reviewer and most of the rest by LLM reviewers. A
68reviewer phrasing issues the way those producers do will match more
69easily. A finding nobody in the producer set made can still score through
70augmented metrics, but not grounded ones. Read our grounded recall as
71"agreement with the producer set", not as absolute coverage.
72
73## What g1t's reviewer does today
74
75Sources:
76
77- `services/runner/src/index.ts` (`review`, `startReviewRun`, `withMemory`,
78 `guidance`, `modelEnv`);
79- `crates/runner/src/review.rs` and `harness.rs`;
80- `services/work/src/reviews.rs` and `confidence.rs`;
81- `services/runner/wrangler.jsonc`.
82
83**Trigger.** A person asks for a review (`@g1t-agent review`, or the API).
84g1t also starts one by itself, from lifecycle and wait queues (`index.ts`
85around line 1870). Admission and guardrails apply. The default time cap
86for a review is 30 minutes, the cost cap comes from the workspace plan, and
87Claude Code enforces `--max-budget-usd`.
88
89**Inputs (the prompt, in order):**
90
911. `Pull request #N: <title>`, then the description.
922. The linked issue's title and body, if there is one.
933. What people have said on the PR (`peopleSaid`: its comments).
944. Repository instructions (`instructionsFor`, task `review`): `AGENTS.md`,
95 `CLAUDE.md`, `.g1t/review.md`, and the same files in directories the
96 change touches. Up to 8,000 characters per file and 24,000 in total,
97 read from the head only for same-repository branches, never from forks.
985. Workspace and project memory (`memoryContext`), and the context hub
99 (`hubContext`: catalog, relevant memory, recent decisions).
1006. The fixed `INSTRUCTIONS` in `review.rs`: read `/work/change.diff`,
101 then surrounding code, run tests if you like, don't modify anything;
102 judge whether it does what it's for, whether it is correct, and whether
103 it would break anything; "Be specific and brief. Comment only on real
104 problems…"; write
105 `{verdict, body, comments: [{path, line, body}]}` to
106 `/work/review.json`.
107
108**The sandbox.**
109
110- Full clone at head, merge-base with the target branch, `git diff base
111 HEAD` (equivalent to ReviewBench's three-dot diff).
112- Claude Code headless (`--print`, `stream-json`, `--max-turns 80`,
113 `--dangerously-skip-permissions`) with the repository's toolchains, so it
114 can run tests.
115- Fork checkouts get `UNTRUSTED_FLAGS`, so the repository's own Claude
116 settings are not loaded.
117- A review gets no g1t MCP tools and no steer hooks, because no
118 `G1T_AGENT_TOKEN` is set.
119
120**What it does not see:**
121
122- Required checks and their results (the review env has no `CHECKS`).
123- The CI status of the head.
124- Other open PRs touching the same files, although the work service
125 computes them (`overlaps`).
126- Previous reviews of the same PR, beyond comments via `peopleSaid`.
127
128**Model.** The production route for `review` is **Claude Sonnet 5.5**
129(`claude-sonnet-5-5`, `AGENT_ROUTES` in `services/runner/wrangler.jsonc`).
130A workspace can route reviews to its own provider through the model proxy
131(`openModelSession`). No effort level is set, so it uses the Claude Code
132default. The model-env test uses Opus 5.5 for review as a fixture only.
133
134**Outputs and post-processing** (`reviews.rs`, `report_review`):
135
136- Verdict `approve` / `request_changes`, or none if invalid.
137- A summary capped at 20,000 characters, signed with the model name.
138- Line comments: empty bodies dropped; **comments on files the PR does
139 not change are dropped**; at most **30** comments; one `line` (no
140 ranges).
141- Posted as `g1t-agent`, with a `review.completed` event.
142- Confidence (`confidence.rs`) uses the result: `request_changes` sinks
143 the change's confidence; `approve` with 3 or more comments costs a
144 point; no review costs a point.
145- The reviewer reports no confidence or severity of its own. `confidence::ASK`
146 is added to implement and revise runs, not to reviews.
147
148### How it would be scored
149
150The adapter (`bench/reviewbench/agent/agent.mjs`) maps each surviving line
151comment to `{file: path, start_line: line, end_line: line, message: body}`.
152The verdict and summary are kept beside the findings for analysis
153(`findings.g1t.json`) but not scored. ReviewBench scores only located
154findings.
155
156So:
157
158- A problem described only in the summary scores nothing.
159- An empty comment list on a PR with golden TPs costs recall and nothing
160 else.
161- Every comment counts toward precision. A comment matching a golden *FP*
162 counts against grounded precision; an unmatched one is classified by the
163 judge.
164
165Benchmark repositories have no g1t memory, issue, comments or
166`.g1t/review.md`. The prompt is title, body and `INSTRUCTIONS`, plus
167whatever `CLAUDE.md` or `AGENTS.md` the repository carries (Claude Code
168reads `CLAUDE.md` itself). This measures the reviewer's core, which is
169what we want. Memory and instructions are product features to measure
170online.
171
172## The harness (`bench/reviewbench/`)
173
174| File | What it does |
175|---|---|
176| `run.mjs` | The driver. `setup` clones ReviewBench at the pinned commit and installs its judge. `build` builds the image. `estimate` prices a run without spending. `review` runs our reviewer through ReviewBench's own `scripts/try-agent.sh`, one fresh container per PR, with the benchmark's mounts and validation. `judge` runs ReviewBench's judge with our key. `report` prints a summary or one JSON line. Nothing that spends runs without `--yes`. There is a per-PR cap (`--budget-usd`, default $5, passed to `--max-budget-usd`) and a per-run cap (`--max-total-usd`, default $150). |
177| `Dockerfile` | `FROM` the production sandbox base image named in `services/runner/base.json` (pinned by digest, same Claude Code CLI version), plus the adapter. |
178| `agent/agent.mjs` | The contract adapter: what `review.rs` does after its clone (writes `/work/change.diff`, runs Claude Code with `harness.rs`'s flags, reads `/work/review.json`), then the filters `reviews.rs` applies (changed files only, at most 30, 20,000 characters), then findings. It also writes a `findings.g1t.json` sidecar: verdict, summary, cost, turns, and comments dropped. |
179| `agent/instructions.txt` | Generated from `review.rs` by `prompt.mjs`. `node bench/reviewbench/prompt.mjs --check` fails if production's prompt has moved, which can run in CI. |
180| `agent/variants/*.txt` | Alternative instructions to A/B against `production` (`--variant`). `findings-first.txt` is the first candidate (see below). |
181| `.cache/` | Gitignored: the ReviewBench clone, repository mirrors, runs, scores. |
182
183Model access goes into each container from the environment:
184
185- `ANTHROPIC_API_KEY` alone sends requests straight to the provider.
186- `ANTHROPIC_BASE_URL=<MODELS_URL>/anthropic` plus a model-proxy session
187 token sends them through g1t's model proxy. The spend then lands on a
188 workspace like any run: use an internal workspace such as flagon-io so
189 it is visible in billing.
190
191The judge needs `ANTHROPIC_API_KEY` (or another provider ReviewBench's
192`pi` registry supports).
193
194```sh
195node bench/reviewbench/run.mjs setup
196node bench/reviewbench/run.mjs build
197node bench/reviewbench/run.mjs estimate --set sample:50
198node bench/reviewbench/run.mjs review --set sample:50 --seed 1 --yes # production prompt
199node bench/reviewbench/run.mjs review --set sample:50 --seed 1 --variant findings-first --yes
200node bench/reviewbench/run.mjs judge --run <id> --yes
201node bench/reviewbench/run.mjs report --run <id>
202```
203
204`sample:N` is stratified by change size in the corpus's proportions
205(≤200 / 201 to 1,000 / >1,000 lines) and fixed by `--seed`, so week-over-week
206runs see the same PRs. The leaderboard's own protocol is `--set full`, three
207times.
208
209Before the first paid run:
210
2111. Run `build` and one PR (`try-agent.sh … --pr 0`) with a capped key, to
212 check that the adapter writes valid findings. This costs about $1.
2132. Check that the image runs as a user that can write `/work`. The base
214 image's default user is used.
2153. Calibrate `estimateOne` in `run.mjs` against the billing ledger's real
216 `review` runs. Every run reports `total_cost_usd`, and spend is kept per
217 pull request.
218
219The adapter re-implements about 40 lines of `review.rs` rather than
220calling it, because `review.rs` clones from g1t and posts back to the API.
221To remove that duplication, a small product change would add
222`MODE=review-local` to `g1t-runner`: read `/work/pr/*`, skip the clone and
223the report, and write `RB_OUT`. The adapter would then be `exec g1t-runner`.
224That is worth doing once the benchmark is in regular use.
225
226### Cost model
227
228Per PR, an agentic review re-reads a growing context from cache every turn.
229`run.mjs` models it as follows:
230
231- turns: 12 + 0.8·√lines + 0.4·files, capped at 80;
232- context: from about 18K tokens plus the diff, growing about 2.5K tokens
233 per turn, capped at 180K;
234- 90%+ of input served as cache reads ($0.20/M on Sonnet 5.5);
235- new context written once ($2.50/M);
236- about 450 output tokens per turn ($10/M).
237
238This gives about $0.90 for a median PR and about $1.25 averaged over the
239corpus, which is skewed by the 36% of PRs over 1,000 lines. The judge is
240estimated at about $0.33 per PR: matching calls plus a tool-using
241classification of each unmatched finding.
242
243## Likely weaknesses against the benchmark
244
245These come from reading the prompt and code; none is measured yet.
246
2471. **Recall is capped by the instructions.** "Be specific and brief.
248 Comment only on real problems" and nothing asking for coverage. The
249 golden set has 12 TPs per PR (median 10) and 59% are low severity, which
250 the benchmark counts as worth fixing: small correctness slips, missing
251 tests, stale docs. The best reviewers emit about 5 findings per PR. We
252 probably emit 0 to 3. Expect grounded recall well under 15%.
2532. **Category blind spots.** The instructions frame the review as
254 "correct, does what it's for, doesn't break anything". That covers
255 correctness (41% of TPs) and reliability (14%). It gives no prompt for
256 maintainability (14%), documentation (12%), testing (9%), security (4%),
257 API design, performance or accessibility.
2583. **Problems that live only in the summary.** Nothing requires every
259 issue in `body` to also be a line comment. Those issues score zero, and
260 in the product they are not anchored where an author fixes them.
2614. **Comments outside the changed files are dropped.** `reviews.rs`
262 discards them. The benchmark anchors some TPs in untouched files, for
263 example a missing dependency in `pyproject.toml`, a caller the change
264 breaks, or 173 pre-existing-scope TPs. That is a product decision (the
265 UI has nowhere to put them), and it costs recall.
2665. **Single-line anchors.** 78% of golden TPs span several lines. Matching
267 is semantic within a file, so this matters less than the wrong file
268 would. A range still helps the matcher and helps a person.
2696. **Large changes.** 36% of the corpus is over 1,000 lines (median 9
270 files, mean 28). One agent with 80 turns, a 30-minute cap and a
271 30-comment cap will read the first files carefully and skim the rest.
272 No fan-out by file group.
2737. **No verification or calibration.** There is no second pass that checks
274 each candidate against the code, so precision rests on the "brief"
275 instruction, which also caps recall. There is no per-comment severity or
276 confidence, so we cannot pick an operating point (β), and
277 `confidence.rs` can only count comments.
2788. **Context the product has but the reviewer does not get.** Required
279 checks and CI results, overlapping PRs, and earlier reviews. These do
280 not matter on the benchmark but matter online.
281
282## Improvements to try, in order
283
284Each is one `--variant` (or a flag) on the same 50-PR sample with the same
285seed. The comparison is grounded precision and recall, by severity and
286category.
287
2881. **Prompt: findings first** (`agent/variants/findings-first.txt`,
289 written). It works in three passes:
290 - map the change and its callers;
291 - hunt by an explicit category checklist;
292 - verify each candidate and drop what isn't true at head.
293
294 It asks for every distinct problem rather than "brief", requires every
295 summary issue as a line comment, and adds `end_line` and `severity`.
296 The adapter already reads `end_line`. Expect the largest recall gain
297 for little cost.
2982. **Self-critique as a separate pass.** A second, cheaper call (Sonnet
299 5.5 at low effort, or Haiku 4.5) gets each candidate plus the exact
300 code lines and answers "true at head? worth the author's time?". Keep
301 the survivors. This lets the first pass be generous (recall) while
302 holding precision. It costs about 10 to 20% more.
3033. **Context retrieval.** Before the agent starts, compute the changed
304 symbols (from diff hunks) and their references (`git grep`, or the
305 context service where indexed), and put the list in the prompt. Point
306 it at the tests that cover the touched files. This targets the 1,529
307 TPs needing "diff plus related files".
3084. **Fan-out on large diffs.** Over about 800 lines or 15 files, split by
309 directory or file group. Run parallel reviewers (Claude Code subagents,
310 or separate runs) with the shared PR context, then merge and dedupe by
311 file, line and meaning. This targets item 6, the long tail where recall
312 collapses.
3135. **Confidence calibration.** Have each comment carry `severity` and
314 `confidence`. On the benchmark, sweep a threshold to draw the
315 precision/recall curve and pick the product's operating point, for
316 example post only medium and above inline and fold the low items into
317 the summary. Feed the same fields to `confidence.rs`, so that one
318 high-severity comment counts for more than three nits.
3196. **Model and effort.** Run the same sample on Sonnet 5.5 at
320 `medium`/`high` effort and on Opus 5.5. The review route can change in
321 `AGENT_ROUTES` without a code change. Opus 5.5 is about 1.4x the cost
322 per review at our token profile.
3237. **Product-side follow-ups** (not benchmark-visible):
324 - allow file-level comments on unchanged files when the change breaks
325 them, instead of dropping them;
326 - add required-check results and overlapping PRs to the review prompt.
327
328## Making review quality a tracked metric
329
330**Offline (ReviewBench).**
331
332- A weekly g1t Actions workflow, `.g1t/workflows/reviewbench.yml`, runs
333 `review --set sample:50 --seed 1` with the production variant and the
334 production model, then `judge` and `report --json`.
335- The runner base image already has the toolchains. The job needs Docker,
336 or, more simply, it can run `agent.mjs` directly in a sandbox that *is*
337 the same image, with `/work` laid out by a small wrapper instead of
338 `try-agent.sh`.
339- Budget: about $80 a week (about $350 a month), billed to the flagon-io
340 workspace through the model proxy so it shows in spend.
341- Also run on demand for any change to `review.rs`'s instructions or to
342 the review route. `prompt.mjs --check` in CI flags such a change.
343- Do a full-set run (about $350) monthly, or before a model switch.
344
345**Where results go.**
346
347- Append each `report --json` line to an internal results store: a D1
348 table in the work service, or a JSON file committed to `docs/research/`.
349- Show it on a **sudo "Review quality"** page, with a trend of:
350 - grounded precision and recall;
351 - recall by severity (high and medium matter most) and by category;
352 - findings per PR;
353 - cost per PR.
354
355 Mark the g1t commit and model on each point.
356- Not on public docs: our own numbers are fine internally, but
357 user-facing copy compares to no one.
358
359**Online, the metric that matters.** Track the *addressed rate* of
360`g1t-agent` line comments: the share followed by a commit touching those
361lines before merge, or resolved by a person. Also track the share of
362`request_changes` verdicts that led to a revision. g1t has the comments,
363the commits and the review runs, so this is a query, not a model call. Put
364it next to the offline number on the same sudo page. The article's lesson
365is that the offline number earns trust only while it moves the same way as
366the online one.
367
368## Open questions
369
370- Run once on the public leaderboard? Onboarding is self-service (a GHCR
371 image and our key). The final run is three rounds on 219 PRs at our
372 inference cost (about $820) with their judge free. A ranking is
373 marketing-adjacent, and the no-comparisons rule applies to anything we
374 say about it.
375- The judge is a Claude model and so is our reviewer. The ReviewBench docs
376 report human agreement for the labels but no per-family judge bias.
377 Treat small differences (under 2 points, inside the leaderboard's
378 round-to-round deviation of about 0.5 to 2) as noise.