g1t/docs/research/reviewbench.md

381 lines22,297 bytesCodeBlame

Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.

Research: shipping CSS, and a ReviewBench harness for g1t's reviewer1# Measuring g1t's reviewer with ReviewBench
2
3Internal research, 2026-10-06. Prompted by GitHub's "ReviewBench: an open
4benchmark for AI code review" (github.blog). It covers:
5
6- what the benchmark is;
7- how g1t's review agent maps onto it;
8- a harness in `bench/reviewbench/`, written but not run, because running
9 it costs model credit;
10- the cost of running it;
11- where our reviewer is likely weak, and what to try;
12- how to make review quality a tracked number.
13
14Internal only. User-facing copy never names other review products or
15leaderboard positions.
16
17## Verdict
18
19- **Worth doing.** ReviewBench is open (the code and the 219-PR corpus with
20 labeled findings are MIT, in one repository). It scores the output shape
21 our reviewer already produces: file, line, message. Its judge runs
22 locally with our own key. It gives us the first quality number for
g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent23 `@g1t review` that is not anecdote.
Research: shipping CSS, and a ReviewBench harness for g1t's reviewer24- **Harness status.** The code skeleton is complete in
25 `bench/reviewbench/`. The image builds from the production sandbox base
26 image. The prompt is generated from `crates/runner/src/review.rs`, and
27 `prompt.mjs --check` passes. Nothing has been run against a model.
28- **Cost, Sonnet 5.5 reviewer and Sonnet 5 judge.** These estimates come
29 from a token model and are plus or minus 50%. Calibrate them against the
30 billing ledger's real review runs.
31
32 | Run | PRs | Review | Judge | Total |
33 |---|---:|---:|---:|---:|
34 | 50-PR sample | 50 | about $62 | about $17 | **about $80** |
35 | Test set | 25 | | | about $38 |
36 | Full set, one round | 219 | about $274 | about $73 | **about $350** |
37 | Full set, three rounds (leaderboard protocol) | 657 | | | **about $1,040** |
38
39 The 50-PR sample on Opus 5.5 is about $103.
40- **Expected result as the prompt stands:** high precision and low recall.
41 The prompt asks for brevity and "only real problems". Production drops
42 comments outside changed files and anchors each comment to a single
43 line. The benchmark rewards several distinct, located findings per PR:
44 the golden set averages 12 true positives per PR, and 59% of them are
45 low severity. The first changes to try are in
46 [Improvements](#improvements-to-try-in-order).
47
48## The benchmark
49
50| | |
51|---|---|
52| Announcement | github.blog, "ReviewBench: an open benchmark for AI code review" |
53| Site, leaderboard | https://review-bench.ai (the site text is CC BY-NC 4.0) |
54| Repository | https://github.com/review-bench/ReviewBench, **MIT**; pinned at `ceb0794a` in the harness |
55| Data in the repository | `corpus/manifest.json` (219 PRs), `corpus/test/test.json` (the 25-PR test set), `golden/<pr_key>.json` (labeled findings, 7.8 MB). Each source repository is mirrored at `github.com/review-bench/<owner>_<repo>` with the base and head commits. |
56| Corpus | 219 PRs from 187 repositories, 19 languages: TypeScript 31%, Python 19%, C# 11%, Go 9%, JS 7%. Sizes: 36% over 1,000 changed lines; median 562 lines and 9 files. Types: 36% features, 27% bug fixes. |
57| Golden set | 4,632 findings, **2,623 TP** and 2,009 FP. The labeled FPs are kept so that a reviewer repeating a known false alarm is penalized. |
58| TP labels | Severity: high 181, medium 891, low **1,551**. Categories: correctness 1,068, maintainability 365, reliability 362, documentation 306, testing 223, security 98, api-architecture 81, performance 63, accessibility 57. Scope: introduced-by-pr 2,391, pre-existing 173, exacerbated 59. Context needed: diff-only 940, diff plus related files 1,529, broader project 154. 78% of TPs span several lines (median span 4 lines). |
59| Producers of the TPs | LLM reviewers 1,391; Copilot code review 1,185; human reviewers 39; deterministic tools 8. This is a source of bias, discussed below. |
60| Agent contract | One container per PR, `linux/amd64`, 15 minutes. Mounts: `/work/repo` (checkout at head, frozen, shallow, no network to GitHub), `/work/pr/diff.patch` (three-dot diff), `/work/pr/pr.json` (title and body). The agent writes `RB_OUT` as `{pr, agent, findings: [{file, start_line, end_line, message, producer}]}`. Egress is allowlisted. |
61| Judge | An LLM matcher compares candidate and golden findings per file, in chunks, semantically, many-to-many. Unmatched findings are then classified TP/FP with the same rubric used to build the golden set. Official judge: Claude Sonnet 5. Local judging uses your own provider key (`npm run judge -- --provider anthropic --model …`). |
62| Metrics | Grounded precision (matched TPs over matched) and grounded recall (golden TPs covered); the comparable pair. Augmented precision and recall, which also credit or penalize unmatched findings by the classifier. Novel TP count. F-beta with an adjustable β. Results can be split by severity and category, and reported macro and micro. Duration is reported but not scored. Leaderboard rows are the mean of 3 rounds on the full set. |
63| Leaderboard snapshot (2026-10-01) | Top rows: grounded precision 84 to 90%, **grounded recall 16 to 26%**, augmented F1 33 to 50, with 490 to 1,190 findings over 219 PRs. Precision is about the same for everyone; rank follows recall, and recall follows how many distinct, correct findings a reviewer emits. Every row is a vendor's product. |
64| Validity notes (from the docs) | The golden set is the union of what its producers found, so issues none of them found are invisible. The labels depend on the classifier (96.6% agreement with independent senior engineers). Augmented recall's denominator differs per agent; compare agents on grounded recall. |
65
66The bias matters for reading our score. Almost half the golden TPs were
67found by one commercial reviewer and most of the rest by LLM reviewers. A
68reviewer phrasing issues the way those producers do will match more
69easily. A finding nobody in the producer set made can still score through
70augmented metrics, but not grounded ones. Read our grounded recall as
71"agreement with the producer set", not as absolute coverage.
72
73## What g1t's reviewer does today
74
75Sources:
76
77- `services/runner/src/index.ts` (`review`, `startReviewRun`, `withMemory`,
78 `guidance`, `modelEnv`);
79- `crates/runner/src/review.rs` and `harness.rs`;
80- `services/work/src/reviews.rs` and `confidence.rs`;
81- `services/runner/wrangler.jsonc`.
82
g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent83**Trigger.** A person asks for a review (`@g1t review`, or the API).
Research: shipping CSS, and a ReviewBench harness for g1t's reviewer84g1t also starts one by itself, from lifecycle and wait queues (`index.ts`
85around line 1870). Admission and guardrails apply. The default time cap
86for a review is 30 minutes, the cost cap comes from the workspace plan, and
87Claude Code enforces `--max-budget-usd`.
88
89**Inputs (the prompt, in order):**
90
911. `Pull request #N: <title>`, then the description.
922. The linked issue's title and body, if there is one.
933. What people have said on the PR (`peopleSaid`: its comments).
944. Repository instructions (`instructionsFor`, task `review`): `AGENTS.md`,
95 `CLAUDE.md`, `.g1t/review.md`, and the same files in directories the
96 change touches. Up to 8,000 characters per file and 24,000 in total,
97 read from the head only for same-repository branches, never from forks.
985. Workspace and project memory (`memoryContext`), and the context hub
99 (`hubContext`: catalog, relevant memory, recent decisions).
1006. The fixed `INSTRUCTIONS` in `review.rs`: read `/work/change.diff`,
101 then surrounding code, run tests if you like, don't modify anything;
102 judge whether it does what it's for, whether it is correct, and whether
103 it would break anything; "Be specific and brief. Comment only on real
104 problems…"; write
105 `{verdict, body, comments: [{path, line, body}]}` to
106 `/work/review.json`.
107
108**The sandbox.**
109
110- Full clone at head, merge-base with the target branch, `git diff base
111 HEAD` (equivalent to ReviewBench's three-dot diff).
112- Claude Code headless (`--print`, `stream-json`, `--max-turns 80`,
113 `--dangerously-skip-permissions`) with the repository's toolchains, so it
114 can run tests.
115- Fork checkouts get `UNTRUSTED_FLAGS`, so the repository's own Claude
116 settings are not loaded.
117- A review gets no g1t MCP tools and no steer hooks, because no
118 `G1T_AGENT_TOKEN` is set.
119
120**What it does not see:**
121
122- Required checks and their results (the review env has no `CHECKS`).
123- The CI status of the head.
124- Other open PRs touching the same files, although the work service
125 computes them (`overlaps`).
126- Previous reviews of the same PR, beyond comments via `peopleSaid`.
127
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier128**Model.** The production route for `review` was **Claude Sonnet 5.5**
129(`claude-sonnet-5-5`) when this was measured. Reviews are now routed by
130tier (`AGENT_ROUTING` in `services/runner/wrangler.jsonc`): small changes
131that touch no sensitive path go to the small tier (Claude Haiku 4.5), the
132rest stay on Sonnet 5.5.
Research: shipping CSS, and a ReviewBench harness for g1t's reviewer133A workspace can route reviews to its own provider through the model proxy
134(`openModelSession`). No effort level is set, so it uses the Claude Code
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier135default.
Research: shipping CSS, and a ReviewBench harness for g1t's reviewer136
137**Outputs and post-processing** (`reviews.rs`, `report_review`):
138
139- Verdict `approve` / `request_changes`, or none if invalid.
140- A summary capped at 20,000 characters, signed with the model name.
141- Line comments: empty bodies dropped; **comments on files the PR does
142 not change are dropped**; at most **30** comments; one `line` (no
143 ranges).
g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent144- Posted as `g1t`, with a `review.completed` event.
Research: shipping CSS, and a ReviewBench harness for g1t's reviewer145- Confidence (`confidence.rs`) uses the result: `request_changes` sinks
146 the change's confidence; `approve` with 3 or more comments costs a
147 point; no review costs a point.
148- The reviewer reports no confidence or severity of its own. `confidence::ASK`
149 is added to implement and revise runs, not to reviews.
150
151### How it would be scored
152
153The adapter (`bench/reviewbench/agent/agent.mjs`) maps each surviving line
154comment to `{file: path, start_line: line, end_line: line, message: body}`.
155The verdict and summary are kept beside the findings for analysis
156(`findings.g1t.json`) but not scored. ReviewBench scores only located
157findings.
158
159So:
160
161- A problem described only in the summary scores nothing.
162- An empty comment list on a PR with golden TPs costs recall and nothing
163 else.
164- Every comment counts toward precision. A comment matching a golden *FP*
165 counts against grounded precision; an unmatched one is classified by the
166 judge.
167
168Benchmark repositories have no g1t memory, issue, comments or
169`.g1t/review.md`. The prompt is title, body and `INSTRUCTIONS`, plus
170whatever `CLAUDE.md` or `AGENTS.md` the repository carries (Claude Code
171reads `CLAUDE.md` itself). This measures the reviewer's core, which is
172what we want. Memory and instructions are product features to measure
173online.
174
175## The harness (`bench/reviewbench/`)
176
177| File | What it does |
178|---|---|
179| `run.mjs` | The driver. `setup` clones ReviewBench at the pinned commit and installs its judge. `build` builds the image. `estimate` prices a run without spending. `review` runs our reviewer through ReviewBench's own `scripts/try-agent.sh`, one fresh container per PR, with the benchmark's mounts and validation. `judge` runs ReviewBench's judge with our key. `report` prints a summary or one JSON line. Nothing that spends runs without `--yes`. There is a per-PR cap (`--budget-usd`, default $5, passed to `--max-budget-usd`) and a per-run cap (`--max-total-usd`, default $150). |
180| `Dockerfile` | `FROM` the production sandbox base image named in `services/runner/base.json` (pinned by digest, same Claude Code CLI version), plus the adapter. |
181| `agent/agent.mjs` | The contract adapter: what `review.rs` does after its clone (writes `/work/change.diff`, runs Claude Code with `harness.rs`'s flags, reads `/work/review.json`), then the filters `reviews.rs` applies (changed files only, at most 30, 20,000 characters), then findings. It also writes a `findings.g1t.json` sidecar: verdict, summary, cost, turns, and comments dropped. |
182| `agent/instructions.txt` | Generated from `review.rs` by `prompt.mjs`. `node bench/reviewbench/prompt.mjs --check` fails if production's prompt has moved, which can run in CI. |
183| `agent/variants/*.txt` | Alternative instructions to A/B against `production` (`--variant`). `findings-first.txt` is the first candidate (see below). |
184| `.cache/` | Gitignored: the ReviewBench clone, repository mirrors, runs, scores. |
185
186Model access goes into each container from the environment:
187
188- `ANTHROPIC_API_KEY` alone sends requests straight to the provider.
189- `ANTHROPIC_BASE_URL=<MODELS_URL>/anthropic` plus a model-proxy session
190 token sends them through g1t's model proxy. The spend then lands on a
191 workspace like any run: use an internal workspace such as flagon-io so
192 it is visible in billing.
193
194The judge needs `ANTHROPIC_API_KEY` (or another provider ReviewBench's
195`pi` registry supports).
196
197```sh
198node bench/reviewbench/run.mjs setup
199node bench/reviewbench/run.mjs build
200node bench/reviewbench/run.mjs estimate --set sample:50
201node bench/reviewbench/run.mjs review --set sample:50 --seed 1 --yes # production prompt
202node bench/reviewbench/run.mjs review --set sample:50 --seed 1 --variant findings-first --yes
203node bench/reviewbench/run.mjs judge --run <id> --yes
204node bench/reviewbench/run.mjs report --run <id>
205```
206
207`sample:N` is stratified by change size in the corpus's proportions
208(≤200 / 201 to 1,000 / >1,000 lines) and fixed by `--seed`, so week-over-week
209runs see the same PRs. The leaderboard's own protocol is `--set full`, three
210times.
211
212Before the first paid run:
213
2141. Run `build` and one PR (`try-agent.sh … --pr 0`) with a capped key, to
215 check that the adapter writes valid findings. This costs about $1.
2162. Check that the image runs as a user that can write `/work`. The base
217 image's default user is used.
2183. Calibrate `estimateOne` in `run.mjs` against the billing ledger's real
219 `review` runs. Every run reports `total_cost_usd`, and spend is kept per
220 pull request.
221
222The adapter re-implements about 40 lines of `review.rs` rather than
223calling it, because `review.rs` clones from g1t and posts back to the API.
224To remove that duplication, a small product change would add
225`MODE=review-local` to `g1t-runner`: read `/work/pr/*`, skip the clone and
226the report, and write `RB_OUT`. The adapter would then be `exec g1t-runner`.
227That is worth doing once the benchmark is in regular use.
228
229### Cost model
230
231Per PR, an agentic review re-reads a growing context from cache every turn.
232`run.mjs` models it as follows:
233
234- turns: 12 + 0.8·√lines + 0.4·files, capped at 80;
235- context: from about 18K tokens plus the diff, growing about 2.5K tokens
236 per turn, capped at 180K;
237- 90%+ of input served as cache reads ($0.20/M on Sonnet 5.5);
238- new context written once ($2.50/M);
239- about 450 output tokens per turn ($10/M).
240
241This gives about $0.90 for a median PR and about $1.25 averaged over the
242corpus, which is skewed by the 36% of PRs over 1,000 lines. The judge is
243estimated at about $0.33 per PR: matching calls plus a tool-using
244classification of each unmatched finding.
245
246## Likely weaknesses against the benchmark
247
248These come from reading the prompt and code; none is measured yet.
249
2501. **Recall is capped by the instructions.** "Be specific and brief.
251 Comment only on real problems" and nothing asking for coverage. The
252 golden set has 12 TPs per PR (median 10) and 59% are low severity, which
253 the benchmark counts as worth fixing: small correctness slips, missing
254 tests, stale docs. The best reviewers emit about 5 findings per PR. We
255 probably emit 0 to 3. Expect grounded recall well under 15%.
2562. **Category blind spots.** The instructions frame the review as
257 "correct, does what it's for, doesn't break anything". That covers
258 correctness (41% of TPs) and reliability (14%). It gives no prompt for
259 maintainability (14%), documentation (12%), testing (9%), security (4%),
260 API design, performance or accessibility.
2613. **Problems that live only in the summary.** Nothing requires every
262 issue in `body` to also be a line comment. Those issues score zero, and
263 in the product they are not anchored where an author fixes them.
2644. **Comments outside the changed files are dropped.** `reviews.rs`
265 discards them. The benchmark anchors some TPs in untouched files, for
266 example a missing dependency in `pyproject.toml`, a caller the change
267 breaks, or 173 pre-existing-scope TPs. That is a product decision (the
268 UI has nowhere to put them), and it costs recall.
2695. **Single-line anchors.** 78% of golden TPs span several lines. Matching
270 is semantic within a file, so this matters less than the wrong file
271 would. A range still helps the matcher and helps a person.
2726. **Large changes.** 36% of the corpus is over 1,000 lines (median 9
273 files, mean 28). One agent with 80 turns, a 30-minute cap and a
274 30-comment cap will read the first files carefully and skim the rest.
275 No fan-out by file group.
2767. **No verification or calibration.** There is no second pass that checks
277 each candidate against the code, so precision rests on the "brief"
278 instruction, which also caps recall. There is no per-comment severity or
279 confidence, so we cannot pick an operating point (β), and
280 `confidence.rs` can only count comments.
2818. **Context the product has but the reviewer does not get.** Required
282 checks and CI results, overlapping PRs, and earlier reviews. These do
283 not matter on the benchmark but matter online.
284
285## Improvements to try, in order
286
287Each is one `--variant` (or a flag) on the same 50-PR sample with the same
288seed. The comparison is grounded precision and recall, by severity and
289category.
290
2911. **Prompt: findings first** (`agent/variants/findings-first.txt`,
292 written). It works in three passes:
293 - map the change and its callers;
294 - hunt by an explicit category checklist;
295 - verify each candidate and drop what isn't true at head.
296
297 It asks for every distinct problem rather than "brief", requires every
298 summary issue as a line comment, and adds `end_line` and `severity`.
299 The adapter already reads `end_line`. Expect the largest recall gain
300 for little cost.
3012. **Self-critique as a separate pass.** A second, cheaper call (Sonnet
302 5.5 at low effort, or Haiku 4.5) gets each candidate plus the exact
303 code lines and answers "true at head? worth the author's time?". Keep
304 the survivors. This lets the first pass be generous (recall) while
305 holding precision. It costs about 10 to 20% more.
3063. **Context retrieval.** Before the agent starts, compute the changed
307 symbols (from diff hunks) and their references (`git grep`, or the
308 context service where indexed), and put the list in the prompt. Point
309 it at the tests that cover the touched files. This targets the 1,529
310 TPs needing "diff plus related files".
3114. **Fan-out on large diffs.** Over about 800 lines or 15 files, split by
312 directory or file group. Run parallel reviewers (Claude Code subagents,
313 or separate runs) with the shared PR context, then merge and dedupe by
314 file, line and meaning. This targets item 6, the long tail where recall
315 collapses.
3165. **Confidence calibration.** Have each comment carry `severity` and
317 `confidence`. On the benchmark, sweep a threshold to draw the
318 precision/recall curve and pick the product's operating point, for
319 example post only medium and above inline and fold the low items into
320 the summary. Feed the same fields to `confidence.rs`, so that one
321 high-severity comment counts for more than three nits.
3226. **Model and effort.** Run the same sample on Sonnet 5.5 at
323 `medium`/`high` effort and on Opus 5.5. The review route can change in
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier324 `AGENT_ROUTING` without a code change. Opus 5.5 is about 1.4x the cost
Research: shipping CSS, and a ReviewBench harness for g1t's reviewer325 per review at our token profile.
3267. **Product-side follow-ups** (not benchmark-visible):
327 - allow file-level comments on unchanged files when the change breaks
328 them, instead of dropping them;
329 - add required-check results and overlapping PRs to the review prompt.
330
331## Making review quality a tracked metric
332
333**Offline (ReviewBench).**
334
335- A weekly g1t Actions workflow, `.g1t/workflows/reviewbench.yml`, runs
336 `review --set sample:50 --seed 1` with the production variant and the
337 production model, then `judge` and `report --json`.
338- The runner base image already has the toolchains. The job needs Docker,
339 or, more simply, it can run `agent.mjs` directly in a sandbox that *is*
340 the same image, with `/work` laid out by a small wrapper instead of
341 `try-agent.sh`.
342- Budget: about $80 a week (about $350 a month), billed to the flagon-io
343 workspace through the model proxy so it shows in spend.
344- Also run on demand for any change to `review.rs`'s instructions or to
345 the review route. `prompt.mjs --check` in CI flags such a change.
346- Do a full-set run (about $350) monthly, or before a model switch.
347
348**Where results go.**
349
350- Append each `report --json` line to an internal results store: a D1
351 table in the work service, or a JSON file committed to `docs/research/`.
352- Show it on a **sudo "Review quality"** page, with a trend of:
353 - grounded precision and recall;
354 - recall by severity (high and medium matter most) and by category;
355 - findings per PR;
356 - cost per PR.
357
358 Mark the g1t commit and model on each point.
359- Not on public docs: our own numbers are fine internally, but
360 user-facing copy compares to no one.
361
362**Online, the metric that matters.** Track the *addressed rate* of
g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent363`g1t` line comments: the share followed by a commit touching those
Research: shipping CSS, and a ReviewBench harness for g1t's reviewer364lines before merge, or resolved by a person. Also track the share of
365`request_changes` verdicts that led to a revision. g1t has the comments,
366the commits and the review runs, so this is a query, not a model call. Put
367it next to the offline number on the same sudo page. The article's lesson
368is that the offline number earns trust only while it moves the same way as
369the online one.
370
371## Open questions
372
373- Run once on the public leaderboard? Onboarding is self-service (a GHCR
374 image and our key). The final run is three rounds on 219 PRs at our
375 inference cost (about $820) with their judge free. A ranking is
376 marketing-adjacent, and the no-comparisons rule applies to anything we
377 say about it.
378- The judge is a Claude model and so is our reviewer. The ReviewBench docs
379 report human agreement for the labels but no per-family judge bias.
380 Treat small differences (under 2 points, inside the leaderboard's
381 round-to-round deviation of about 0.5 to 2) as noise.