Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 1 | # Measuring g1t's reviewer with ReviewBench |
| 2 | ||
| 3 | Internal research, 2026-10-06. Prompted by GitHub's "ReviewBench: an open | |
| 4 | benchmark for AI code review" (github.blog). It covers: | |
| 5 | ||
| 6 | - what the benchmark is; | |
| 7 | - how g1t's review agent maps onto it; | |
| 8 | - a harness in `bench/reviewbench/`, written but not run, because running | |
| 9 | it costs model credit; | |
| 10 | - the cost of running it; | |
| 11 | - where our reviewer is likely weak, and what to try; | |
| 12 | - how to make review quality a tracked number. | |
| 13 | ||
| 14 | Internal only. User-facing copy never names other review products or | |
| 15 | leaderboard positions. | |
| 16 | ||
| 17 | ## Verdict | |
| 18 | ||
| 19 | - **Worth doing.** ReviewBench is open (the code and the 219-PR corpus with | |
| 20 | labeled findings are MIT, in one repository). It scores the output shape | |
| 21 | our reviewer already produces: file, line, message. Its judge runs | |
| 22 | locally with our own key. It gives us the first quality number for | |
| g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent | 23 | `@g1t review` that is not anecdote. |
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 24 | - **Harness status.** The code skeleton is complete in |
| 25 | `bench/reviewbench/`. The image builds from the production sandbox base | |
| 26 | image. The prompt is generated from `crates/runner/src/review.rs`, and | |
| 27 | `prompt.mjs --check` passes. Nothing has been run against a model. | |
| 28 | - **Cost, Sonnet 5.5 reviewer and Sonnet 5 judge.** These estimates come | |
| 29 | from a token model and are plus or minus 50%. Calibrate them against the | |
| 30 | billing ledger's real review runs. | |
| 31 | ||
| 32 | | Run | PRs | Review | Judge | Total | | |
| 33 | |---|---:|---:|---:|---:| | |
| 34 | | 50-PR sample | 50 | about $62 | about $17 | **about $80** | | |
| 35 | | Test set | 25 | | | about $38 | | |
| 36 | | Full set, one round | 219 | about $274 | about $73 | **about $350** | | |
| 37 | | Full set, three rounds (leaderboard protocol) | 657 | | | **about $1,040** | | |
| 38 | ||
| 39 | The 50-PR sample on Opus 5.5 is about $103. | |
| 40 | - **Expected result as the prompt stands:** high precision and low recall. | |
| 41 | The prompt asks for brevity and "only real problems". Production drops | |
| 42 | comments outside changed files and anchors each comment to a single | |
| 43 | line. The benchmark rewards several distinct, located findings per PR: | |
| 44 | the golden set averages 12 true positives per PR, and 59% of them are | |
| 45 | low severity. The first changes to try are in | |
| 46 | [Improvements](#improvements-to-try-in-order). | |
| 47 | ||
| 48 | ## The benchmark | |
| 49 | ||
| 50 | | | | | |
| 51 | |---|---| | |
| 52 | | Announcement | github.blog, "ReviewBench: an open benchmark for AI code review" | | |
| 53 | | Site, leaderboard | https://review-bench.ai (the site text is CC BY-NC 4.0) | | |
| 54 | | Repository | https://github.com/review-bench/ReviewBench, **MIT**; pinned at `ceb0794a` in the harness | | |
| 55 | | Data in the repository | `corpus/manifest.json` (219 PRs), `corpus/test/test.json` (the 25-PR test set), `golden/<pr_key>.json` (labeled findings, 7.8 MB). Each source repository is mirrored at `github.com/review-bench/<owner>_<repo>` with the base and head commits. | | |
| 56 | | Corpus | 219 PRs from 187 repositories, 19 languages: TypeScript 31%, Python 19%, C# 11%, Go 9%, JS 7%. Sizes: 36% over 1,000 changed lines; median 562 lines and 9 files. Types: 36% features, 27% bug fixes. | | |
| 57 | | Golden set | 4,632 findings, **2,623 TP** and 2,009 FP. The labeled FPs are kept so that a reviewer repeating a known false alarm is penalized. | | |
| 58 | | TP labels | Severity: high 181, medium 891, low **1,551**. Categories: correctness 1,068, maintainability 365, reliability 362, documentation 306, testing 223, security 98, api-architecture 81, performance 63, accessibility 57. Scope: introduced-by-pr 2,391, pre-existing 173, exacerbated 59. Context needed: diff-only 940, diff plus related files 1,529, broader project 154. 78% of TPs span several lines (median span 4 lines). | | |
| 59 | | Producers of the TPs | LLM reviewers 1,391; Copilot code review 1,185; human reviewers 39; deterministic tools 8. This is a source of bias, discussed below. | | |
| 60 | | Agent contract | One container per PR, `linux/amd64`, 15 minutes. Mounts: `/work/repo` (checkout at head, frozen, shallow, no network to GitHub), `/work/pr/diff.patch` (three-dot diff), `/work/pr/pr.json` (title and body). The agent writes `RB_OUT` as `{pr, agent, findings: [{file, start_line, end_line, message, producer}]}`. Egress is allowlisted. | | |
| 61 | | Judge | An LLM matcher compares candidate and golden findings per file, in chunks, semantically, many-to-many. Unmatched findings are then classified TP/FP with the same rubric used to build the golden set. Official judge: Claude Sonnet 5. Local judging uses your own provider key (`npm run judge -- --provider anthropic --model …`). | | |
| 62 | | Metrics | Grounded precision (matched TPs over matched) and grounded recall (golden TPs covered); the comparable pair. Augmented precision and recall, which also credit or penalize unmatched findings by the classifier. Novel TP count. F-beta with an adjustable β. Results can be split by severity and category, and reported macro and micro. Duration is reported but not scored. Leaderboard rows are the mean of 3 rounds on the full set. | | |
| 63 | | Leaderboard snapshot (2026-10-01) | Top rows: grounded precision 84 to 90%, **grounded recall 16 to 26%**, augmented F1 33 to 50, with 490 to 1,190 findings over 219 PRs. Precision is about the same for everyone; rank follows recall, and recall follows how many distinct, correct findings a reviewer emits. Every row is a vendor's product. | | |
| 64 | | Validity notes (from the docs) | The golden set is the union of what its producers found, so issues none of them found are invisible. The labels depend on the classifier (96.6% agreement with independent senior engineers). Augmented recall's denominator differs per agent; compare agents on grounded recall. | | |
| 65 | ||
| 66 | The bias matters for reading our score. Almost half the golden TPs were | |
| 67 | found by one commercial reviewer and most of the rest by LLM reviewers. A | |
| 68 | reviewer phrasing issues the way those producers do will match more | |
| 69 | easily. A finding nobody in the producer set made can still score through | |
| 70 | augmented metrics, but not grounded ones. Read our grounded recall as | |
| 71 | "agreement with the producer set", not as absolute coverage. | |
| 72 | ||
| 73 | ## What g1t's reviewer does today | |
| 74 | ||
| 75 | Sources: | |
| 76 | ||
| 77 | - `services/runner/src/index.ts` (`review`, `startReviewRun`, `withMemory`, | |
| 78 | `guidance`, `modelEnv`); | |
| 79 | - `crates/runner/src/review.rs` and `harness.rs`; | |
| 80 | - `services/work/src/reviews.rs` and `confidence.rs`; | |
| 81 | - `services/runner/wrangler.jsonc`. | |
| 82 | ||
| g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent | 83 | **Trigger.** A person asks for a review (`@g1t review`, or the API). |
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 84 | g1t also starts one by itself, from lifecycle and wait queues (`index.ts` |
| 85 | around line 1870). Admission and guardrails apply. The default time cap | |
| 86 | for a review is 30 minutes, the cost cap comes from the workspace plan, and | |
| 87 | Claude Code enforces `--max-budget-usd`. | |
| 88 | ||
| 89 | **Inputs (the prompt, in order):** | |
| 90 | ||
| 91 | 1. `Pull request #N: <title>`, then the description. | |
| 92 | 2. The linked issue's title and body, if there is one. | |
| 93 | 3. What people have said on the PR (`peopleSaid`: its comments). | |
| 94 | 4. Repository instructions (`instructionsFor`, task `review`): `AGENTS.md`, | |
| 95 | `CLAUDE.md`, `.g1t/review.md`, and the same files in directories the | |
| 96 | change touches. Up to 8,000 characters per file and 24,000 in total, | |
| 97 | read from the head only for same-repository branches, never from forks. | |
| 98 | 5. Workspace and project memory (`memoryContext`), and the context hub | |
| 99 | (`hubContext`: catalog, relevant memory, recent decisions). | |
| 100 | 6. The fixed `INSTRUCTIONS` in `review.rs`: read `/work/change.diff`, | |
| 101 | then surrounding code, run tests if you like, don't modify anything; | |
| 102 | judge whether it does what it's for, whether it is correct, and whether | |
| 103 | it would break anything; "Be specific and brief. Comment only on real | |
| 104 | problems…"; write | |
| 105 | `{verdict, body, comments: [{path, line, body}]}` to | |
| 106 | `/work/review.json`. | |
| 107 | ||
| 108 | **The sandbox.** | |
| 109 | ||
| 110 | - Full clone at head, merge-base with the target branch, `git diff base | |
| 111 | HEAD` (equivalent to ReviewBench's three-dot diff). | |
| 112 | - Claude Code headless (`--print`, `stream-json`, `--max-turns 80`, | |
| 113 | `--dangerously-skip-permissions`) with the repository's toolchains, so it | |
| 114 | can run tests. | |
| 115 | - Fork checkouts get `UNTRUSTED_FLAGS`, so the repository's own Claude | |
| 116 | settings are not loaded. | |
| 117 | - A review gets no g1t MCP tools and no steer hooks, because no | |
| 118 | `G1T_AGENT_TOKEN` is set. | |
| 119 | ||
| 120 | **What it does not see:** | |
| 121 | ||
| 122 | - Required checks and their results (the review env has no `CHECKS`). | |
| 123 | - The CI status of the head. | |
| 124 | - Other open PRs touching the same files, although the work service | |
| 125 | computes them (`overlaps`). | |
| 126 | - Previous reviews of the same PR, beyond comments via `peopleSaid`. | |
| 127 | ||
| Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier | 128 | **Model.** The production route for `review` was **Claude Sonnet 5.5** |
| 129 | (`claude-sonnet-5-5`) when this was measured. Reviews are now routed by | |
| 130 | tier (`AGENT_ROUTING` in `services/runner/wrangler.jsonc`): small changes | |
| 131 | that touch no sensitive path go to the small tier (Claude Haiku 4.5), the | |
| 132 | rest stay on Sonnet 5.5. | |
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 133 | A workspace can route reviews to its own provider through the model proxy |
| 134 | (`openModelSession`). No effort level is set, so it uses the Claude Code | |
| Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier | 135 | default. |
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 136 | |
| 137 | **Outputs and post-processing** (`reviews.rs`, `report_review`): | |
| 138 | ||
| 139 | - Verdict `approve` / `request_changes`, or none if invalid. | |
| 140 | - A summary capped at 20,000 characters, signed with the model name. | |
| 141 | - Line comments: empty bodies dropped; **comments on files the PR does | |
| 142 | not change are dropped**; at most **30** comments; one `line` (no | |
| 143 | ranges). | |
| g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent | 144 | - Posted as `g1t`, with a `review.completed` event. |
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 145 | - Confidence (`confidence.rs`) uses the result: `request_changes` sinks |
| 146 | the change's confidence; `approve` with 3 or more comments costs a | |
| 147 | point; no review costs a point. | |
| 148 | - The reviewer reports no confidence or severity of its own. `confidence::ASK` | |
| 149 | is added to implement and revise runs, not to reviews. | |
| 150 | ||
| 151 | ### How it would be scored | |
| 152 | ||
| 153 | The adapter (`bench/reviewbench/agent/agent.mjs`) maps each surviving line | |
| 154 | comment to `{file: path, start_line: line, end_line: line, message: body}`. | |
| 155 | The verdict and summary are kept beside the findings for analysis | |
| 156 | (`findings.g1t.json`) but not scored. ReviewBench scores only located | |
| 157 | findings. | |
| 158 | ||
| 159 | So: | |
| 160 | ||
| 161 | - A problem described only in the summary scores nothing. | |
| 162 | - An empty comment list on a PR with golden TPs costs recall and nothing | |
| 163 | else. | |
| 164 | - Every comment counts toward precision. A comment matching a golden *FP* | |
| 165 | counts against grounded precision; an unmatched one is classified by the | |
| 166 | judge. | |
| 167 | ||
| 168 | Benchmark repositories have no g1t memory, issue, comments or | |
| 169 | `.g1t/review.md`. The prompt is title, body and `INSTRUCTIONS`, plus | |
| 170 | whatever `CLAUDE.md` or `AGENTS.md` the repository carries (Claude Code | |
| 171 | reads `CLAUDE.md` itself). This measures the reviewer's core, which is | |
| 172 | what we want. Memory and instructions are product features to measure | |
| 173 | online. | |
| 174 | ||
| 175 | ## The harness (`bench/reviewbench/`) | |
| 176 | ||
| 177 | | File | What it does | | |
| 178 | |---|---| | |
| 179 | | `run.mjs` | The driver. `setup` clones ReviewBench at the pinned commit and installs its judge. `build` builds the image. `estimate` prices a run without spending. `review` runs our reviewer through ReviewBench's own `scripts/try-agent.sh`, one fresh container per PR, with the benchmark's mounts and validation. `judge` runs ReviewBench's judge with our key. `report` prints a summary or one JSON line. Nothing that spends runs without `--yes`. There is a per-PR cap (`--budget-usd`, default $5, passed to `--max-budget-usd`) and a per-run cap (`--max-total-usd`, default $150). | | |
| 180 | | `Dockerfile` | `FROM` the production sandbox base image named in `services/runner/base.json` (pinned by digest, same Claude Code CLI version), plus the adapter. | | |
| 181 | | `agent/agent.mjs` | The contract adapter: what `review.rs` does after its clone (writes `/work/change.diff`, runs Claude Code with `harness.rs`'s flags, reads `/work/review.json`), then the filters `reviews.rs` applies (changed files only, at most 30, 20,000 characters), then findings. It also writes a `findings.g1t.json` sidecar: verdict, summary, cost, turns, and comments dropped. | | |
| 182 | | `agent/instructions.txt` | Generated from `review.rs` by `prompt.mjs`. `node bench/reviewbench/prompt.mjs --check` fails if production's prompt has moved, which can run in CI. | | |
| 183 | | `agent/variants/*.txt` | Alternative instructions to A/B against `production` (`--variant`). `findings-first.txt` is the first candidate (see below). | | |
| 184 | | `.cache/` | Gitignored: the ReviewBench clone, repository mirrors, runs, scores. | | |
| 185 | ||
| 186 | Model access goes into each container from the environment: | |
| 187 | ||
| 188 | - `ANTHROPIC_API_KEY` alone sends requests straight to the provider. | |
| 189 | - `ANTHROPIC_BASE_URL=<MODELS_URL>/anthropic` plus a model-proxy session | |
| 190 | token sends them through g1t's model proxy. The spend then lands on a | |
| 191 | workspace like any run: use an internal workspace such as flagon-io so | |
| 192 | it is visible in billing. | |
| 193 | ||
| 194 | The judge needs `ANTHROPIC_API_KEY` (or another provider ReviewBench's | |
| 195 | `pi` registry supports). | |
| 196 | ||
| 197 | ```sh | |
| 198 | node bench/reviewbench/run.mjs setup | |
| 199 | node bench/reviewbench/run.mjs build | |
| 200 | node bench/reviewbench/run.mjs estimate --set sample:50 | |
| 201 | node bench/reviewbench/run.mjs review --set sample:50 --seed 1 --yes # production prompt | |
| 202 | node bench/reviewbench/run.mjs review --set sample:50 --seed 1 --variant findings-first --yes | |
| 203 | node bench/reviewbench/run.mjs judge --run <id> --yes | |
| 204 | node bench/reviewbench/run.mjs report --run <id> | |
| 205 | ``` | |
| 206 | ||
| 207 | `sample:N` is stratified by change size in the corpus's proportions | |
| 208 | (≤200 / 201 to 1,000 / >1,000 lines) and fixed by `--seed`, so week-over-week | |
| 209 | runs see the same PRs. The leaderboard's own protocol is `--set full`, three | |
| 210 | times. | |
| 211 | ||
| 212 | Before the first paid run: | |
| 213 | ||
| 214 | 1. Run `build` and one PR (`try-agent.sh … --pr 0`) with a capped key, to | |
| 215 | check that the adapter writes valid findings. This costs about $1. | |
| 216 | 2. Check that the image runs as a user that can write `/work`. The base | |
| 217 | image's default user is used. | |
| 218 | 3. Calibrate `estimateOne` in `run.mjs` against the billing ledger's real | |
| 219 | `review` runs. Every run reports `total_cost_usd`, and spend is kept per | |
| 220 | pull request. | |
| 221 | ||
| 222 | The adapter re-implements about 40 lines of `review.rs` rather than | |
| 223 | calling it, because `review.rs` clones from g1t and posts back to the API. | |
| 224 | To remove that duplication, a small product change would add | |
| 225 | `MODE=review-local` to `g1t-runner`: read `/work/pr/*`, skip the clone and | |
| 226 | the report, and write `RB_OUT`. The adapter would then be `exec g1t-runner`. | |
| 227 | That is worth doing once the benchmark is in regular use. | |
| 228 | ||
| 229 | ### Cost model | |
| 230 | ||
| 231 | Per PR, an agentic review re-reads a growing context from cache every turn. | |
| 232 | `run.mjs` models it as follows: | |
| 233 | ||
| 234 | - turns: 12 + 0.8·√lines + 0.4·files, capped at 80; | |
| 235 | - context: from about 18K tokens plus the diff, growing about 2.5K tokens | |
| 236 | per turn, capped at 180K; | |
| 237 | - 90%+ of input served as cache reads ($0.20/M on Sonnet 5.5); | |
| 238 | - new context written once ($2.50/M); | |
| 239 | - about 450 output tokens per turn ($10/M). | |
| 240 | ||
| 241 | This gives about $0.90 for a median PR and about $1.25 averaged over the | |
| 242 | corpus, which is skewed by the 36% of PRs over 1,000 lines. The judge is | |
| 243 | estimated at about $0.33 per PR: matching calls plus a tool-using | |
| 244 | classification of each unmatched finding. | |
| 245 | ||
| 246 | ## Likely weaknesses against the benchmark | |
| 247 | ||
| 248 | These come from reading the prompt and code; none is measured yet. | |
| 249 | ||
| 250 | 1. **Recall is capped by the instructions.** "Be specific and brief. | |
| 251 | Comment only on real problems" and nothing asking for coverage. The | |
| 252 | golden set has 12 TPs per PR (median 10) and 59% are low severity, which | |
| 253 | the benchmark counts as worth fixing: small correctness slips, missing | |
| 254 | tests, stale docs. The best reviewers emit about 5 findings per PR. We | |
| 255 | probably emit 0 to 3. Expect grounded recall well under 15%. | |
| 256 | 2. **Category blind spots.** The instructions frame the review as | |
| 257 | "correct, does what it's for, doesn't break anything". That covers | |
| 258 | correctness (41% of TPs) and reliability (14%). It gives no prompt for | |
| 259 | maintainability (14%), documentation (12%), testing (9%), security (4%), | |
| 260 | API design, performance or accessibility. | |
| 261 | 3. **Problems that live only in the summary.** Nothing requires every | |
| 262 | issue in `body` to also be a line comment. Those issues score zero, and | |
| 263 | in the product they are not anchored where an author fixes them. | |
| 264 | 4. **Comments outside the changed files are dropped.** `reviews.rs` | |
| 265 | discards them. The benchmark anchors some TPs in untouched files, for | |
| 266 | example a missing dependency in `pyproject.toml`, a caller the change | |
| 267 | breaks, or 173 pre-existing-scope TPs. That is a product decision (the | |
| 268 | UI has nowhere to put them), and it costs recall. | |
| 269 | 5. **Single-line anchors.** 78% of golden TPs span several lines. Matching | |
| 270 | is semantic within a file, so this matters less than the wrong file | |
| 271 | would. A range still helps the matcher and helps a person. | |
| 272 | 6. **Large changes.** 36% of the corpus is over 1,000 lines (median 9 | |
| 273 | files, mean 28). One agent with 80 turns, a 30-minute cap and a | |
| 274 | 30-comment cap will read the first files carefully and skim the rest. | |
| 275 | No fan-out by file group. | |
| 276 | 7. **No verification or calibration.** There is no second pass that checks | |
| 277 | each candidate against the code, so precision rests on the "brief" | |
| 278 | instruction, which also caps recall. There is no per-comment severity or | |
| 279 | confidence, so we cannot pick an operating point (β), and | |
| 280 | `confidence.rs` can only count comments. | |
| 281 | 8. **Context the product has but the reviewer does not get.** Required | |
| 282 | checks and CI results, overlapping PRs, and earlier reviews. These do | |
| 283 | not matter on the benchmark but matter online. | |
| 284 | ||
| 285 | ## Improvements to try, in order | |
| 286 | ||
| 287 | Each is one `--variant` (or a flag) on the same 50-PR sample with the same | |
| 288 | seed. The comparison is grounded precision and recall, by severity and | |
| 289 | category. | |
| 290 | ||
| 291 | 1. **Prompt: findings first** (`agent/variants/findings-first.txt`, | |
| 292 | written). It works in three passes: | |
| 293 | - map the change and its callers; | |
| 294 | - hunt by an explicit category checklist; | |
| 295 | - verify each candidate and drop what isn't true at head. | |
| 296 | ||
| 297 | It asks for every distinct problem rather than "brief", requires every | |
| 298 | summary issue as a line comment, and adds `end_line` and `severity`. | |
| 299 | The adapter already reads `end_line`. Expect the largest recall gain | |
| 300 | for little cost. | |
| 301 | 2. **Self-critique as a separate pass.** A second, cheaper call (Sonnet | |
| 302 | 5.5 at low effort, or Haiku 4.5) gets each candidate plus the exact | |
| 303 | code lines and answers "true at head? worth the author's time?". Keep | |
| 304 | the survivors. This lets the first pass be generous (recall) while | |
| 305 | holding precision. It costs about 10 to 20% more. | |
| 306 | 3. **Context retrieval.** Before the agent starts, compute the changed | |
| 307 | symbols (from diff hunks) and their references (`git grep`, or the | |
| 308 | context service where indexed), and put the list in the prompt. Point | |
| 309 | it at the tests that cover the touched files. This targets the 1,529 | |
| 310 | TPs needing "diff plus related files". | |
| 311 | 4. **Fan-out on large diffs.** Over about 800 lines or 15 files, split by | |
| 312 | directory or file group. Run parallel reviewers (Claude Code subagents, | |
| 313 | or separate runs) with the shared PR context, then merge and dedupe by | |
| 314 | file, line and meaning. This targets item 6, the long tail where recall | |
| 315 | collapses. | |
| 316 | 5. **Confidence calibration.** Have each comment carry `severity` and | |
| 317 | `confidence`. On the benchmark, sweep a threshold to draw the | |
| 318 | precision/recall curve and pick the product's operating point, for | |
| 319 | example post only medium and above inline and fold the low items into | |
| 320 | the summary. Feed the same fields to `confidence.rs`, so that one | |
| 321 | high-severity comment counts for more than three nits. | |
| 322 | 6. **Model and effort.** Run the same sample on Sonnet 5.5 at | |
| 323 | `medium`/`high` effort and on Opus 5.5. The review route can change in | |
| Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier | 324 | `AGENT_ROUTING` without a code change. Opus 5.5 is about 1.4x the cost |
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 325 | per review at our token profile. |
| 326 | 7. **Product-side follow-ups** (not benchmark-visible): | |
| 327 | - allow file-level comments on unchanged files when the change breaks | |
| 328 | them, instead of dropping them; | |
| 329 | - add required-check results and overlapping PRs to the review prompt. | |
| 330 | ||
| 331 | ## Making review quality a tracked metric | |
| 332 | ||
| 333 | **Offline (ReviewBench).** | |
| 334 | ||
| 335 | - A weekly g1t Actions workflow, `.g1t/workflows/reviewbench.yml`, runs | |
| 336 | `review --set sample:50 --seed 1` with the production variant and the | |
| 337 | production model, then `judge` and `report --json`. | |
| 338 | - The runner base image already has the toolchains. The job needs Docker, | |
| 339 | or, more simply, it can run `agent.mjs` directly in a sandbox that *is* | |
| 340 | the same image, with `/work` laid out by a small wrapper instead of | |
| 341 | `try-agent.sh`. | |
| 342 | - Budget: about $80 a week (about $350 a month), billed to the flagon-io | |
| 343 | workspace through the model proxy so it shows in spend. | |
| 344 | - Also run on demand for any change to `review.rs`'s instructions or to | |
| 345 | the review route. `prompt.mjs --check` in CI flags such a change. | |
| 346 | - Do a full-set run (about $350) monthly, or before a model switch. | |
| 347 | ||
| 348 | **Where results go.** | |
| 349 | ||
| 350 | - Append each `report --json` line to an internal results store: a D1 | |
| 351 | table in the work service, or a JSON file committed to `docs/research/`. | |
| 352 | - Show it on a **sudo "Review quality"** page, with a trend of: | |
| 353 | - grounded precision and recall; | |
| 354 | - recall by severity (high and medium matter most) and by category; | |
| 355 | - findings per PR; | |
| 356 | - cost per PR. | |
| 357 | ||
| 358 | Mark the g1t commit and model on each point. | |
| 359 | - Not on public docs: our own numbers are fine internally, but | |
| 360 | user-facing copy compares to no one. | |
| 361 | ||
| 362 | **Online, the metric that matters.** Track the *addressed rate* of | |
| g1t is one name: its agent's work, commits and comments show as @g1t, and nobody can claim g1t or g1t-agent | 363 | `g1t` line comments: the share followed by a commit touching those |
| Research: shipping CSS, and a ReviewBench harness for g1t's reviewer | 364 | lines before merge, or resolved by a person. Also track the share of |
| 365 | `request_changes` verdicts that led to a revision. g1t has the comments, | |
| 366 | the commits and the review runs, so this is a query, not a model call. Put | |
| 367 | it next to the offline number on the same sudo page. The article's lesson | |
| 368 | is that the offline number earns trust only while it moves the same way as | |
| 369 | the online one. | |
| 370 | ||
| 371 | ## Open questions | |
| 372 | ||
| 373 | - Run once on the public leaderboard? Onboarding is self-service (a GHCR | |
| 374 | image and our key). The final run is three rounds on 219 PRs at our | |
| 375 | inference cost (about $820) with their judge free. A ranking is | |
| 376 | marketing-adjacent, and the no-comparisons rule applies to anything we | |
| 377 | say about it. | |
| 378 | - The judge is a Claude model and so is our reviewer. The ReviewBench docs | |
| 379 | report human agreement for the labels but no per-family judge bias. | |
| 380 | Treat small differences (under 2 points, inside the leaderboard's | |
| 381 | round-to-round deviation of about 0.5 to 2) as noise. |