Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 1 | # Performance |
| 2 | ||
| 3 | How g1t.sh answers a page: where the code runs, where the data lives, how | |
| 4 | reads and writes travel, what is cached, and the budget pages are held to. | |
| 5 | Internal. The tools: `scripts/perf/measure.ps1` (time pages from your | |
| 6 | machine), `scripts/perf/placement-probe.mjs` (measure placements without | |
| 7 | touching production), and the Server-Timing header on every page. | |
| 8 | ||
| 9 | ## The shape of a request | |
| 10 | ||
| 11 | ``` | |
| 12 | browser ──► Cloudflare edge (nearest data centre) | |
| 13 | │ | |
| 14 | ▼ | |
| 15 | g1t (apps/web, React Router) root loader + layout + page loaders, in parallel | |
| 16 | │ service bindings: POST /rpc/<method>, JSON | |
| 17 | ▼ | |
| 18 | g1t-identity, g1t-repos, g1t-work, g1t-projects, g1t-billing, … (services/*) | |
| 19 | │ | |
| 20 | ▼ | |
| 21 | D1 (one SQLite database per service; primary in WNAM, US West) | |
| 22 | Artifacts (git objects and refs, for g1t-repos) | |
| 23 | ``` | |
| 24 | ||
| 25 | - **The site holds no data.** Every loader calls services; every service | |
| 26 | owns one D1 database. A page is a few rounds of service calls; each call | |
| 27 | is a few D1 queries. | |
| 28 | - **Latency is round trips times distance.** A query from a Worker next to | |
| 29 | its database takes 1 to 5 ms. The same query from Amsterdam to WNAM took | |
| 30 | about 150 ms. A page with three rounds of calls, each with two or three | |
| 31 | queries in turn, costs about a second and a half when the code and the | |
| 32 | data are on different continents. That was the slowness. | |
| 33 | ||
| 34 | ## Where the code runs | |
| 35 | ||
| 36 | | Worker | Placement | Why | | |
| 37 | | --- | --- | --- | | |
| 38 | | `g1t` (site), `g1t-api`, `g1t-sudo` | `off` | Run next to the person. With D1 replicas (below), most reads are local too. | | |
| 39 | | identity, repos, work, search, billing, projects, deployments | `off` | They read with D1 sessions: the nearest replica, or the primary when consistency needs it. | | |
| 40 | | actions, events, webhooks, integrations, context, security | `off` now; a region near WNAM once the probe has picked one | They read the primary only. Pinned beside it, a page from Europe pays one ocean crossing per call instead of one per query. | | |
| 41 | | pages, models, og, runner, status, docs | none | Edge-serving or no data of their own. | | |
| 42 | ||
| 43 | ### What happened with Smart Placement | |
| 44 | ||
| 45 | Every data-holding Worker and the site had `"placement": { "mode": "smart" }`. | |
| 46 | Production answered with `cf-placement: remote-AMS` for a visitor in Denver: | |
| 47 | the site ran in Amsterdam, and so did the services it called through | |
| 48 | bindings (a binding runs the callee where the caller is unless the callee | |
| 49 | is placed). Every D1 query then crossed the Atlantic. Smart Placement only | |
| 50 | considers locations where the Worker has already run and needs traffic | |
| 51 | from many places to decide, so it settled on a poor spot and stayed there. | |
| 52 | Measured from Colorado, signed out, warm: | |
| 53 | ||
| 54 | | Page | Smart (AMS) | Off | | |
| 55 | | --- | --- | --- | | |
| 56 | | `/` | 270 ms | 140 ms | | |
| 57 | | `/explore` | 850 ms | 170 ms | | |
| 58 | | `/flagon-io/g1t/pulls` | 600 ms | 220 ms | | |
| 59 | | `/flagon-io/g1t/issues` | 780 ms | 230 ms | | |
| 60 | | `/pricing` | 500 ms | 120 ms | | |
| 61 | | `/flagon-io/g1t` (overview) | 1.4 s | 1.67 s (its own problem, below) | | |
| 62 | ||
| 63 | Placement now applies only to `fetch` handlers; the services are reached | |
| 64 | through `fetch` (`/rpc/<method>`), so it applies to them. | |
| 65 | ||
| 66 | ### Choosing a region (the probe) | |
| 67 | ||
| 68 | Cloudflare's placement hints (`"placement": { "region": "aws:us-west-1" }`, | |
| 69 | GCP and Azure regions too, Wrangler 4.146 accepts them) run a Worker next | |
| 70 | to a cloud region. D1 is not a cloud region, and Cloudflare does not say | |
| 71 | which city WNAM is, so measure: | |
| 72 | ||
| 73 | ```powershell | |
| 74 | node scripts/perf/placement-probe.mjs deploy # g1t-probe-* on workers.dev, SELECT 1 against g1t-repos | |
| 75 | node scripts/perf/placement-probe.mjs measure # a table: where each ran, ms per D1 query, through a binding too | |
| 76 | node scripts/perf/placement-probe.mjs delete | |
| 77 | ``` | |
| 78 | ||
| 79 | Pick the region with the lowest **D1 primary ms/query**. To pin the | |
| 80 | primary-only services there, edit their configs (or use | |
| 81 | `node scripts/perf/placement-probe.mjs apply '{"region":"<it>"}'`, which | |
| 82 | sets every config it lists, then put `"mode": "off"` back on the site, API, | |
| 83 | sudo and the session services). `apply '{"mode":"off"}'` is what the | |
| 84 | working tree has now. | |
| 85 | ||
| 86 | ## Where the data lives, and how reads travel | |
| 87 | ||
| 88 | Every database's primary is in WNAM. D1 read replication puts read-only | |
| 89 | copies in every region (ENAM, WNAM, WEUR, EEUR, APAC, OC) at no extra | |
| 90 | cost. A copy trails the primary, so reading one needs care. | |
| 91 | ||
| 92 | ### Sessions and bookmarks | |
| 93 | ||
| 94 | Seven services read through D1's Sessions API when asked | |
| 95 | (`crates/kit/src/d1.rs`, `packages/contracts/src/d1.ts`). The caller asks | |
| 96 | with the `x-d1-bookmark` request header: | |
| 97 | ||
| 98 | | Header | Reads go to | | |
| 99 | | --- | --- | | |
| 100 | | absent | the primary, with no session: exactly as before | | |
| 101 | | `first-primary` | the primary, then any copy at least as new | | |
| 102 | | `first-unconstrained` | the nearest copy | | |
| 103 | | a bookmark | any copy at least as new as the bookmark | | |
| 104 | ||
| 105 | Writes always go to the primary. A session is sequentially consistent: it | |
| 106 | reads its own writes. The service returns the session's latest bookmark | |
| 107 | in `x-d1-bookmark`. | |
| 108 | ||
| 109 | **Only the site asks for anything.** Service-to-service calls | |
| 110 | (`g1t_kit::call`), queues, crons, the API and MCP send no header, so they | |
| 111 | read the primary as they always did. Billing's `can_start`/`start_run`, | |
| 112 | credential checks for git, and everything agents do stay on the primary. | |
| 113 | ||
| 114 | The site decides per call (`apps/web/app/lib/perf.ts`, `sessionFor`): | |
| 115 | ||
| 116 | 1. A request that writes (any method but GET and HEAD) starts every session | |
| 117 | on the primary, so what an action checks before writing is current. | |
| 118 | 2. Within 30 seconds of the person's last write, every service reads its | |
| 119 | primary. A write can reach a service the site did not call (work | |
| 120 | writing to repos during a merge) whose bookmark the site never sees; | |
| 121 | replicas trail by well under a second, so 30 seconds is a wide margin. | |
| 122 | 3. Otherwise, the bookmark that service returned after the last write. | |
| 123 | 4. Otherwise, the nearest copy. | |
| 124 | ||
| 125 | After a request that may have written, the site sets the `g1t_d1` cookie: | |
| 126 | `at:<unix seconds>` and `service:<bookmark>` pairs, HttpOnly, five minutes. | |
| 127 | "May have written" is: a non-GET request; a GET that called a method not | |
| 128 | on the known-read list (`READS` in `perf.ts`; anything new counts as a | |
| 129 | write until listed); or a GET that started a session (signing in with | |
| 130 | GitHub). GETs that only read set no cookie, so public pages stay cacheable. | |
| 131 | ||
| Git storage hardened, pages in tens of milliseconds, honest security alerts, and costs reconciled daily | 132 | Keep `READS` complete. Until 2026-10-06 it lacked `get` (repos and |
| 133 | projects), `list`, `queue` and `pulls_for_repos`, so every project page | |
| 134 | and Mission control looked like a write: each set the cookie, which kept | |
| 135 | signed-out project pages out of the public cache (every view rendered, | |
| 136 | 0.4 to 0.6 s, crawlers included), sent the person's next 30 seconds of | |
| 137 | reads to the primary, and turned off the sidebar cache | |
| 138 | (`mustReadFresh`). `scripts/perf/measure.ps1` shows a **Sets g1t_d1** | |
| 139 | column: it should say False for every page it measures. | |
| 140 | ||
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 141 | What a person can still see out of date: something someone else (or an |
| 142 | agent, or the API) changed in the last fraction of a second, which a page | |
| 143 | would have missed by loading a moment earlier anyway; and a session | |
| 144 | revoked from another device working for that same fraction of a second | |
| 145 | on reads (sign-out from this browser is a write, so it is immediate). | |
| 146 | ||
| 147 | ### Turning replication on | |
| 148 | ||
| 149 | Not on yet: the code above works the same without it (every read is then | |
| 150 | the primary). Turn it on per database once the site and the seven services | |
| 151 | are deployed with sessions. There is no Wrangler command; use the | |
| 152 | dashboard (**D1 → database → Settings → Read replication → Enable**) or | |
| 153 | the API with a token that has D1 Edit: | |
| 154 | ||
| 155 | ```powershell | |
| 156 | $token = $env:CLOUDFLARE_D1_TOKEN # D1: Edit on account syntaqx | |
| 157 | $account = "1e6f2cffa3f445920836e8ebe446bb58" | |
| 158 | $databases = @{ | |
| 159 | "g1t" = "b7d49c93-2666-4006-a3c3-073a01838dc9" # identity | |
| 160 | "g1t-repos" = "f9544c51-c3bf-4621-a96f-8a6d5cf24a97" | |
| 161 | "g1t-work" = "f35a9022-e36b-4547-b596-9ab9d5f1c47a" | |
| 162 | "g1t-projects" = "0af698be-2b60-4bd9-aa93-81e84828e991" | |
| 163 | "g1t-deployments" = "aa935a8f-845b-4132-8fa3-98f1afba1db5" | |
| 164 | "g1t-search" = "9d04cbf3-cd02-4544-b3e3-2d78767e13bf" | |
| 165 | "g1t-billing" = "695a6979-fd07-4850-bc97-904a6b4b7a04" | |
| 166 | } | |
| 167 | foreach ($name in $databases.Keys) { | |
| 168 | curl.exe -s -X PUT "https://api.cloudflare.com/client/v4/accounts/$account/d1/database/$($databases[$name])" ` | |
| 169 | -H "Authorization: Bearer $token" -H "Content-Type: application/json" ` | |
| 170 | --data '{\"read_replication\":{\"mode\":\"auto\"}}' | |
| 171 | Write-Host "" | |
| 172 | } | |
| 173 | npx wrangler d1 info g1t-repos # read_replication: { mode: "auto" } | |
| 174 | ``` | |
| 175 | ||
| 176 | Turning it off is `{"read_replication":{"mode":"disabled"}}` and takes up | |
| 177 | to a day to finish. Read-heavy over the last 24 hours: g1t-repos (60,103 | |
| 178 | reads to 111 writes), g1t-work (43,219 / 1,462), g1t-projects (28,861 / | |
| 179 | 22), g1t-billing (20,600 / 507), g1t (5,434 / 202). g1t-search writes more | |
| 180 | than it reads (indexing), so replicas help it least. | |
| 181 | ||
| 182 | ## Caching | |
| 183 | ||
| 184 | | What | Where | For how long | Rules | | |
| 185 | | --- | --- | --- | --- | | |
| 186 | | Static assets (`/assets/*`) | browser and edge | a year, immutable | hashed file names | | |
| 187 | | Avatars | edge cache | a year, immutable | by content hash | | |
| Signed-out page cache: a repository's kept page is served only while the repository is still public, so one made private or deleted never shows from any data centre's copy | 188 | | Public pages for people signed out | the data centre's cache (`workers/app.ts`, `servePublic`) | fresh 30 s, then served once more while a new copy is made, up to 5 min | GET, no `g1t_session` cookie, an allowlisted path (home, pricing, explore, policies, a project's pages), status 200 or 404, no `Set-Cookie`, nothing private. Reserved first segments and workspace pages (`-`) are never kept. A project's kept page is served only after repos' `visibility` says the repository is still there and public (one indexed read, alongside the cache lookup); a repository made private or deleted is never served from any data centre's copy, and the copy is dropped. The answer says `server-timing: cache;desc="hit, Ns old"`. | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 189 | | Sidebar data (projects, spend, limit, entitlements) | per isolate (`lib/cache.server.ts`) | 15 s, per person and workspace | skipped during a write and for 30 s after the person's last one; failures not kept; only settled answers kept | |
| 190 | | Registration mode | per isolate | 60 s | | | |
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 191 | | A commit's log by hash | repos' data-centre cache | for good | history from a commit never changes. One of 100 commits or more is put together from a 16-commit read and the history kept from any of those commits, when there is one (`store.rs` `spliced_log`): a default branch that moved by a merge costs 16 commits, not 120 or 1,000 | |
| Merge branch drift: count across merges the way git does; v2 cache key | 192 | | A branch's drift from the default branch (Active branches) | repos' data-centre cache (`branch_drift`) | a count for good; "too far to count" a day | by repository and the pair of head commits (key version `v2`); a failed read is not kept | |
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 193 | | A repository's tags | repos' data-centre cache | until the refs move, 5 min at most | as the branch list; not kept when a tag's commit could not be read | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 194 | | Git objects, trees, refs | repos' caches | see services/repos | | |
| Git storage hardened, pages in tens of milliseconds, honest security alerts, and costs reconciled daily | 195 | | A branch's log, the branch list, a file by branch and path | repos' data-centre cache | until the repository's refs change (`refs_version`), 5 min at most | only while no handed-out push credential is live; by commit hash for good (docs/ARTIFACTS.md R9) | |
| 196 | | A target branch's history, for mergeability | the repos isolate | 60 s, per target head | 100 pull requests checked after a push walk it once (R10) | | |
| 197 | | Git store credentials | repos isolate and KV | reused 50 min (1 h tokens); 3 min for ones handed out | (R3) | | |
| React Router's handler is made once per isolate, and docs/PERFORMANCE.md says where the site's CPU went and how to measure it | 198 | | A file's highlighted lines (blob, blame) | the site's isolate (about 8 MB), then the data centre's cache (`content.g1t.internal/highlight-lines/`) | for good (30 days in the data centre) | by SHA-256 of the language and the text, nothing else: the same text on any branch, commit or page is one entry. `HIGHLIGHT_VERSION` in `lib/highlight.server.ts` is in the key; bump it when the theme, grammars, Shiki or `linesToHtml` change. Nothing kept for a file without a language or over 200,000 characters | |
| 199 | | A pull request's first highlighted files | as above (`highlight-diff/`, about 4 MB per isolate) | for good | by the language and each line's side and text (`diffContent` in `lib/diff.ts`); line numbers and the path are not in it | | |
| 200 | | Parsed markdown | the isolate, or the browser tab (400,000 characters of source, about 8 MB) | until pushed out | by the text and the repository its references point into (`components/markdown.tsx`); the tree is rendered with the page's components each time | | |
| 201 | | React Router's route tables | the isolate | its life | the build is loaded once and the handler made once (`workers/app.ts`); development still reloads it per request | | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 202 | |
| 203 | ## Server-Timing | |
| 204 | ||
| 205 | Every page and `.data` response from the site carries a `Server-Timing` | |
| 206 | header (DevTools → Network → the request → Timing): | |
| 207 | ||
| 208 | ``` | |
| 209 | total;dur=180;desc="web to first byte", | |
| 210 | loader.root;dur=40, loader.repo.layout;dur=60, loader.repo.pull;dur=150, | |
| 211 | rpc;dur=140;desc="9 service calls, overlap counted once", | |
| 212 | work;dur=120;desc="4 calls, 70ms inside", repos;dur=30;desc="2 calls, 12ms inside", …, | |
| 213 | d1;desc="work=unconstrained repos=bookmark identity=primary" | |
| 214 | ``` | |
| 215 | ||
| 216 | - `total`: from the request reaching the site to the response headers. | |
| 217 | Streamed panels finish after it. | |
| 218 | - `loader.<route>` / `action.<route>`: each route's loader or action | |
| 219 | (React Router instrumentation in `app/entry.server.tsx`). | |
| 220 | - `rpc`: time waiting on services, overlapping calls counted once. | |
| 221 | - One entry per service: summed wall time of its calls from the site, | |
| 222 | and how much of that the service itself reported (`svc;dur` from | |
| 223 | `Served::finish`). The difference is the trip between them. | |
| 224 | - `d1`: how each session-capable service was asked to read. | |
| Git storage hardened, pages in tens of milliseconds, honest security alerts, and costs reconciled daily | 225 | - Inside a service entry, `db Nms in T round trips`: what the service |
| 226 | reported waiting on its own database (`db;dur`, from | |
| 227 | `g1t_kit::d1::Timing` and `Served::finish_timed`). The work service | |
| 228 | reports it, and `rpc;dur` for its own calls to other services; a service | |
| 229 | call's own response carries both beside `svc;dur`. | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 230 | |
| 231 | Git requests keep their own header (`repos;dur` plus the repos service's | |
| Git storage hardened, pages in tens of milliseconds, honest security alerts, and costs reconciled daily | 232 | steps). Counting git operations writes nothing on the way: the meters are |
| 233 | written after the answer (`wait_until`), so `kept` no longer includes a D1 | |
| 234 | upsert (63–98 ms before; docs/ARTIFACTS.md R13). Mission control keeps its per-section timings. | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 235 | |
| 236 | ## What a page does, in rounds | |
| 237 | ||
| 238 | Rounds are what cost: calls in the same round overlap. | |
| 239 | ||
| 240 | | Page | Before | Now | | |
| 241 | | --- | --- | --- | | |
| 242 | | Any in-app navigation | root (sidebar: 6 calls, then up to 20 `get_by_id` for shared repositories) and the project layout re-ran when moving between pages of a project | root re-runs only when the workspace or project changes, or after a form; the project layout likewise; shared repositories are one `readable` call, in the same round; the sidebar's workspace data is cached for 15 s; open counts are read once per request for both | | |
| 243 | | Pull request ("Review and respond") | access, then 8 calls, then checks' runs / comparison / session, then up to 5 more deployments lookups for stacked previews | one round of 9 (access-dependent ones start as soon as the repository lookup returns), then the comparison on Changes; the workflow jobs and stacked previews stream in | | |
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 244 | | Project overview | access, then 17 calls, one of which (Active branches) read the default branch's last 120 commits and up to 10 branches' last 40 | one round; Active branches streams in with a skeleton, in one `branch_drift` call (below) | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 245 | | Mission control | per project: open pulls, closed pulls and events (3 × up to 10), then `get_by_id` per unknown repository | one `pulls_for_repos` call for every project (one access check, one query), events per project alongside, one `readable` for the rest | |
| 246 | | Issue, issues | access, then the rest | one round | | |
| 247 | ||
| Git storage hardened, pages in tens of milliseconds, honest security alerts, and costs reconciled daily | 248 | ### Inside the work service |
| 249 | ||
| 250 | Before 2026-10-06 a pull request's page cost the work service about | |
| 251 | twenty D1 round trips one after another, plus two calls to repos: the | |
| 252 | access check (`get`), then the pull request, its issue, then its | |
| 253 | lifecycle read progress, latest review, settings, statuses, review | |
| 254 | comments, the confidence signals (three in turn), approvals, requests | |
| 255 | for changes, the queue entry, and wrote the stage back on every view, | |
| 256 | then statuses and settings again, messages and earlier checks. It also | |
| 257 | asked repos `behind` on every view, which walks up to `MAX_ANCESTRY` | |
| 258 | commits in Artifacts. About 350 ms of the page's 0.5 s. | |
| 259 | ||
| 260 | Now (`services/work/src/prefetch.rs`): | |
| 261 | ||
| 262 | - **One batch.** Everything the page and the lifecycle read is one D1 | |
| 263 | batch of 20 statements keyed by repository id and number (subqueries | |
| 264 | find the pull request's id, issue and head). The helpers that decide | |
| 265 | the lifecycle (`settings`, `statuses`, `review_pending`, | |
| 266 | `approvals_gap`, `signals`, …) read from it when it is there, so the | |
| 267 | decision is the same code either way. | |
| 268 | - **Beside the access check.** `repo_then` starts the batch with the | |
| 269 | repository id this isolate last saw for the path, at the same time as | |
| 270 | repos' `get`; the rows are used only if `get` then allows that same | |
| 271 | repository, and read again otherwise. Issues, lists, counts, labels and | |
| 272 | settings do the same. `pulls_for_repos` reads its rows beside | |
| 273 | `readable` and drops those of repositories the viewer cannot read. | |
| 274 | - **Precomputed `behind`.** `pulls.behind` is written with mergeability | |
| 275 | on every push to either side (migration 0023); a view reads it when it | |
| 276 | was worked out for the current head and asks repos only otherwise | |
| 277 | (then keeps the answer). | |
| 278 | - **No write on a view** unless the stage changed. | |
| 279 | - `list_active_pulls` reads its pull requests' issues in the same batch | |
| 280 | instead of one query each. | |
| 281 | ||
| 282 | A pull request is now the repos `get` (about 40 ms) and one batch beside | |
| 283 | it. The indexes were checked with `EXPLAIN QUERY PLAN` against the | |
| 284 | migrations: every statement is an index search; 0023 adds | |
| 285 | `agent_messages_by_sender` for the unanswered-questions count. | |
| 286 | ||
| 287 | ### The overview streams | |
| 288 | ||
| 289 | `routes/repo/overview.tsx` returns its seventeen calls as one deferred | |
| 290 | promise. The layout's header and tabs (repository and project, two | |
| 291 | cheap calls) and a skeleton go out first; the sections follow in the | |
| 292 | same response. Crawlers still get the whole page (`entry.server.tsx` | |
| 293 | waits for `allReady` for bots), and signed out it is kept in the public | |
| 294 | cache like any other project page. | |
| 295 | ||
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 296 | ### Active branches (2026-10-08) |
| 297 | ||
| 298 | Measured on production, `/flagon-io/g1t` uncached, as a crawler, soon | |
| 299 | after pushes: **3,606 ms** to the first byte, `rpc` 3,605 ms over 38 | |
| 300 | service calls, **repos 25 calls, 17,953 ms inside**. Nearly all of it was | |
| 301 | Active branches: | |
| 302 | ||
| 303 | - The site measured each branch itself: a `log` call per branch per | |
| 304 | depth (40, then 1,000), plus the default branch's (120, then 1,000), plus | |
| 305 | one more for the default branch's head, each call with its own access | |
| 306 | check and store handle. The answer was kept by the pair of heads in the | |
| 307 | site's cache, so every push to a branch, and every merge to the default | |
| 308 | branch (which changes every pair), started the walks again, in every | |
| 309 | data centre. | |
| 310 | - A branch far from the default branch read 1,000 commits of both, and | |
| 311 | the default branch's 1,000 again after each merge. | |
| 312 | - The page waits 3.5 s for the section, then shows a link to Branches. | |
| 313 | The walks were not in `waitUntil`, so when the page stopped waiting | |
| 314 | they were dropped with it: an answer that took longer than 3.5 s was | |
| 315 | never kept, and the next view started over. 3,606 ms is that timeout. | |
| 316 | ||
| 317 | Now: | |
| 318 | ||
| 319 | - **One call.** `branch_drift` (services/repos/src/drift.rs) takes the | |
| 320 | default branch's head and every branch head, checks access once, opens | |
| Merge branch drift: count across merges the way git does; v2 cache key | 321 | the store once, and reads the default branch's last 120 commits once |
| 322 | for all of them. Each count is kept in repos' data-centre cache by the | |
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 323 | pair of hashes; only pairs that changed are walked. It also returns the |
| 324 | default branch's head commit, which the site read with its own call. | |
| Merge branch drift: count across merges the way git does; v2 cache key | 325 | - **Only what the count needs.** The walk goes newest commit first from |
| 326 | both heads, as `git rev-list --left-right --count` does, and stops once | |
| 327 | everything left is reached by both. The store lists first parents only, | |
| 328 | so a merge's other parent is read on its own (16 commits, by hash, kept | |
| 329 | for good) when the walk reaches it; a branch a few commits from the | |
| 330 | default branch costs one read of its own. Past 128 reads or 4,000 | |
| 331 | commits there is no count, and that answer is kept a day, not for good. | |
| 332 | - **Fixed 2026-10-09: counts never showed.** The first version (and the | |
| 333 | site's before it) gave up when a commit on one side only had a parent | |
| 334 | not read, which every merge on the default branch has, and kept "no | |
| 335 | count" for good: no branch of flagon-io/hello or flagon-io/g1t showed | |
| 336 | counts. The cache key moved to `drift.g1t.internal/v2/`, so those | |
| 337 | answers are not read again. | |
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 338 | - **Long histories spliced.** A log by hash of 100 commits or more is a |
| 339 | 16-commit read plus the log kept from one of those commits (the | |
| 340 | first-parent chain from a commit never changes), so the default branch | |
| Merge branch drift: count across merges the way git does; v2 cache key | 341 | after a merge costs 16 commits instead of 120 or 1,000. A kept log is |
| 342 | used only when it starts at that commit and goes on to the one the | |
| 343 | 16-commit read lists next (`splice_first`), so the join neither repeats | |
| 344 | nor skips a commit. | |
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 345 | - **Finished after the page.** The call runs in `waitUntil`, so repos |
| 346 | keeps the answer even when the page stopped waiting for it. | |
| 347 | - **Crawlers wait 0.7 s** for the section (browsers 3.5 s, streamed); | |
| 348 | past that they get the link to Branches, as a slow browser does. | |
| 349 | - `tags` is kept until the refs move (it listed the refs from the store | |
| 350 | on every call), and the project is looked up once per request for the | |
| 351 | layout and the overview (`projectFor`, beside `repoFor`). | |
| 352 | - `tags`, `commit_checks`, `shortcuts` and the Files page's reads were | |
| 353 | missing from `READS`, so every signed-in overview counted as a write: | |
| 354 | it set `g1t_d1`, sent the next 30 seconds of reads to the primary and | |
| 355 | turned off the sidebar cache. | |
| 356 | ||
| 357 | | Overview, signed out, uncached | Before | After | | |
| 358 | | --- | --- | --- | | |
| 359 | | repos calls | 25 (6 when nothing had moved) | 6 whatever moved: `get`, `stars`, `branches`, `log`, `tags`, `branch_drift` | | |
| 360 | | projects calls | 3 (`get` twice) | 2 | | |
| 361 | | Crawler, after a push | 3,606 ms (the 3.5 s timeout) | at most about 0.8 s: the rest of the page, or 0.7 s for Active branches | | |
| 362 | | Crawler, nothing moved | 460 to 570 ms | not yet measured on production | | |
| 363 | | Browser, first byte | 115 to 160 ms (`total`), 200 to 245 ms measured from Colorado | unchanged: the page does not wait for any of this | | |
| 364 | ||
| React Router's handler is made once per isolate, and docs/PERFORMANCE.md says where the site's CPU went and how to measure it | 365 | ## CPU per page (2026-10-09) |
| 366 | ||
| 367 | Workers bill CPU time past 30 million ms a cycle. From 1 to 9 October the | |
| 368 | site (`g1t`) used 83.3 million CPU-ms over 1.93 million requests, about 43 | |
| 369 | ms a request and seven tenths of all g1t's Workers CPU; `g1t-repos` used | |
| 370 | 28.3 million over 6.58 million (about 4 ms). Signed-out pages are kept | |
| 371 | for 30 s (above), but a crawler reads each file once, so most of its | |
| 372 | requests render. | |
| 373 | ||
| 374 | ### Measuring it | |
| 375 | ||
| 376 | `Server-Timing` cannot show CPU: a Worker's clock does not move while it | |
| 377 | computes, so `total;dur=0` on a page that rendered for 20 ms is normal. | |
| 378 | Measure locally instead, with the built site and fake services: | |
| 379 | ||
| 380 | 1. `npm run build -w apps/web`. | |
| 381 | 2. Load `build/server/index.js` in Node with `cloudflare:workers` pointed | |
| 382 | at a stub whose `env` has a service binding per service, answering | |
| 383 | `POST /rpc/<method>` from fixtures. A page's `.data` (fetched signed | |
| 384 | out from production) decoded with React Router's turbo-stream decoder | |
| 385 | gives realistic answers; files and READMEs can come from the working | |
| 386 | tree. | |
| 387 | 3. Call the worker's `fetch` with a crawler's user agent (the whole page | |
| 388 | renders before the answer) and read `process.cpuUsage()` over 100 | |
| 389 | requests after a few to warm up. Run Node with `--single-threaded` so | |
| 390 | garbage collection and compilation count on the one thread, as in a | |
| 391 | Worker; on Windows `cpuUsage` moves in 15.6 ms steps, so divide a long | |
| 392 | run, never time one request. | |
| 393 | 4. `node --cpu-prof` on the same loop says where it goes. | |
| 394 | ||
| 395 | ### Where it went | |
| 396 | ||
| 397 | Profiled with fixtures from flagon-io/g1t: | |
| 398 | ||
| 399 | | Page | CPU a request | Where | | |
| 400 | | --- | --- | --- | | |
| 401 | | A 130-line TypeScript file | 30 ms | 63% highlighting (Shiki's tokenizer), 17% rendering | | |
| 402 | | A 1,800-line TSX file (2.4 MB page) | 250 ms | 85% highlighting; the rest rendering and encoding the page | | |
| 403 | | The Files page with a README | 32 ms | 46% parsing the README (remark, rehype-raw, sanitize, the plugins) | | |
| 404 | | Any page | 1 to 1.5 ms more | React Router rebuilt its route tables for every request: given the build as a function, it wraps every route for the timings and flattens and ranks the route table again each time | | |
| 405 | | `package-lock.json` (3.8 MB page, too large to highlight) | 170 ms | rendering one row per line and the file again in the page's data | | |
| 406 | ||
| 407 | The CSP nonce, `isbot`, Server-Timing bookkeeping and the signed-out | |
| 408 | cache's own work were each under 1% of a page's CPU in the profiles. | |
| 409 | ||
| 410 | ### What changed | |
| 411 | ||
| 412 | - **Highlighting is kept by content** (`lib/highlight.server.ts`, | |
| 413 | `lib/content-cache.ts`): the isolate first, then the data centre's | |
| 414 | cache, then Shiki. The key is a SHA-256 of the language and the text, | |
| 415 | with a version, so a file that is the same on another branch or commit, | |
| 416 | in blame, or for the next crawler is highlighted once per data centre. | |
| 417 | A pull request's first files are kept the same way. | |
| 418 | - **Markdown is parsed once per text** (`lib/markdown-tree.ts`, | |
| 419 | `components/markdown.tsx`): the steps `react-markdown` runs on every | |
| 420 | render are split, and the parsed tree is kept per isolate (and per | |
| 421 | browser tab). `lib/markdown-tree.test.ts` renders README.md, this file | |
| 422 | and a set of edge cases (raw HTML, scripts, `javascript:` links, alerts, | |
| 423 | references) both ways and checks the HTML is identical. | |
| 424 | - **The request handler is made once per isolate** (`workers/app.ts`). | |
| 425 | - Shiki's module is asked for once per isolate, not on every highlight. | |
| 426 | ||
| 427 | ### Measured | |
| 428 | ||
| 429 | Locally, CPU a request, median of three runs of 100 requests, signed out | |
| 430 | as a crawler. "Seen" is a file or README this isolate (or data centre) | |
| 431 | has highlighted or parsed before; "new" is one it has not. | |
| 432 | ||
| 433 | | Page | Before | After, seen | After, new | | |
| 434 | | --- | --- | --- | --- | | |
| 435 | | `/` (landing) | 12.7 ms | 10.6 ms | 10.6 ms | | |
| 436 | | `/pricing` | 4.8 ms | 3.3 ms | 3.3 ms | | |
| 437 | | Files, root with README.md | 30.8 ms | 9.5 ms | 23.7 ms | | |
| 438 | | Files, `docs/` (a 30,000-character README) | 44.7 ms | 8.9 ms | 33.1 ms | | |
| 439 | | Files, no README | 9.2 ms | 7.2 ms | 6.9 ms | | |
| 440 | | A 130-line TypeScript file | 27.3 ms | 7.8 ms | 24.5 ms | | |
| 441 | | A 1,800-line TSX file | 255.6 ms | 41.3 ms | 246.7 ms | | |
| 442 | | A 560-line Rust file | 34.7 ms | 14.8 ms | 30.0 ms | | |
| 443 | | A markdown file's source | 18.4 ms | 11.6 ms | 18.8 ms | | |
| 444 | | `package-lock.json` | 200.2 ms | 134.1 ms | 127.8 ms | | |
| 445 | ||
| 446 | A pull request's first screens: a 288-line diff took 18.8 ms to | |
| 447 | highlight and 0.15 ms to read back from the data centre's cache. Hashing | |
| 448 | the text for the key is part of every "new" figure above. Small | |
| 449 | differences (a few ms) are within the noise of these runs; the large | |
| 450 | saving on `package-lock.json`, which nothing here caches, is partly that | |
| 451 | noise and partly less garbage per request. | |
| 452 | ||
| 453 | What is left on large files is rendering: a row per line, and the same | |
| 454 | lines again in the page's data for hydration. Files over 200,000 | |
| 455 | characters (lock files) are not highlighted and still cost about 130 ms | |
| 456 | for a crawler. | |
| 457 | ||
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 458 | ## Client navigation |
| 459 | ||
| 460 | - `<Link prefetch="intent">` on the sidebar, project tabs, breadcrumbs, | |
| 461 | list rows and Mission control rows: hovering loads the next page's code | |
| 462 | and data. | |
| 463 | - A thin progress bar while a navigation is pending (`Progress` in | |
| 464 | `components/shell.tsx`). | |
| 465 | - Mission control's code loads while the browser is idle after the app | |
| 466 | shell paints; its skeleton is rarely seen. | |
| 467 | - Mission control's quick actions show as done when sent and undo on | |
| 468 | failure. | |
| 469 | - Skeletons (`components/ui/skeleton.tsx`: `Skeleton`, `SkeletonText`, | |
| 470 | `SkeletonRows`) stand in, at the same size, wherever something arrives | |
| 471 | after the page: the account menu's name, address and invites (now also | |
| 472 | fetched when the pointer reaches the button), Active branches, the | |
| 473 | statement's next entries, the command palette's results, blame's "why". | |
| 474 | ||
| 475 | ## Static assets | |
| 476 | ||
| 477 | Hashed, immutable, a year. Signed in, the first page loads about 120 KB of | |
| 478 | JavaScript gzipped for React and the router, plus the page's own (the pull | |
| 479 | request page about 240 KB in all, most of it the markdown renderer and | |
| 480 | shared components); later pages load only what they add, usually on hover. | |
| 481 | Shiki's grammars load only when code is highlighted, on the server too | |
| 482 | (`lib/highlight.server.ts`), so a Worker starting up no longer evaluates | |
| 483 | them. | |
| 484 | ||
| 485 | ## Budget | |
| 486 | ||
| 487 | | | Target (from the US) | | |
| 488 | | --- | --- | | |
| 489 | | Public page, signed out (cached) | p50 time to first byte under 200 ms | | |
| 490 | | Signed-in page | p50 under 400 ms to first byte | | |
| 491 | | In-app navigation | under 300 ms until the new page shows (data prefetched on hover) | | |
| 492 | | Streamed panels | within 1 s | | |
| 493 | ||
| 494 | The status page's **Page speed** part checks a public project page and | |
| Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders | 495 | Explore every minute, timed to the first byte (the answer's headers), and |
| 496 | shows them as degraded over 800 ms (`apps/status/src/components.ts`, | |
| 497 | `SPEED_BUDGET_MS`). It asks as a browser does: a crawler's user agent makes | |
| 498 | the site render the whole page before the first byte (`isbot` in | |
| 499 | `apps/web/app/entry.server.tsx`), so the check sends a browser's user | |
| 500 | agent ending in `g1t-status/1.0 (+status.g1t.sh)`, which isbot reads as a | |
| 501 | browser (`apps/status/src/probe.ts`, `BROWSER_USER_AGENT`; the sign-in page | |
| 502 | of **Website and sign-in** is loaded the same way). A slow answer is asked | |
| 503 | again at once and counts only if the second is slow too, at the faster of | |
| 504 | the two times; every check is kept for 7 days with the data centre it ran | |
| 505 | from, and sudo's incident page charts them. | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 506 | |
| 507 | ## Measuring | |
| 508 | ||
| 509 | ```powershell | |
| Git storage hardened, pages in tens of milliseconds, honest security alerts, and costs reconciled daily | 510 | # Signed out, as a browser (streamed). Without -BrowserUA curl's own |
| 511 | # user agent counts as a crawler, which waits for the whole page. | |
| 512 | powershell -File scripts/perf/measure.ps1 -BrowserUA -Runs 7 -Out before.csv | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 513 | # Signed in: your g1t_session cookie's value, from DevTools; never printed |
| 514 | $env:G1T_SESSION = "<64 hex>" | |
| 515 | powershell -File scripts/perf/measure.ps1 -Runs 7 -Pull 12 -Issue 11 -Out before-signed-in.csv | |
| 516 | ``` | |
| 517 | ||
| 518 | It prints p50 and p90 of the server's share (TLS handshake done to first | |
| Git storage hardened, pages in tens of milliseconds, honest security alerts, and costs reconciled daily | 519 | byte), where the Worker ran, whether the answer set `g1t_d1` (it should |
| 520 | not, for a page that only reads), and the slowest Server-Timing entries. | |
| Merge project overview: one branch_drift call, spliced histories, cached tags, 6 repos calls instead of 25 | 521 | |
| 522 | Signed out, a public page is usually answered from the data centre's | |
| 523 | cache (`server-timing: cache;desc="hit, …"`), and `cache-control: | |
| 524 | no-cache` does not change that. To time a render, add a query string the | |
| 525 | page ignores: the cache is keyed by the whole URL, so | |
| 526 | `/flagon-io/g1t?nc=<random>` is always a miss. A crawler's user agent | |
| 527 | (`Googlebot/2.1`) waits for the whole page; a browser's gets the first | |
| 528 | byte and the streamed rest. |
This file's history is long; its oldest lines are credited to the oldest commit read.