g1t/docs/PERFORMANCE.md

343 lines19,075 bytesCodeBlame
1# Performance
2
3How g1t.sh answers a page: where the code runs, where the data lives, how
4reads and writes travel, what is cached, and the budget pages are held to.
5Internal. The tools: `scripts/perf/measure.ps1` (time pages from your
6machine), `scripts/perf/placement-probe.mjs` (measure placements without
7touching production), and the Server-Timing header on every page.
8
9## The shape of a request
10
11```
12browser ──► Cloudflare edge (nearest data centre)
13 │
14 ▼
15 g1t (apps/web, React Router) root loader + layout + page loaders, in parallel
16 │ service bindings: POST /rpc/<method>, JSON
17 ▼
18 g1t-identity, g1t-repos, g1t-work, g1t-projects, g1t-billing, … (services/*)
19 │
20 ▼
21 D1 (one SQLite database per service; primary in WNAM, US West)
22 Artifacts (git objects and refs, for g1t-repos)
23```
24
25- **The site holds no data.** Every loader calls services; every service
26 owns one D1 database. A page is a few rounds of service calls; each call
27 is a few D1 queries.
28- **Latency is round trips times distance.** A query from a Worker next to
29 its database takes 1 to 5 ms. The same query from Amsterdam to WNAM took
30 about 150 ms. A page with three rounds of calls, each with two or three
31 queries in turn, costs about a second and a half when the code and the
32 data are on different continents. That was the slowness.
33
34## Where the code runs
35
36| Worker | Placement | Why |
37| --- | --- | --- |
38| `g1t` (site), `g1t-api`, `g1t-sudo` | `off` | Run next to the person. With D1 replicas (below), most reads are local too. |
39| identity, repos, work, search, billing, projects, deployments | `off` | They read with D1 sessions: the nearest replica, or the primary when consistency needs it. |
40| actions, events, webhooks, integrations, context, security | `off` now; a region near WNAM once the probe has picked one | They read the primary only. Pinned beside it, a page from Europe pays one ocean crossing per call instead of one per query. |
41| pages, models, og, runner, status, docs | none | Edge-serving or no data of their own. |
42
43### What happened with Smart Placement
44
45Every data-holding Worker and the site had `"placement": { "mode": "smart" }`.
46Production answered with `cf-placement: remote-AMS` for a visitor in Denver:
47the site ran in Amsterdam, and so did the services it called through
48bindings (a binding runs the callee where the caller is unless the callee
49is placed). Every D1 query then crossed the Atlantic. Smart Placement only
50considers locations where the Worker has already run and needs traffic
51from many places to decide, so it settled on a poor spot and stayed there.
52Measured from Colorado, signed out, warm:
53
54| Page | Smart (AMS) | Off |
55| --- | --- | --- |
56| `/` | 270 ms | 140 ms |
57| `/explore` | 850 ms | 170 ms |
58| `/flagon-io/g1t/pulls` | 600 ms | 220 ms |
59| `/flagon-io/g1t/issues` | 780 ms | 230 ms |
60| `/pricing` | 500 ms | 120 ms |
61| `/flagon-io/g1t` (overview) | 1.4 s | 1.67 s (its own problem, below) |
62
63Placement now applies only to `fetch` handlers; the services are reached
64through `fetch` (`/rpc/<method>`), so it applies to them.
65
66### Choosing a region (the probe)
67
68Cloudflare's placement hints (`"placement": { "region": "aws:us-west-1" }`,
69GCP and Azure regions too, Wrangler 4.146 accepts them) run a Worker next
70to a cloud region. D1 is not a cloud region, and Cloudflare does not say
71which city WNAM is, so measure:
72
73```powershell
74node scripts/perf/placement-probe.mjs deploy # g1t-probe-* on workers.dev, SELECT 1 against g1t-repos
75node scripts/perf/placement-probe.mjs measure # a table: where each ran, ms per D1 query, through a binding too
76node scripts/perf/placement-probe.mjs delete
77```
78
79Pick the region with the lowest **D1 primary ms/query**. To pin the
80primary-only services there, edit their configs (or use
81`node scripts/perf/placement-probe.mjs apply '{"region":"<it>"}'`, which
82sets every config it lists, then put `"mode": "off"` back on the site, API,
83sudo and the session services). `apply '{"mode":"off"}'` is what the
84working tree has now.
85
86## Where the data lives, and how reads travel
87
88Every database's primary is in WNAM. D1 read replication puts read-only
89copies in every region (ENAM, WNAM, WEUR, EEUR, APAC, OC) at no extra
90cost. A copy trails the primary, so reading one needs care.
91
92### Sessions and bookmarks
93
94Seven services read through D1's Sessions API when asked
95(`crates/kit/src/d1.rs`, `packages/contracts/src/d1.ts`). The caller asks
96with the `x-d1-bookmark` request header:
97
98| Header | Reads go to |
99| --- | --- |
100| absent | the primary, with no session: exactly as before |
101| `first-primary` | the primary, then any copy at least as new |
102| `first-unconstrained` | the nearest copy |
103| a bookmark | any copy at least as new as the bookmark |
104
105Writes always go to the primary. A session is sequentially consistent: it
106reads its own writes. The service returns the session's latest bookmark
107in `x-d1-bookmark`.
108
109**Only the site asks for anything.** Service-to-service calls
110(`g1t_kit::call`), queues, crons, the API and MCP send no header, so they
111read the primary as they always did. Billing's `can_start`/`start_run`,
112credential checks for git, and everything agents do stay on the primary.
113
114The site decides per call (`apps/web/app/lib/perf.ts`, `sessionFor`):
115
1161. A request that writes (any method but GET and HEAD) starts every session
117 on the primary, so what an action checks before writing is current.
1182. Within 30 seconds of the person's last write, every service reads its
119 primary. A write can reach a service the site did not call (work
120 writing to repos during a merge) whose bookmark the site never sees;
121 replicas trail by well under a second, so 30 seconds is a wide margin.
1223. Otherwise, the bookmark that service returned after the last write.
1234. Otherwise, the nearest copy.
124
125After a request that may have written, the site sets the `g1t_d1` cookie:
126`at:<unix seconds>` and `service:<bookmark>` pairs, HttpOnly, five minutes.
127"May have written" is: a non-GET request; a GET that called a method not
128on the known-read list (`READS` in `perf.ts`; anything new counts as a
129write until listed); or a GET that started a session (signing in with
130GitHub). GETs that only read set no cookie, so public pages stay cacheable.
131
132Keep `READS` complete. Until 2026-10-06 it lacked `get` (repos and
133projects), `list`, `queue` and `pulls_for_repos`, so every project page
134and Mission control looked like a write: each set the cookie, which kept
135signed-out project pages out of the public cache (every view rendered,
1360.4 to 0.6 s, crawlers included), sent the person's next 30 seconds of
137reads to the primary, and turned off the sidebar cache
138(`mustReadFresh`). `scripts/perf/measure.ps1` shows a **Sets g1t_d1**
139column: it should say False for every page it measures.
140
141What a person can still see out of date: something someone else (or an
142agent, or the API) changed in the last fraction of a second, which a page
143would have missed by loading a moment earlier anyway; and a session
144revoked from another device working for that same fraction of a second
145on reads (sign-out from this browser is a write, so it is immediate).
146
147### Turning replication on
148
149Not on yet: the code above works the same without it (every read is then
150the primary). Turn it on per database once the site and the seven services
151are deployed with sessions. There is no Wrangler command; use the
152dashboard (**D1 → database → Settings → Read replication → Enable**) or
153the API with a token that has D1 Edit:
154
155```powershell
156$token = $env:CLOUDFLARE_D1_TOKEN # D1: Edit on account syntaqx
157$account = "1e6f2cffa3f445920836e8ebe446bb58"
158$databases = @{
159 "g1t" = "b7d49c93-2666-4006-a3c3-073a01838dc9" # identity
160 "g1t-repos" = "f9544c51-c3bf-4621-a96f-8a6d5cf24a97"
161 "g1t-work" = "f35a9022-e36b-4547-b596-9ab9d5f1c47a"
162 "g1t-projects" = "0af698be-2b60-4bd9-aa93-81e84828e991"
163 "g1t-deployments" = "aa935a8f-845b-4132-8fa3-98f1afba1db5"
164 "g1t-search" = "9d04cbf3-cd02-4544-b3e3-2d78767e13bf"
165 "g1t-billing" = "695a6979-fd07-4850-bc97-904a6b4b7a04"
166}
167foreach ($name in $databases.Keys) {
168 curl.exe -s -X PUT "https://api.cloudflare.com/client/v4/accounts/$account/d1/database/$($databases[$name])" `
169 -H "Authorization: Bearer $token" -H "Content-Type: application/json" `
170 --data '{\"read_replication\":{\"mode\":\"auto\"}}'
171 Write-Host ""
172}
173npx wrangler d1 info g1t-repos # read_replication: { mode: "auto" }
174```
175
176Turning it off is `{"read_replication":{"mode":"disabled"}}` and takes up
177to a day to finish. Read-heavy over the last 24 hours: g1t-repos (60,103
178reads to 111 writes), g1t-work (43,219 / 1,462), g1t-projects (28,861 /
17922), g1t-billing (20,600 / 507), g1t (5,434 / 202). g1t-search writes more
180than it reads (indexing), so replicas help it least.
181
182## Caching
183
184| What | Where | For how long | Rules |
185| --- | --- | --- | --- |
186| Static assets (`/assets/*`) | browser and edge | a year, immutable | hashed file names |
187| Avatars | edge cache | a year, immutable | by content hash |
188| Public pages for people signed out | the data centre's cache (`workers/app.ts`, `servePublic`) | fresh 30 s, then served once more while a new copy is made, up to 5 min | GET, no `g1t_session` cookie, an allowlisted path (home, pricing, explore, policies, a project's pages), status 200 or 404, no `Set-Cookie`, nothing private. Reserved first segments and workspace pages (`-`) are never kept. A project's kept page is served only after repos' `visibility` says the repository is still there and public (one indexed read, alongside the cache lookup); a repository made private or deleted is never served from any data centre's copy, and the copy is dropped. The answer says `server-timing: cache;desc="hit, Ns old"`. |
189| Sidebar data (projects, spend, limit, entitlements) | per isolate (`lib/cache.server.ts`) | 15 s, per person and workspace | skipped during a write and for 30 s after the person's last one; failures not kept; only settled answers kept |
190| Registration mode | per isolate | 60 s | |
191| A commit's log by hash | repos' data-centre cache | for good | history from a commit never changes; Active branches asks by hash |
192| Git objects, trees, refs | repos' caches | see services/repos | |
193| A branch's log, the branch list, a file by branch and path | repos' data-centre cache | until the repository's refs change (`refs_version`), 5 min at most | only while no handed-out push credential is live; by commit hash for good (docs/ARTIFACTS.md R9) |
194| A target branch's history, for mergeability | the repos isolate | 60 s, per target head | 100 pull requests checked after a push walk it once (R10) |
195| Git store credentials | repos isolate and KV | reused 50 min (1 h tokens); 3 min for ones handed out | (R3) |
196
197## Server-Timing
198
199Every page and `.data` response from the site carries a `Server-Timing`
200header (DevTools → Network → the request → Timing):
201
202```
203total;dur=180;desc="web to first byte",
204loader.root;dur=40, loader.repo.layout;dur=60, loader.repo.pull;dur=150,
205rpc;dur=140;desc="9 service calls, overlap counted once",
206work;dur=120;desc="4 calls, 70ms inside", repos;dur=30;desc="2 calls, 12ms inside", …,
207d1;desc="work=unconstrained repos=bookmark identity=primary"
208```
209
210- `total`: from the request reaching the site to the response headers.
211 Streamed panels finish after it.
212- `loader.<route>` / `action.<route>`: each route's loader or action
213 (React Router instrumentation in `app/entry.server.tsx`).
214- `rpc`: time waiting on services, overlapping calls counted once.
215- One entry per service: summed wall time of its calls from the site,
216 and how much of that the service itself reported (`svc;dur` from
217 `Served::finish`). The difference is the trip between them.
218- `d1`: how each session-capable service was asked to read.
219- Inside a service entry, `db Nms in T round trips`: what the service
220 reported waiting on its own database (`db;dur`, from
221 `g1t_kit::d1::Timing` and `Served::finish_timed`). The work service
222 reports it, and `rpc;dur` for its own calls to other services; a service
223 call's own response carries both beside `svc;dur`.
224
225Git requests keep their own header (`repos;dur` plus the repos service's
226steps). Counting git operations writes nothing on the way: the meters are
227written after the answer (`wait_until`), so `kept` no longer includes a D1
228upsert (63–98 ms before; docs/ARTIFACTS.md R13). Mission control keeps its per-section timings.
229
230## What a page does, in rounds
231
232Rounds are what cost: calls in the same round overlap.
233
234| Page | Before | Now |
235| --- | --- | --- |
236| Any in-app navigation | root (sidebar: 6 calls, then up to 20 `get_by_id` for shared repositories) and the project layout re-ran when moving between pages of a project | root re-runs only when the workspace or project changes, or after a form; the project layout likewise; shared repositories are one `readable` call, in the same round; the sidebar's workspace data is cached for 15 s; open counts are read once per request for both |
237| Pull request ("Review and respond") | access, then 8 calls, then checks' runs / comparison / session, then up to 5 more deployments lookups for stacked previews | one round of 9 (access-dependent ones start as soon as the repository lookup returns), then the comparison on Changes; the workflow jobs and stacked previews stream in |
238| Project overview | access, then 17 calls, one of which (Active branches) read the default branch's last 120 commits and up to 10 branches' last 40 | one round; Active branches streams in with a skeleton, reading logs by commit hash so a branch that has not moved costs nothing |
239| Mission control | per project: open pulls, closed pulls and events (3 × up to 10), then `get_by_id` per unknown repository | one `pulls_for_repos` call for every project (one access check, one query), events per project alongside, one `readable` for the rest |
240| Issue, issues | access, then the rest | one round |
241
242### Inside the work service
243
244Before 2026-10-06 a pull request's page cost the work service about
245twenty D1 round trips one after another, plus two calls to repos: the
246access check (`get`), then the pull request, its issue, then its
247lifecycle read progress, latest review, settings, statuses, review
248comments, the confidence signals (three in turn), approvals, requests
249for changes, the queue entry, and wrote the stage back on every view,
250then statuses and settings again, messages and earlier checks. It also
251asked repos `behind` on every view, which walks up to `MAX_ANCESTRY`
252commits in Artifacts. About 350 ms of the page's 0.5 s.
253
254Now (`services/work/src/prefetch.rs`):
255
256- **One batch.** Everything the page and the lifecycle read is one D1
257 batch of 20 statements keyed by repository id and number (subqueries
258 find the pull request's id, issue and head). The helpers that decide
259 the lifecycle (`settings`, `statuses`, `review_pending`,
260 `approvals_gap`, `signals`, …) read from it when it is there, so the
261 decision is the same code either way.
262- **Beside the access check.** `repo_then` starts the batch with the
263 repository id this isolate last saw for the path, at the same time as
264 repos' `get`; the rows are used only if `get` then allows that same
265 repository, and read again otherwise. Issues, lists, counts, labels and
266 settings do the same. `pulls_for_repos` reads its rows beside
267 `readable` and drops those of repositories the viewer cannot read.
268- **Precomputed `behind`.** `pulls.behind` is written with mergeability
269 on every push to either side (migration 0023); a view reads it when it
270 was worked out for the current head and asks repos only otherwise
271 (then keeps the answer).
272- **No write on a view** unless the stage changed.
273- `list_active_pulls` reads its pull requests' issues in the same batch
274 instead of one query each.
275
276A pull request is now the repos `get` (about 40 ms) and one batch beside
277it. The indexes were checked with `EXPLAIN QUERY PLAN` against the
278migrations: every statement is an index search; 0023 adds
279`agent_messages_by_sender` for the unanswered-questions count.
280
281### The overview streams
282
283`routes/repo/overview.tsx` returns its seventeen calls as one deferred
284promise. The layout's header and tabs (repository and project, two
285cheap calls) and a skeleton go out first; the sections follow in the
286same response. Crawlers still get the whole page (`entry.server.tsx`
287waits for `allReady` for bots), and signed out it is kept in the public
288cache like any other project page.
289
290## Client navigation
291
292- `<Link prefetch="intent">` on the sidebar, project tabs, breadcrumbs,
293 list rows and Mission control rows: hovering loads the next page's code
294 and data.
295- A thin progress bar while a navigation is pending (`Progress` in
296 `components/shell.tsx`).
297- Mission control's code loads while the browser is idle after the app
298 shell paints; its skeleton is rarely seen.
299- Mission control's quick actions show as done when sent and undo on
300 failure.
301- Skeletons (`components/ui/skeleton.tsx`: `Skeleton`, `SkeletonText`,
302 `SkeletonRows`) stand in, at the same size, wherever something arrives
303 after the page: the account menu's name, address and invites (now also
304 fetched when the pointer reaches the button), Active branches, the
305 statement's next entries, the command palette's results, blame's "why".
306
307## Static assets
308
309Hashed, immutable, a year. Signed in, the first page loads about 120 KB of
310JavaScript gzipped for React and the router, plus the page's own (the pull
311request page about 240 KB in all, most of it the markdown renderer and
312shared components); later pages load only what they add, usually on hover.
313Shiki's grammars load only when code is highlighted, on the server too
314(`lib/highlight.server.ts`), so a Worker starting up no longer evaluates
315them.
316
317## Budget
318
319| | Target (from the US) |
320| --- | --- |
321| Public page, signed out (cached) | p50 time to first byte under 200 ms |
322| Signed-in page | p50 under 400 ms to first byte |
323| In-app navigation | under 300 ms until the new page shows (data prefetched on hover) |
324| Streamed panels | within 1 s |
325
326The status page's **Page speed** part checks a public project page and
327Explore every minute and shows them as degraded over 800 ms
328(`apps/status/src/components.ts`, `SPEED_BUDGET_MS`).
329
330## Measuring
331
332```powershell
333# Signed out, as a browser (streamed). Without -BrowserUA curl's own
334# user agent counts as a crawler, which waits for the whole page.
335powershell -File scripts/perf/measure.ps1 -BrowserUA -Runs 7 -Out before.csv
336# Signed in: your g1t_session cookie's value, from DevTools; never printed
337$env:G1T_SESSION = "<64 hex>"
338powershell -File scripts/perf/measure.ps1 -Runs 7 -Pull 12 -Issue 11 -Out before-signed-in.csv
339```
340
341It prints p50 and p90 of the server's share (TLS handshake done to first
342byte), where the Worker ran, whether the answer set `g1t_d1` (it should
343not, for a page that only reads), and the slowest Server-Timing entries.