Commit

Pull request forks are removed a day after they merge or close, not a week

Their head stays in the repository as refs/pull/<id>/head, so the pull request reads the same; pushing or reopening makes the fork again.

syntaqxcommitted Parent70b3e66Browse files
4 files+6−60/4 viewed
+2−2
142142 in your workspace; pushing creates it.
143143 - **Status.** Not scheduled.
144144
145−### Pull request forks are removed a week after they close
145+### Pull request forks are removed a day after they close
146146
147−Seven days after a pull request merges or closes, its fork's git data is
147+A day after a pull request merges or closes, its fork's git data is
148148 removed. The pull request's changes stay readable: its head is kept in the
149149 repository as `refs/pull/<pull request id>/head`. Pushing to the fork, or
150150 reopening the pull request, makes the fork again from there.
+1−1
3838 listed to every client on every fetch, and each needing cleanup once the
3939 work is merged or closed.
4040
41−Forks add nothing to the repository's branches. A week after a pull request
41+Forks add nothing to the repository's branches. A day after a pull request
4242 merges or closes, its fork is removed, and only its head is kept in the
4343 repository, as `refs/pull/<pull request id>/head`, which clones and fetches
4444 do not download. The repository's branches stay the handful that describe
+2−2
356356 | --- | --- | --- |
357357 | R1 | Built; Cloudflare's answer still needed | Every interaction with the store is metered raw (`meters.rs` → `artifacts_meters`, per day, namespace, repository, workspace and meter, with bytes where known): client git (`git.info_refs`, `git.ls_refs`, `git.fetch`, `git.receive_pack`), g1t's own git (`internal.git.*`: landing, catch-up, mirrors, branch listings, fork retirement), every binding call (`binding.get`, `binding.create_token`, `binding.log`, `binding.read_tree`, …, each retry included), and answers g1t served from its own cache (`cache.*`, never operations). Which meters are operations is data: `operation_mapping` (`cost_operations` for g1t's bill, `billable_operations` for workspaces), read every 5 minutes, changed with `set_operation_mapping` without a deploy. Default: `git.fetch`, `git.receive_pack`, `internal.git.fetch`, `internal.git.receive_pack`, `binding.create`, `binding.fork`, `binding.delete` = 1, everything else 0. `git_operations` (what billing reads) is filled from the meters × `billable_operations`, by the hour. RPCs: `artifacts_usage { from, to, workspace?, by_repo? }` (raw meters and the mapping, for the reconciler), `operation_mapping`, `set_operation_mapping { meter, cost_operations, billable_operations, note? }`. Script: `scripts/ops/artifacts-usage.mjs`. |
358358 | R13 | Built | Nothing on the request path writes D1 for counting. Meters add up per isolate and are written in one batch from `ctx.wait_until` after every request (and at the end of the cron and queue handlers); a failed write is kept for the next. The free-workspace slow-down decides from counts the isolate read back after its last write plus what it added since (`git_ops::standing`, at most 10 minutes old) and billing's plan answer kept 5 minutes. The 63–98 ms `kept` step's D1 upsert is gone. A workspace whose counts this isolate never read is not slowed: nothing slows anyone on a guess. |
359−| R2 | Built; the fork storage test is yours to run | `pull.merged` and `pull.closed` set the fork's `retire_after` (`FORK_RETENTION_DAYS`, default 7); `pull.reopened` clears it, or makes the fork again. The hourly sweep (`23 * * * *`, 25 a run) keeps the fork's head in its repository as `refs/pull/<pull id>/head` (only missing objects travel; an empty pack when merged), records `retired_at` and `retired_head`, then deletes the fork from the store (a failed delete puts the row back). Reads of a retired fork (the pull request's changes, divergence, tree, blob, log, branches) are answered from the repository with the fork's branch mapped to the kept head (`forks.rs` `Viewed`). Anything that writes or uses git on it (git over HTTPS, `git_access`, catch-up, land, `delete_branch`) makes it again first (`revive`: fork, then move its branch to the head) and schedules it to go again. Work never emits `pull.reopened` today; the handler is ready for it. Script: `scripts/ops/fork-storage-test.mjs`. |
359+| R2 | Built; the fork storage test is yours to run | `pull.merged` and `pull.closed` set the fork's `retire_after` (`FORK_RETENTION_DAYS`, 1 day in production; 7 when unset); `pull.reopened` clears it, or makes the fork again. The hourly sweep (`23 * * * *`, 25 a run) keeps the fork's head in its repository as `refs/pull/<pull id>/head` (only missing objects travel; an empty pack when merged), records `retired_at` and `retired_head`, then deletes the fork from the store (a failed delete puts the row back). Reads of a retired fork (the pull request's changes, divergence, tree, blob, log, branches) are answered from the repository with the fork's branch mapped to the kept head (`forks.rs` `Viewed`). Anything that writes or uses git on it (git over HTTPS, `git_access`, catch-up, land, `delete_branch`) makes it again first (`revive`: fork, then move its branch to the head) and schedules it to go again. Work never emits `pull.reopened` today; the handler is ready for it. Script: `scripts/ops/fork-storage-test.mjs`. |
360360 | R3 | Built | Credentials g1t uses itself: TTL 3,600 s, reused for 50 minutes (isolate and KV, key `cred2:<key>:<scope>:internal`). `git_access` hands out its own: TTL 300 s, reused 180 s (`…:handout`); `refs_open` still uses 300 s. Every internal path (land, catch-up, mirrors, branch listing, commits, deleting a branch) now reuses kept credentials instead of minting each time. The remote is worked out as `https://<account>.artifacts.cloudflare.net/git/<namespace>/<name>.git` (the documented format, `api/git-protocol`), learned per namespace from the first `info()` an isolate makes, which runs alongside `createToken` and so costs no time; after that a mint is `get` and `createToken`. Optional `ARTIFACTS_REMOTE_BASE` skips even the first `info()`. |
361361 | R4 | Built | Pushes are read as they arrive (`request.stream()`) and walked by `pack_limits::PackSizer` (each object inflated into a 32 KiB window and thrown away): an object over 32 MB (a delta measured by the object it makes), or a push taking the repository and its forks (`stored_bytes`) past `REPO_STORAGE_LIMIT_BYTES` (950 MB), is declined with `ng` lines and `remote:` text; a repository already at the limit is refused at the push's `info/refs` in plain text. Up to 24 MiB is kept, scanned and sent on as one copy, not three. Past 24 MiB push protection cannot read the push, so it is declined (`LARGE_PUSHES=refuse`, failing closed) with a command to push in parts, the 100 MB network limit named; `LARGE_PUSHES=unscanned` streams it to the store instead, still size-checked (a violation ends the stream before the pack's checksum, so the store keeps nothing). A pack too large for the scanner to inflate (48 MB inflated) is declined the same way instead of let through. Landing streams: upload-pack's side-band answer is taken apart chunk by chunk (`pack_limits::Sideband`) straight into the receive-pack body. |
362362 | R5 | Built | `resilience.rs` sorts errors into rate limited, transient (`INTERNAL_ERROR`, `UPSTREAM_UNAVAILABLE`, `*_IN_PROGRESS`, no code, HTTP 5xx) and permanent. Binding reads, `get`, `info`, `createToken`, `create` and `delete` try up to 3 times with exponential backoff and jitter (80 ms base, 400 ms for rate limits, 2 s cap); `fork` and every receive-pack never retry. Git reads (`info/refs`, upload-pack) retry on 429 and 5xx. Per isolate, each namespace has a breaker that opens after 5 transient failures in a row, for 10 s, then lets one probe through. Busy answers reach git as 429 (rate limited) or 503, with `Retry-After: 5`; the site's read RPCs (`tree`, `blob`, `log`, `branches`, `blame`, `compare`) answer an `Outcome` failure saying the git storage is busy; other RPCs answer 503 with the same words. Health is counted by the minute (`store_health`) and served by the `store_health { minutes }` RPC; status.g1t.sh lists **Git storage** through a new `REPOS` service binding (down: 25% or more of at least 5 calls failed, or the breaker refused calls; degraded: rate limited, or a mean call over 1.5 s). |
453453 the six, not about 106 MB. Whatever Artifacts shares underneath, storage billing and the 1 TB
454454 account limit see full copies. So:
455455
456−- `FORK_RETENTION_DAYS` should drop from 7 to 1 (not yet changed).
456+- `FORK_RETENTION_DAYS` is 1 in production (changed 2026-10-07).
457457 - g1t meters a workspace's storage once per repository (`stored_bytes`), so an open pull
458458 request's working copy is Cloudflare cost g1t absorbs: about $0.05 a month per 100 MB per open
459459 pull request. Small now; decide whether open working copies count toward a workspace's
+1−1
6666 "FREE_PRIVATE_STORAGE_BYTES": "1000000000",
6767 "REPO_STORAGE_LIMIT_BYTES": "950000000",
6868 "LARGE_PUSHES": "unscanned",
69− "FORK_RETENTION_DAYS": "7",
69+ "FORK_RETENTION_DAYS": "1",
7070 "ARTIFACTS_NAMESPACES": "{\"ARTIFACTS\":\"g1t\"}",
7171 "ARTIFACTS_NEW_REPOS": ""
7272 },