Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 1 | # Building a git platform on Cloudflare: a field report |
| 2 | ||
| 3 | Status: feedback document, 2026-10-06. For Cloudflare's Artifacts, Workers, D1 and Containers teams, | |
| 4 | and for the organisers of "Build the Next-Gen Git Platform on Cloudflare". | |
| 5 | ||
| 6 | g1t (g1t.sh) is a git hosting platform where people and coding agents work in the same issues, pull | |
| 7 | requests and merge queue. It runs entirely on Cloudflare: about twenty Workers (Rust compiled to | |
| 8 | WebAssembly, plus TypeScript), one D1 database per service, Artifacts for every repository and pull | |
| 9 | request fork, Containers for agent sandboxes and CI, KV, R2, Queues and Cloudflare for SaaS for custom | |
| 10 | domains. The source lives on g1t itself at [g1t.sh/flagon-io/g1t](https://g1t.sh/flagon-io/g1t); code | |
| 11 | pointers below are paths in that repository. | |
| 12 | ||
| 13 | We went from an empty repository to an invite-only launch on this stack, and most of it worked the | |
| 14 | first time. This report is about where it did not, written so that each point can become a ticket: | |
| 15 | what we tried to do, what we hit (with numbers, dates and doc quotes), what we built around it and what | |
| 16 | that costs us, and what we would ask for. The full Artifacts due diligence, with every citation and the | |
| 17 | capacity model, is in [ARTIFACTS.md](ARTIFACTS.md). Latency work is in [PERFORMANCE.md](PERFORMANCE.md); | |
| 18 | deploys in [DEPLOYING.md](DEPLOYING.md). | |
| 19 | ||
| 20 | Severity: | |
| 21 | ||
| 22 | - **Blocking**: we cannot price, launch or scale a feature safely until it is answered or changed. | |
| 23 | - **Costly**: we ship, but only with a workaround that costs us engineering time, latency or money on | |
| 24 | every request. | |
| 25 | - **Friction**: a sharp edge that cost hours once, or a doc that disagrees with behaviour. | |
| 26 | ||
| 27 | All quotes are from developers.cloudflare.com, retrieved 2026-10-06, unless stated. | |
| 28 | ||
| 29 | ## Executive summary: the five asks | |
| 30 | ||
| 31 | 1. **Define a billable Artifacts operation, and meter exactly that.** Pricing says "repo operations, | |
| 32 | such as `create`, `push`, `pull`, and `clone`"; metrics count `create`, `fork`, `push`, `pull`, | |
| 33 | `delete`. Depending on whether binding reads, token mints and ref advertisements count, our cost at | |
| 34 | 3,000 workspaces is about $1,800 or about $31,000 a month. Billing starts 2026-10-14. (A1) | |
| 35 | 2. **Document what a fork stores, and give forks a lifecycle.** Forks are the obvious per-pull-request | |
| 36 | primitive, but nothing says whether they copy or share objects. If they copy, an agent-heavy | |
| 37 | workload reaches the 1 TB account limit in about two days and every push in the account fails. Ask: | |
| 38 | documented semantics, per-repository storage in the binding, fork expiry, and a warning before the | |
| 39 | account cliff. (A2) | |
| 40 | 3. **A write path and a policy hook in the binding.** List refs, compare-and-swap ref updates | |
| 41 | (atomic across refs), object and pack writes, and a pre-receive hook a Worker can answer. Today we | |
| 42 | speak smart HTTP to our own storage, buffer packs in 128 MB isolates, and parse every push in the | |
| 43 | Worker before forwarding it. (A3, A5) | |
| 44 | 4. **Account-level ref-change events and cheaper git credentials.** One subscription per repository | |
| 45 | does not scale to tens of thousands of repositories, so we version refs ourselves and guard it with a | |
| 46 | test that scans our source. Every git access costs three binding calls and about 0.8 s to mint a | |
| 47 | token. (A4, A6) | |
| 48 | 5. **Placement that follows the data.** Smart Placement ran our site in Amsterdam for a visitor in | |
| 49 | Denver while every D1 primary is in WNAM. Explore took 0.85 s that way and 0.17 s with placement off. Ask: placement that | |
| 50 | understands D1 primaries and service-binding chains, D1 regions as placement hints, and D1 Sessions | |
| 51 | that work across service bindings. (W1, W2, D1) | |
| 52 | ||
| 53 | ## Artifacts | |
| 54 | ||
| 55 | Artifacts is the reason g1t exists in this form. One Durable Object per repository, smart HTTP that | |
| 56 | stock git understands, forks, scoped tokens and a binding that reads commits, trees and blobs is the | |
| 57 | right shape for a git platform. Everything below is about taking it from "it works" to "we can bill | |
| 58 | for it and sleep at night at a few thousand workspaces". | |
| 59 | ||
| 60 | ### A1. What is a billable operation? (blocking) | |
| 61 | ||
| 62 | - **What we tried.** Price g1t from cost: pass Artifacts through to each workspace with a modest | |
| 63 | uniform overhead, so a workspace's bill tracks what it really costs us. | |
| 64 | - **What we hit.** [Pricing](https://developers.cloudflare.com/artifacts/platform/pricing/) defines | |
| 65 | operations only as "the number of repo operations, such as `create`, `push`, `pull`, and `clone`", | |
| 66 | $0.15 per 1,000 after 10,000 a month. [Metrics](https://developers.cloudflare.com/artifacts/observability/metrics/) | |
| 67 | (`artifactsEventsAdaptiveGroups`) list `eventType` values `create`, `fork`, `push`, `pull`, `delete`: | |
| 68 | a different list, with no `clone`. Neither says whether binding calls (`get`, `info`, `createToken`, | |
| 69 | `log`, `readTree`, `readBlob`, `readCommit`, `readFile`), `info/refs`, or protocol v2 `ls-refs` count. | |
| 70 | - **What it means.** Our capacity model (ARTIFACTS.md section 5) at 3,000 workspaces gives about | |
| 71 | 12 million operations a month if only git data transfers count (about $1,800), and about 207 million | |
| 72 | if every call counts (about $31,000). A 15x spread, with billing starting 2026-10-14. | |
| 73 | - **What we built.** `services/repos/src/git_ops.rs` counts every upload-pack and receive-pack POST per | |
| 74 | workspace, including `ls-refs` answers served from our own cache that never reach Artifacts. It is a | |
| 75 | guess. We also had to size the free tier (`GIT_OPERATIONS_FREE_CAP`, 50,000 a month) against a | |
| 76 | number we cannot see. | |
| 77 | - **Ask.** A table: each binding method and each git endpoint, billable or not, and how many operations | |
| 78 | it is. Metrics with the same event names as the invoice, per repository, so a platform can attribute | |
| 79 | cost to its own customers. A usage endpoint for the current billing period. | |
| 80 | ||
| 81 | ### A2. Fork storage, fork lifecycle, and the 1 TB cliff (blocking) | |
| 82 | ||
| 83 | - **What we tried.** One fork per pull request (`pulls--<id>` in `services/repos/src/store.rs`), so an | |
| 84 | agent's work is isolated until it lands. Forks are the most natural primitive Artifacts offers for | |
| 85 | this, and agents open pull requests by the thousand. | |
| 86 | - **What we hit.** The docs do not say whether `fork` copies objects or shares them with its source. | |
| 87 | The fork response returns `objects: N`, which reads like a copy count. Pricing says "Repos remain | |
| 88 | stored until you explicitly delete them", and [limits](https://developers.cloudflare.com/artifacts/platform/limits/) | |
| 89 | put account storage at 1 TB ("can be raised on request"). There is no repository size in the binding, | |
| 90 | no fork expiry and no guidance on fork lifecycle. | |
| 91 | - **What it means.** At 15,000 agent pull requests a day on 25 MB repositories: if forks share objects, | |
| 92 | storage grows by about 270 GB a month; if they copy, by about 13.5 TB a month, and the account limit | |
| 93 | is reached in about two days. At that point every push in the account fails with | |
| 94 | `storageLimitReached`, for every customer at once. | |
| 95 | - **What we built.** Nothing yet that we trust. We must add a sweep that deletes forks after a pull | |
| 96 | request merges or closes plus a grace period (`lifecycle.rs`), and measure semantics ourselves with a | |
| 97 | 100 MB test repository and the storage metric. | |
| 98 | - **Ask.** Document fork storage (copy, copy-on-write, or shared packs, and how it is billed). Report a | |
| 99 | repository's stored bytes in `info()`. Allow an expiry on `fork()` (or on any repository). Warn by | |
| 100 | event or email at 80% and 95% of the account limit, and fail per repository rather than account-wide. | |
| 101 | ||
| 102 | ### A3. The binding cannot write (costly) | |
| 103 | ||
| 104 | - **What we tried.** Land pull requests, catch branches up, mirror external repositories, commit a | |
| 105 | file from the web, delete and rename branches: all server-side operations a platform does without a | |
| 106 | git client. | |
| 107 | - **What we hit.** The [Workers binding](https://developers.cloudflare.com/artifacts/api/workers-binding/) | |
| 108 | reads (`log`, `readCommit`, `readTree`, `readBlob`, `readFile`) but cannot list refs, update refs, | |
| 109 | write objects or packs, or update several refs atomically. The | |
| 110 | [isomorphic-git example](https://developers.cloudflare.com/artifacts/examples/isomorphic-git/) still | |
| 111 | says the binding "cannot read or write files inside them", which is half out of date. | |
| 112 | - **What we built.** A smart HTTP client inside the Worker. `services/repos/src/land.rs` asks the fork | |
| 113 | for a pack (`fetch_pack`), strips the side-band (`unpack_sideband`), and pushes it to the target | |
| 114 | (`push_pack`, `fast_forward`, `delete_ref`), each a hand-built receive-pack command with pkt-lines. | |
| 115 | Branch listing parses `info/refs` (`refs.rs`). Mirrors and imports do the same (`mirror.rs`, | |
| 116 | `import.rs`), capped at 40 MB a pack (`MAX_PACK_BYTES`) because everything is buffered. | |
| 117 | - **What it costs.** Every landing is two token mints (six binding calls), an upload-pack and a | |
| 118 | receive-pack. Each pack is held in memory two or three times in an isolate with 128 MB shared by | |
| 119 | concurrent requests, so a large pull request can fail to land, and can take other requests in the | |
| 120 | same isolate with it. Ref updates are one ref per request; there is no documented `atomic` support. | |
| 121 | - **Ask.** `repo.refs()` (with peeled tags), `repo.updateRefs([{ name, old, new }], { atomic: true })`, | |
| 122 | `repo.copyObjects(from, wants, haves)` or `repo.writePack(stream)` that stream rather than buffer, | |
| 123 | and confirmation that receive-pack `atomic` is supported. | |
| 124 | ||
| 125 | ### A4. Three calls and 0.8 s for every credential (costly) | |
| 126 | ||
| 127 | - **What we tried.** Proxy git over HTTPS: authenticate the person or agent ourselves, then forward to | |
| 128 | the repository's Artifacts remote with a short-lived token. | |
| 129 | - **What we hit.** Each credential needs `namespace.get()` (a lookup that throws `NOT_FOUND` or | |
| 130 | `*_IN_PROGRESS`), then `info()` for the remote URL ("Each call performs a fresh lookup") and | |
| 131 | `createToken`. Measured from Denver on 2026-10-06: 807 to 856 ms per mint, most of a cold `info/refs` | |
| 132 | that took 1.3 to 1.4 s end to end. | |
| 133 | - **What we built.** Two credential caches in `services/repos/src/store.rs`: per isolate (`Credentials`) | |
| 134 | and shared across isolates through KV, sealed with AES-256-GCM (`shared.rs`). Tokens live 300 s and | |
| 135 | are reused for 180 s. That bounds staleness and leak exposure but still means about 20 mints an hour | |
| 136 | per busy repository and scope, and a KV round trip on many requests. | |
| 137 | - **Ask.** Return `remote` from `get()` (it is fixed for a repository) or document its format so | |
| 138 | `info()` is not needed. A way for a Worker with the binding to forward git requests without a bearer | |
| 139 | token at all, for example `repo.fetch(request)`, authenticated by the binding. Failing that, document | |
| 140 | whether `get`, `info` and `createToken` are billed (A1). | |
| 141 | ||
| 142 | ### A5. No server-side hooks (costly) | |
| 143 | ||
| 144 | - **What we tried.** Branch protection, push protection (secret scanning) and blocking pushes that | |
| 145 | would publish a person's private email, as every git host does, before refs move. | |
| 146 | - **What we hit.** Artifacts has no pre-receive or update hook, and tokens are bearer credentials. | |
| 147 | Anything holding a write token can push past every policy. | |
| 148 | - **What we built.** `services/repos/src/git_http.rs` `forward` reads the whole push body | |
| 149 | (`request.bytes()`), checks protected branches (`refusal`), then `secret_scan.rs` `scan_push` parses | |
| 150 | the pack in WebAssembly, resolves delta bases by reading objects back through the binding | |
| 151 | (`supply_bases`, up to 500 bases), walks trees and scans up to 24 MB (`MAX_SCANNED_PUSH`) before the | |
| 152 | body is copied again and forwarded. Pushes larger than that are let through unscanned, and we say so in | |
| 153 | the logs. Write tokens handed to sandboxes (`run_access.rs`) bypass all of it, so we keep their life | |
| 154 | minimal. | |
| 155 | - **What it costs.** A second implementation of git's pack format in our code, memory pressure on every | |
| 156 | push, and a policy that only holds for pushes that come through our proxy. | |
| 157 | - **Ask.** A pre-receive hook: Artifacts calls a Worker (service binding) with the ref updates and a | |
| 158 | streaming view of the new objects, and the Worker answers accept, or reject with a message git shows | |
| 159 | the user. Even a hook that sees only ref updates, with object reads through the binding, would remove | |
| 160 | most of this code. | |
| 161 | ||
| 162 | ### A6. No account-level push or ref-change events (costly) | |
| 163 | ||
| 164 | - **What we tried.** React to pushes: run checks and Actions, recompute mergeability of open pull | |
| 165 | requests, reindex search, and invalidate caches of the ref advertisement. | |
| 166 | - **What we hit.** [Event subscriptions](https://developers.cloudflare.com/artifacts/guides/event-subscriptions/) | |
| 167 | give account-level `repo.created|deleted|forked|imported`, but `pushed`, `cloned`, `fetched` and token | |
| 168 | events need one subscription per repository. At tens of thousands of repositories and forks that is not | |
| 169 | practical to manage. Read-after-write for refs is not documented either. | |
| 170 | - **What we built.** We record pushes ourselves (`record_push`), reading `log(branch, 1)` right after | |
| 171 | the receive-pack response and assuming it is final. Every code path that moves refs bumps a | |
| 172 | `refs_version` in D1, and `services/repos/src/refs_cache.rs` keys cached ref advertisements by it. | |
| 173 | Because a missed bump would serve stale refs, a unit test (`every_ref_writer_records_the_change`) | |
| 174 | scans our own source files for every ref writer and fails the build if one does not call | |
| 175 | `refs_moved`. | |
| 176 | - **What it costs.** Correctness depends on a convention enforced by text search. Pushes made directly | |
| 177 | to Artifacts with a sandbox token are invisible to it until the 60 s cache TTL runs out. | |
| 178 | - **Ask.** An account- or namespace-level `refs.updated` event with repository, ref, old and new ids, | |
| 179 | delivered to a Queue. A documented guarantee that refs are visible to `log` and `info/refs` everywhere | |
| 180 | once receive-pack responds. A cheap `refsVersion` or ref-advertisement ETag on `info()`. | |
| 181 | ||
| 182 | ### A7. Every full clone rebuilds its pack; partial clone docs disagree with behaviour (costly, friction) | |
| 183 | ||
| 184 | - **What we tried.** Agent sandboxes, checks and the merge queue each clone the repository, often the | |
| 185 | same commit several times per pull request. | |
| 186 | - **What we hit.** Measured 2026-10-06 from Denver: a full upload-pack of `flagon-io/g1t` (6,187 | |
| 187 | objects, 5.5 MB) takes 1.96 to 2.22 s to the first byte, every time; the pack is built before anything | |
| 188 | is sent and nothing is cached. Separately, the [git protocol page](https://developers.cloudflare.com/artifacts/api/git-protocol/) | |
| 189 | marks `filter` "not supported" for v1, yet a protocol v2 `git clone --filter=blob:none` returned a real | |
| 190 | partial clone (2,789 objects, 903 KB). We cannot tell whether to build on it. | |
| 191 | - **What we built.** Ref advertisement caching (above). A pack cache in R2 keyed by repository, | |
| 192 | `refs_version` and request hash is planned (ARTIFACTS.md R6), as is moving checks to `--depth=1`. | |
| 193 | - **Ask.** Cache or reuse packs for identical wants with no haves (or bitmaps, so pack building is | |
| 194 | cheap). State partial clone support for protocol v2 explicitly, including which filters. | |
| 195 | ||
| 196 | ### A8. Namespace rate limit and sharding left to the customer (costly) | |
| 197 | ||
| 198 | - **What we tried.** One namespace (`g1t`) for every repository and fork, one `ARTIFACTS` binding. | |
| 199 | - **What we hit.** Control-plane requests are limited to "2,000 requests per 10 seconds per namespace". | |
| 200 | It is not documented whether binding calls count. Our model peaks at 250 to 400 calls a second against | |
| 201 | a ceiling of 200. [Best practices](https://developers.cloudflare.com/artifacts/concepts/best-practices/) | |
| 202 | say "Do not keep every repo in one default namespace once usage grows", but a binding names exactly one | |
| 203 | namespace, and [data location](https://developers.cloudflare.com/artifacts/guides/data-localization/) | |
| 204 | (`eu` or `us`) is fixed per namespace at creation. | |
| 205 | - **What we have to build.** Bindings `ARTIFACTS_0..N` plus `ARTIFACTS_EU`, a shard column per | |
| 206 | repository in our registry, least-loaded placement of new repositories, and forks in their own | |
| 207 | namespaces (ARTIFACTS.md R7). Every platform on Artifacts will build this same router. | |
| 208 | - **Ask.** Say whether binding calls count toward the namespace limit, and what a client sees past it | |
| 209 | (429, `rateLimited`, queueing, `Retry-After`). One binding that can address many namespaces | |
| 210 | (`env.ARTIFACTS.namespace("g1t-eu-3")`). Either automatic sharding behind one namespace or a | |
| 211 | per-namespace limit that can be raised on request like the account storage limit. | |
| 212 | ||
| 213 | ### A9. Limits whose failure modes are undocumented (friction) | |
| 214 | ||
| 215 | - **What we hit.** 1 GB per repository and 32 MB per file, but nothing on what a git client sees when a | |
| 216 | push crosses them, whether storage is garbage-collected or repacked, and when deleted branches or | |
| 217 | force pushes free space. `MEMORY_LIMIT` (10402) is raised "if the object cannot be buffered safely", | |
| 218 | with no stated size. No repository size in the binding. | |
| 219 | - **What we have to build.** Our own checks in front of Artifacts that refuse objects over 32 MB and | |
| 220 | pushes that would cross about 950 MB, with an `ng` line git can show (ARTIFACTS.md R4), without knowing | |
| 221 | the true current size of a repository. | |
| 222 | - **Ask.** Document the client-visible behaviour at each limit, GC and repack policy, and expose stored | |
| 223 | bytes per repository. | |
| 224 | ||
| 225 | ### A10. Durability, export and support during the beta (costly) | |
| 226 | ||
| 227 | - **What we hit.** [How Artifacts works](https://developers.cloudflare.com/artifacts/concepts/how-artifacts-works/): | |
| 228 | "Cloudflare replicates repo data synchronously across multiple data centers and copies it | |
| 229 | asynchronously to object storage and snapshots." Good. But there is no SLA, no support path for the | |
| 230 | beta, no export or snapshot restore, and no statement of how long snapshots are kept. | |
| 231 | - **What we have to build.** Nightly `git bundle` backups of every changed repository to R2, made in a | |
| 232 | Container because Workers cannot run git, and a restore drill (ARTIFACTS.md R11). | |
| 233 | - **Ask.** Snapshot restore (point in time, per repository) and export to an R2 bucket we own, a beta | |
| 234 | support channel, and a GA date with an SLA. | |
| 235 | ||
| 236 | ### A11. No region or placement hint for a repository (friction) | |
| 237 | ||
| 238 | - **What we hit.** A repository is "a single logical instance that Cloudflare can route to from any | |
| 239 | region". We cannot see or influence where its Durable Object lives, so we cannot put a busy team's | |
| 240 | repository next to them, or our `g1t-repos` Worker next to the repository. | |
| 241 | - **Ask.** A `locationHint` on `create`/`fork`, as Durable Objects have, and the chosen location in | |
| 242 | `info()`. | |
| 243 | ||
| 244 | ## Workers and placement | |
| 245 | ||
| 246 | ### W1. Smart Placement chose Amsterdam for a Denver visitor (costly) | |
| 247 | ||
| 248 | - **What we tried.** `"placement": { "mode": "smart" }` on the site and every data-holding service, as | |
| 249 | recommended for Workers that talk to a backend. | |
| 250 | - **What we hit.** On 2026-10-06 production answered with `cf-placement: remote-AMS` for a visitor in | |
| 251 | Denver, and the git Worker had earlier run as `remote-ATL`, while every D1 primary is in WNAM. Services | |
| 252 | reached through service bindings run where their caller runs, so every D1 query crossed the Atlantic | |
| 253 | (about 85 to 150 ms each), and loaders that chain a few rounds of calls took 1.4 s or more. Pages | |
| 254 | measured from Colorado, signed out, warm (PERFORMANCE.md): | |
| 255 | ||
| 256 | | Page | Smart (AMS) | Placement off | | |
| 257 | | --- | ---: | ---: | | |
| 258 | | `/` | 270 ms | 140 ms | | |
| 259 | | `/explore` | 850 ms | 170 ms | | |
| 260 | | `/flagon-io/g1t/pulls` | 600 ms | 220 ms | | |
| 261 | | `/flagon-io/g1t/issues` | 780 ms | 230 ms | | |
| 262 | | `/pricing` | 500 ms | 120 ms | | |
| 263 | ||
| 264 | The [placement docs](https://developers.cloudflare.com/workers/configuration/placement/) say Smart | |
| 265 | Placement weighs "the Worker's performance and the network latency added by forwarding the request", | |
| 266 | and that it "only affects the execution of fetch event handlers. It does not affect RPC methods or | |
| 267 | named entrypoints." Neither the D1 primary's location nor the chain of service bindings behind the | |
| 268 | Worker seems to enter the decision, and the decision stuck. | |
| 269 | - **What we built.** Placement off everywhere, D1 read replicas through the Sessions API (D1 below), a | |
| 270 | probe script (`scripts/perf/placement-probe.mjs`) that deploys throwaway Workers to measure D1 latency | |
| 271 | per placement, and a `Server-Timing` header on every page so we can see the next regression. | |
| 272 | - **Ask.** Let Smart Placement account for D1 primaries and the bindings a Worker calls, re-evaluate | |
| 273 | placement when it is clearly worse than local, and keep `cf-placement` (the docs say it "may be removed | |
| 274 | before Smart Placement exits beta"; please keep it, it is how we found this). | |
| 275 | ||
| 276 | ### W2. Placement hints name cloud regions, not Cloudflare data (friction) | |
| 277 | ||
| 278 | - **What we hit.** Hints like `"region": "aws:us-west-1"` place a Worker next to a cloud region. D1 is | |
| 279 | not one, and Cloudflare does not say which city WNAM is, so pinning a primary-only service beside its | |
| 280 | database means measuring candidates by hand. | |
| 281 | - **Ask.** Accept a D1 database (or a Durable Object, or an Artifacts namespace) as a placement target, | |
| 282 | for example `"placement": { "near": { "d1": "g1t-repos" } }`. | |
| 283 | ||
| 284 | ### W3. Git over SSH needs inbound TCP (friction) | |
| 285 | ||
| 286 | - **What we hit.** Git users expect `git@host:owner/repo.git`. Workers have no inbound TCP, so g1t | |
| 287 | serves HTTPS only. We submitted the inbound TCP request form on 2026-10-05 and are waiting. | |
| 288 | - **Ask.** A path to inbound TCP on a Worker or Container for SSH, even rate-limited, and a published | |
| 289 | timeline. | |
| 290 | ||
| 291 | ## D1 | |
| 292 | ||
| 293 | ### D1-1. Read replication across service bindings needs hand-built plumbing (costly) | |
| 294 | ||
| 295 | - **What we tried.** One D1 database per service, read replicas near people, read-your-writes after a | |
| 296 | form submits. | |
| 297 | - **What we hit.** D1 has one primary region. Replication is free and good, but the Sessions API's | |
| 298 | bookmarks live in one Worker's code. Our site calls seven services through service bindings, and each | |
| 299 | service owns its database, so a bookmark must travel from the database, through the service, back to | |
| 300 | the site, into a cookie, and back on the next request. | |
| 301 | - **What we built.** An `x-d1-bookmark` header protocol (`crates/kit/src/d1.rs`, | |
| 302 | `packages/contracts/src/d1.ts`), per-service bookmarks in an HttpOnly cookie, a 30 s "read the primary | |
| 303 | after a write" window to cover services the site never sees write, and a list of known-read RPC methods | |
| 304 | so new ones are treated as writes by default (`apps/web/app/lib/perf.ts`). | |
| 305 | - **Ask.** First-class session propagation across service bindings (a bookmark carried with the call, | |
| 306 | like trace context), and a `wrangler d1` command to turn replication on and off. Today it is dashboard | |
| 307 | or REST only, and turning it off "takes up to a day". | |
| 308 | ||
| 309 | ### D1-2. Transient and inconsistent errors from Wrangler (friction) | |
| 310 | ||
| 311 | - **What we hit.** During launch week, `wrangler d1 migrations apply` intermittently failed with 7403 | |
| 312 | ("account not valid or not authorized") with the same credentials that worked minutes later. On | |
| 313 | 2026-10-05 `wrangler d1 execute --file` failed with authentication error 10000 while | |
| 314 | `wrangler d1 execute --command` with the same SQL and the same login worked; we applied a migration | |
| 315 | with `--command` to get unblocked. | |
| 316 | - **Ask.** Make `--file` and `--command` use the same authorization, and make 7403 say which permission | |
| 317 | or account it was checking. | |
| 318 | ||
| 319 | ## Containers | |
| 320 | ||
| 321 | ### C1. No Docker inside a sandbox (friction) | |
| 322 | ||
| 323 | - **What we tried.** Run g1t's own CI on g1t: the workflow that rebuilds our sandbox base image weekly | |
| 324 | runs in a Containers sandbox. | |
| 325 | - **What we hit.** We found no supported way to run Docker or BuildKit inside a Container, so a sandbox | |
| 326 | cannot build or push an image. `.g1t/workflows/runner-base.yml` waits for a self-hosted runner with a | |
| 327 | `docker` label, and image builds happen on a laptop (DEPLOYING.md, "The runner's images"). | |
| 328 | - **Ask.** Rootless BuildKit in a Container, or a managed image build service that takes a Dockerfile | |
| 329 | and a context and pushes to `registry.cloudflare.com`. | |
| 330 | ||
| 331 | ### C2. Ephemeral disk rules Containers out as a durable store (costly) | |
| 332 | ||
| 333 | - **What we tried.** Make our self-hostable git store (`deploy/self-host/gitstore`) a warm fallback for | |
| 334 | Artifacts, running on Cloudflare. | |
| 335 | - **What we hit.** Container disk is ephemeral and at most 20 GB | |
| 336 | ([limits](https://developers.cloudflare.com/containers/platform-details/limits/)). A git store needs | |
| 337 | persistent disk, so the fallback has to live off Cloudflare. | |
| 338 | - **Ask.** Persistent volumes for Containers (or R2-backed volumes with local caching). | |
| 339 | ||
| 340 | ### C3. Image rebuilds without a remote cache (friction) | |
| 341 | ||
| 342 | - **What we hit.** After a local prune of Docker's build cache, rebuilding the sandbox base image (Debian, Node, | |
| 343 | Python, Go, Rust, toolchains) took about 20 minutes, because nothing on the registry side served as a | |
| 344 | build cache. | |
| 345 | - **What we built.** `BUILDKIT_INLINE_CACHE` and `--cache-from` the previous base, and a two-layer image | |
| 346 | so a runner change pushes about 5 MB (`scripts/deploy/image.mjs`). | |
| 347 | - **Ask.** Document `--cache-to type=registry` support on `registry.cloudflare.com`, or a hosted build | |
| 348 | cache. | |
| The runner's base image, pushed as Wrangler builds images, with pushes that fail loudly | 349 | - **Also hit (2026-10-06).** The first push of the 2.7 GB base to `registry.cloudflare.com` failed after |
| 350 | several layers with `error from registry: blob unknown to registry`, from `docker push` and from | |
| 351 | `wrangler containers push` alike. The same push, run again an hour later, found every layer there and | |
| 352 | finished in 17 s. The image was an OCI index carrying a BuildKit provenance attestation (an | |
| 353 | `unknown/unknown` manifest), the default for `docker buildx build --load`; Wrangler's own builds pass | |
| 354 | `--provenance=false`. We now build as Wrangler does (one manifest, no provenance or SBOM) and push up | |
| 355 | to three times, failing loudly when no digest comes back. Ask: say whether manifests may reference | |
| 356 | just-uploaded blobs at once, and whether indexes with attestation manifests are supported. | |
| Fast pages, required checks on the branch, self-hosted runners, honest incidents | 357 | |
| 358 | ## Tooling and account | |
| 359 | ||
| 360 | ### T1. Wrangler credentials are easy to get wrong silently (friction) | |
| 361 | ||
| 362 | - **What we hit.** The global API key needs the account email as well as the key. Wrangler loads the | |
| 363 | repository's `.env` itself, so a `CLOUDFLARE_API_TOKEN` there (with fewer permissions) overrides a | |
| 364 | working `wrangler login`, and `unset` in the shell does not help; only an explicitly empty variable | |
| 365 | does. A missing token permission fails with `Authentication error [code: 10000]` and no permission | |
| 366 | name, and there is no published list of what `wrangler deploy` needs for each binding type (we wrote | |
| 367 | our own, DEPLOYING.md "The API token"). | |
| 368 | - **Ask.** Print which credential source Wrangler chose. Name the missing permission in 10000 errors. | |
| 369 | Publish the permissions each binding needs at deploy time. | |
| 370 | ||
| 371 | ### T2. Two accounts and non-interactive mode (friction) | |
| 372 | ||
| 373 | - **What we hit.** A login that sees two accounts makes `wrangler d1` and `wrangler tail` refuse to pick | |
| 374 | one in non-interactive mode (CI, agents), even when `account_id` is in `wrangler.jsonc`. Every command | |
| 375 | needs `CLOUDFLARE_ACCOUNT_ID` set. | |
| 376 | - **Ask.** Honour `account_id` from the config for every command, as `deploy` does. | |
| 377 | ||
| 378 | ### T3. Cloudflare for SaaS route refused right after enabling (friction) | |
| 379 | ||
| 380 | - **What we hit.** On 2026-10-05, right after enabling Cloudflare for SaaS on `g1t.page`, adding the | |
| 381 | `*/*` Worker route failed with 10022 and 100327. The same request succeeded a few minutes later. | |
| 382 | - **Ask.** Return a "still provisioning, retry in N seconds" error instead of a refusal, or block until | |
| 383 | ready. | |
| 384 | ||
| 385 | ## What we'd love to see | |
| 386 | ||
| 387 | In rough order of how much code it would delete for us: | |
| 388 | ||
| 389 | 1. A published definition of a billable Artifacts operation, with matching per-repository metrics. | |
| 390 | 2. Fork storage semantics, fork expiry, repository size in `info()`, and an account-limit warning. | |
| 391 | 3. Ref listing, atomic compare-and-swap ref updates and streaming object or pack writes in the binding. | |
| 392 | 4. A pre-receive hook answered by a Worker. | |
| 393 | 5. Account-level `refs.updated` events to a Queue, and a read-after-write guarantee for refs. | |
| 394 | 6. Binding-authenticated git forwarding (`repo.fetch(request)`) with no token mint. | |
| 395 | 7. Pack reuse for identical full clones, and documented protocol v2 partial clone. | |
| 396 | 8. One binding over many namespaces, clear namespace rate-limit behaviour, raisable limits. | |
| 397 | 9. Snapshot restore and export to our own R2 bucket; a beta support path; a GA date and SLA. | |
| 398 | 10. Location hints for repositories. | |
| 399 | 11. Smart Placement aware of D1 and service bindings; D1 as a placement target; keep `cf-placement`. | |
| 400 | 12. D1 sessions that cross service bindings; replication in Wrangler. | |
| 401 | 13. Persistent volumes and image builds for Containers. | |
| 402 | 14. Inbound TCP for git over SSH. | |
| 403 | 15. Wrangler: visible credential source, named missing permissions, config `account_id` everywhere. | |
| 404 | ||
| 405 | ## What worked well | |
| 406 | ||
| 407 | - **One Durable Object per repository.** The right isolation unit. A hot repository does not affect | |
| 408 | its neighbours, and there is nothing to shard by hand at the repository level. | |
| 409 | - **Smart HTTP compatibility.** Stock git, protocol v0 and v2, clones, fetches, shallow clones and | |
| 410 | pushes worked through our proxy from the start. Because Artifacts speaks real git, our proxy only adds | |
| 411 | authentication, policy and caching. | |
| 412 | - **Speed when warm.** With a kept credential and a cached ref advertisement, `git fetch` with nothing | |
| 413 | new answers in 0.35 to 0.42 s from Denver, and a cached `info/refs` in 3 to 6 ms on our side. | |
| 414 | - **Forks as a primitive.** `fork()` is one call. Per-pull-request isolation for agents fell out of it | |
| 415 | almost for free. | |
| 416 | - **Scoped, expiring tokens.** `read` and `write` scopes and TTLs from 60 s let us hand sandboxes a | |
| 417 | credential that dies on its own. | |
| 418 | - **Durability design.** Synchronous replication plus R2 snapshots is the right design, clearly | |
| 419 | explained in the launch post and docs. | |
| 420 | - **The binding reads.** `readTree`, `readBlob`, `log` and `readCommit` power every web page, blame, | |
| 421 | mergeability and search indexing without a git client. | |
| 422 | - **Typed errors.** `ArtifactsError` codes (`NOT_FOUND`, `*_IN_PROGRESS`, `MEMORY_LIMIT`) are easy to | |
| 423 | handle precisely. | |
| 424 | - **The rest of the platform.** Workers with Rust and WebAssembly, D1 with free read replication, KV, | |
| 425 | R2, Queues, Containers with three instance sizes, and Cloudflare for SaaS gave a small team a global | |
| 426 | git platform with no servers to run. With placement off, most pages answer in under 250 ms. | |
| 427 | ||
| 428 | We are glad to share traces, test repositories or a call with the teams involved. Contact us through | |
| 429 | g1t.sh. |