Skip to content
910 linesCodeBlameRaw

Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.

Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow1# Deploying g1t
2
3How g1t.sh gets to Cloudflare: one manifest that lists every deployable
4unit, one tool that deploys only what changed, and a g1t Actions workflow
5that runs that tool on every push to `main`. Internal: the public
6self-hosting guide is `apps/docs/src/content/docs/guides/self-hosting.md`.
7
8| Piece | Where |
9| --- | --- |
10| The manifest | `deploy/stack.jsonc` |
11| The tool | `scripts/deploy.mjs` (library and tests in `scripts/deploy/`) |
12| Rust Worker builds | `scripts/build-rust-worker.mjs`, every Rust unit's build command |
Fast pages, required checks on the branch, self-hosted runners, honest incidents13| The runner's images | `services/runner/base/Dockerfile`, `services/runner/Dockerfile`, `services/runner/base.json`, `scripts/build-runner.mjs`, `scripts/deploy/image.mjs` |
14| The workflows | `.g1t/workflows/deploy.yml`, `.g1t/workflows/runner-base.yml` |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow15| The old entry point | `scripts/deploy.sh`, now a wrapper |
Merge platform pause and the hourly usage watcher: staff can pause compute, schedules, indexing or renders for everyone, the watcher emails on a breach and is never blind quietly, and the models proxy holds each run to its cap (billing 0051, integrations 0006)16| Spend guardrails (platform pause, hourly usage watch) | [SPEND-GUARDRAILS.md](SPEND-GUARDRAILS.md) |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow17
18## The manifest
19
20`deploy/stack.jsonc` names every unit: each folder with a `wrangler.jsonc`.
21It is JSONC rather than TOML so that Node reads it with no dependency,
22with the same parser as the Wrangler configs, and comments stay possible.
23
24| Field | |
25| --- | --- |
26| `path` | The unit's folder. |
27| `kind` | `rust-worker` (built by worker-build), `ts-worker` (Wrangler bundles it), `react-router` (`vite build` first), `astro` (`astro build` first). |
28| `worker` | The Worker's name. Must match its `wrangler.jsonc`. |
29| `d1` | `{ database, migrations }`, when it has a database. Must match its `wrangler.jsonc`. |
30| `stage` | `core`, `edge` or `front` (below). |
31| `secrets` | The Wrangler secrets it needs, by name. `node scripts/deploy.mjs doctor` checks they are set. |
32| `setup` | One-time steps no config can say, for a first deploy. |
33| `inputs` | Files outside its folder it is built from that no workspace metadata names. A test finds such imports. |
Fast pages, required checks on the branch, self-hosted runners, honest incidents34| `image` | A Containers image: `dockerfile` (the image deployed), `crate` (the binary it adds), `base` (`{ context, lock }`: the base image's folder and the file recording the base that was pushed) and `repository` (where both are pushed). See [the runner's images](#the-runners-images). |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow35| `self_host` | `run`, `off`, `separate` or `none`: what `deploy/self-host/configs.mjs` does with it. |
36
37What a unit is **built from** is never listed by hand. The tool reads it:
38Rust path dependencies from `cargo metadata`, workspace packages from each
39`package.json`, closed transitively. `node scripts/deploy.mjs manifest`
40prints the result:
41
42```
43unit stage kind worker d1 built from (besides its folder)
44events core rust-worker g1t-events g1t-events crates/contracts crates/kit
45repos core rust-worker g1t-repos g1t-repos crates/contracts crates/kit crates/scan crates/secrets
46runner core ts-worker g1t-runner crates/actions crates/runner packages/contracts
47og core ts-worker g1t-og packages/contracts
48web front react-router g1t packages/contracts packages/theme
49...
50```
51
52Root files count too: `Cargo.toml`, `Cargo.lock` and
53`scripts/build-rust-worker.mjs` for Rust units; `package-lock.json` and
54`tsconfig.base.json` for the others; `package.json` for all. A lockfile
55change counts for a unit only if a package in that unit's graph changed,
56read from the lockfile itself (`scripts/deploy/lockfiles.mjs`), so bumping
57`sharp` for the docs does not redeploy the Rust services.
58
59### Stages
60
61| Stage | What | Units |
62| --- | --- | --- |
63| `migrations` | Every pending D1 migration, in parallel, before any code | each unit's `d1` |
64| `core` | Services reached through bindings | the services, `og` |
65| `edge` | Public endpoints other than the site | `api`, `models`, `pages`, `status` |
66| `front` | The site, sudo, the docs | `web`, `sudo`, `docs` |
67
68A stage starts only when the one before it succeeded. Inside a stage units
69deploy in parallel. The rule the tests enforce: **a unit binds only to units
70in its own stage or an earlier one**, so new code never calls a service
71that has not shipped. (Services in `core` bind to each other in cycles,
72which is why they share a stage.)
73
74### What reads the manifest
75
76- `scripts/deploy.mjs`: everything below.
77- `deploy/self-host/configs.mjs`: which Workers a self-hosted installation
78 runs (`self_host: "run"`, the site first) and which are bound to the off
79 Worker (`"off"`). Its output is identical to before, apart from the order
80 of `workers.txt` after the site.
81- Tests (`npm run test:deploy`) check that: every `wrangler.jsonc` in the
82 repository has a unit; `worker`, `d1` and `image` match the configs;
83 every KV id is named under `resources.kv`; stages follow bindings; the
84 derived dependencies agree with Cargo's own resolved graph; every import
85 that leaves a unit's folder is covered; `deploy/self-host/Dockerfile`
86 builds exactly the Rust units self-hosting runs; and the service table in
87 `docs/SELF_HOSTING.md` names every unit.
88
89## The tool
90
91```sh
92node scripts/deploy.mjs plan # what would deploy, and why (read-only)
93node scripts/deploy.mjs deploy # migrations, then every changed unit
94node scripts/deploy.mjs deploy --only web,api # just these, if they changed
95node scripts/deploy.mjs deploy --only web --force # just this, changed or not
96node scripts/deploy.mjs deploy --all # everything
97node scripts/deploy.mjs build --only events # build as a deploy would; upload nothing
98node scripts/deploy.mjs migrate # pending migrations only
99node scripts/deploy.mjs manifest [--check|--json]
100node scripts/deploy.mjs doctor # secrets each unit lacks
Fast pages, required checks on the branch, self-hosted runners, honest incidents101node scripts/deploy.mjs build-base # build and push the runner's base image (Docker)
102node scripts/deploy.mjs image # build and push the runner's image for this checkout (Docker)
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow103scripts/deploy.sh [units...] # the old entry point: all, or those named, always
104```
105
106| Flag | |
107| --- | --- |
108| `--only a,b` / `--skip a,b` | Units by short name, folder or Worker name. |
109| `--all` | Every unit, changed or not. |
110| `--force` | Deploy the selected units even if unchanged. |
111| `--concurrency N` | Units at once inside a stage, and migrations at once (default 4). |
112| `--stage core` | One stage only. |
113| `--no-migrations` | Skip the migrations step (the workflow runs it as its own job). |
114| `--allow-dirty` | Deploy with uncommitted changes in what deploys. The version records no commit, so the next plan deploys it again. |
Fast pages, required checks on the branch, self-hosted runners, honest incidents115| `--rebuild-image` | Build the runner's image even if one for this source is already in the registry. |
116| `--rebuild-base` | Build and push a new base first (needs Docker), then the runner's image on it. Writes `services/runner/base.json`: commit it. |
117| `--no-push` | `build-base` and `image`: build locally, push nothing. |
118| `--no-cache` | `build-base`: build every layer again. |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow119| `--since REV` | Treat Workers with no recorded commit as running `REV`. Used once to adopt Workers deployed before this tool. |
120| `--json`, `--out FILE`, `--github-output` | The plan as data, for the workflow. |
121
122### Where the deployed commit is kept
123
124On the Worker itself. Every deploy runs `wrangler deploy --message
125"g1t-deploy <40-char sha> <subject>" --tag g1t-<12-char sha>`, which
126Cloudflare keeps as the version's `workers/message` and `workers/tag`
Merge a faster deploy plan: Cloudflare's API read directly and in parallel (2 s against 4 min), and the plan runs beside the check127annotations. `plan` reads them back with two read-only requests per unit
128to Cloudflare's API, the ones `wrangler deployments status` and `wrangler
129versions list` make: `GET /accounts/{account_id}/workers/scripts/{worker}/deployments`
130(the live version) and `GET .../versions?deployable=true` (its
131annotations). Pending migrations are one more per database: `POST
132/accounts/{account_id}/d1/database/{database_id}/query` with `SELECT name
133FROM "d1_migrations"` (the table Wrangler keeps; a unit's
134`migrations_table` if it names one), compared with the `.sql` files in the
135unit's migrations folder. Every request starts at once, at most 16 in
136flight, each tried once more after a 403, 429, 5xx or network error, so a
137plan of every unit takes a few seconds and needs no Wrangler. No KV
138namespace or other infrastructure is needed.
139
140The API is used when the tool has the token Wrangler would be given (in CI,
141`CLOUDFLARE_API_TOKEN`; on a laptop, `CLOUDFLARE_DEPLOY_TOKEN`) and
142`CLOUDFLARE_ACCOUNT_ID` (or g1t's account by default). Without a token
143(your `wrangler login`) the plan asks Wrangler instead, with the same
144answers, a few units at a time; that takes minutes. Either way a Worker
145that does not exist (404, or Cloudflare's code 10007) is "never deployed",
146and any other failure is a reason in the plan, never a crash.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow147
148- A version made by `wrangler secret put` keeps the code of the one before
149 it, so the tool looks through those to the deploy before.
150- A version deployed any other way (by hand, from the dashboard) has no
151 commit, and the unit is deployed again.
152- During a gradual rollout the version with the most traffic counts.
153- After `wrangler rollback`, the plan sees the older commit and deploys
154 what changed since.
155
156### What a deploy does
157
1581. Plans: for each unit, the live commit; `git diff` from it to `HEAD`;
159 whether the changed files touch the unit (its folder, the crates and
160 packages it is built from, its inputs, lockfile changes that reach it).
Merge remote-tracking branch 'origin/main' into workspace-chat161 A Rust crate's `tests/`, `benches/` and `examples/` are not what it is
162 built from, and neither is a source file compiled only for tests: one
163 whose every `mod` declaration is under `#[cfg(test)]` (with or without
164 `#[path]`), or inside a module that is. A change to a crate's tests
165 alone deploys nothing (`scripts/deploy/stack.mjs`, `testOnlySource`).
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow1662. Refuses if uncommitted changes touch what would deploy (`--allow-dirty`).
1673. Applies every pending migration (`wrangler d1 migrations apply --remote`),
168 in parallel. Any failure stops the deploy before code.
1694. Installs worker-build once if any Rust unit is deploying.
1705. Each stage in turn: units in parallel (`--concurrency`), each `npm run
171 build` first for React Router and Astro, then `wrangler deploy` in the
172 unit's folder with the annotation. If any unit fails, later stages are
173 not started.
1746. Prints a table: unit, stage, result, version id, time. Each unit's full
175 output is kept in `$TMPDIR/g1t-deploy/<unit>.log`.
176
Deploying: a migration keeps the live code right, because it runs a minute before the new code177#### Migrations run while the old code is still live
178
179Migrations apply before any code, and a stage takes a minute or more, so
180for that long the code in production is the old code reading the new
181schema. A migration must keep the old code right:
182
183- **Add, never change meaning.** New tables and columns (with defaults) are
184 safe. Rewriting what existing rows mean is not: on 2026-10-08,
185 `billing/0039` turned comped accounts into `custom` terms at 100%, and
186 for about 40 seconds the old billing code, which knew only `comped`, saw
187 flagon-io as a free workspace and refused its workflows.
188- **Change meaning in two deploys.** First ship code that reads both the old
189 and the new form (and keeps writing the old one); then, in a later
190 deploy, the migration that rewrites the rows; then, if you like, code
191 that drops the old form.
192- **Never drop or rename** a column or table the live code reads in the same
193 deploy that stops reading it.
194
Fast pages, required checks on the branch, self-hosted runners, honest incidents195#### Telling the status page about a deploy
196
197Restarts during a deploy can make a part slow for a minute, which the
Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders198status page's checks would otherwise draft as an incident. So
199`scripts/deploy.mjs deploy` says when it starts its stages and when they
200end (`withDeployWindow` in `scripts/deploy/status-window.mjs`; never in a
201dry run, and only once there is something to ship), by hand and in g1t
202Actions alike. It posts to `POST https://status.g1t.sh/deploys` with
Fast pages, required checks on the branch, self-hosted runners, honest incidents203`Authorization: Bearer $STATUS_DEPLOY_TOKEN`, the same value as the status
Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders204Worker's `STATUS_DEPLOY_TOKEN` secret, and `{"phase": "started" |
205"finished", "id": "<commit>"}`. During the deploy and for 3 minutes after
206it, detection keeps counting failed and slow checks but makes no new
207draft; trouble that outlasts that is drafted with its true start. Deploys
208that overlap (the jobs of one stage run at once, each announcing itself)
209are one window: the status Worker counts the starts, and the window closes
210when the last one finishes. A start with no finish stops counting 30
211minutes after the latest start. Without the token nothing is sent, and an
212announcement never fails a deploy: a refusal or network error is one
213warning line. By hand: `node scripts/deploy/status-window.mjs
214started|finished [id]`.
215
216To turn it on (once; until then deploys are not announced):
Fast pages, required checks on the branch, self-hosted runners, honest incidents217
Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders2181. Make a token and set it as the status Worker's secret:
219 `npx wrangler secret put STATUS_DEPLOY_TOKEN` in `apps/status` (as of
220 2026-10-08 the Worker has only `STATUS_SECRET`). Without it,
221 `POST /deploys` answers 404.
2222. Set the same value as the **`STATUS_DEPLOY_TOKEN`** Actions secret on
223 flagon-io/g1t (Settings, Secrets and variables, or
224 `PUT /repos/flagon-io/g1t/actions/secrets/STATUS_DEPLOY_TOKEN`);
225 `.g1t/workflows/deploy.yml` passes it to every deploy job.
2263. Add `status.g1t.sh | deploy.yml | production` to the project's
227 **Workflow-only domains** (see Network below), or the job's request is
228 refused by the guardrails (the deploy still goes on, with a warning).
2294. For deploys by hand, set `STATUS_DEPLOY_TOKEN` in your shell's
230 environment (the tool does not read `.env`).
231
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow232On a laptop the tool uses your `wrangler login` (or `CLOUDFLARE_DEPLOY_TOKEN`
233if set), as `scripts/deploy.sh` always did: a `CLOUDFLARE_API_TOKEN` or
234global API key in your shell, or in the repository's `.env`, is ignored.
Merge a faster deploy plan: Cloudflare's API read directly and in parallel (2 s against 4 min), and the plan runs beside the check235With `CI=true` it uses `CLOUDFLARE_API_TOKEN`. With a token, `plan` reads
236Cloudflare's API itself; with only `wrangler login`, it asks Wrangler (see
237"Where the deployed commit is kept").
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow238
Fast pages, required checks on the branch, self-hosted runners, honest incidents239### The runner's images
240
241`services/runner` runs every sandbox (agents, checks, the merge queue,
242workflow jobs, g1t.page builds) from one Containers image, made in two
243parts:
244
245| Image | Built from | Holds | Rebuilt |
246| --- | --- | --- | --- |
Merge branch 'main' into actions-toolkit-oidc-artifacts247| **Base**, `g1t-runner:base-<date>-<inputs>` | `services/runner/base/Dockerfile` | Debian bookworm, Node 24, Python 3.11, Docker (Engine, Buildx, Compose, from Docker's apt repository), Go (from go.dev), Rust stable for the `node` user with rustfmt, clippy and the `wasm32-unknown-unknown` target, build-essential, musl-tools, git, ripgrep, jq, zstd, sudo, and the pinned Claude Code CLI on top | When its folder changes, weekly, or by hand (`build-base`) |
Fast pages, required checks on the branch, self-hosted runners, honest incidents248| **Runner**, `g1t-runner:<content hash>` | `services/runner/Dockerfile`: `FROM` the base, plus one file | The g1t runner, a static binary | When the binary or the base changes |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow249
Fast pages, required checks on the branch, self-hosted runners, honest incidents250Both are pushed to one repository of Cloudflare's registry,
251`registry.cloudflare.com/<account>/g1t-runner`, so pushing the runner's
252image uploads only its own layer (about 5 MB): the base's layers are
253already there.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow254
The runner's base image, pushed as Wrangler builds images, with pushes that fail loudly255Both are built as Wrangler builds images (`--platform linux/amd64
256--provenance=false --sbom=false`): one manifest, not an OCI index with a
257BuildKit attestation beside it. A push is tried up to three times and counts
258only when Docker reports the digest; the first push of the base once failed
259with `blob unknown to registry` and went through when run again (see
260`docs/CLOUDFLARE_FEEDBACK.md`, C3). `build-base` and `image` exit non-zero
261when a push fails, and `base.json` is written only after the push.
262
Fast pages, required checks on the branch, self-hosted runners, honest incidents263**The base** is recorded in `services/runner/base.json`: its reference, its
264digest, a hash of its folder (`inputs`), when it was built, its size and
265each toolchain's version. `node scripts/deploy.mjs build-base` builds it,
266pushes it and rewrites the file; commit the file, and the next deploy
267builds the runner's image on it. `npm run test:deploy` fails while the
268folder and the file disagree, so a change to the base's Dockerfile cannot
269merge without the base it describes. Layers go from what changes least to
Merge runner image: Java 21, .NET 8, Ruby 3.3270most (system packages, Go, Rust, Java, .NET and Ruby, the Claude Code
271CLI), and the apt and npm caches stay in BuildKit's cache, out of the
272image. Ruby is compiled from source in a single step (a few minutes on a
273cold cache) that removes its source tree before the layer is written.
Fast pages, required checks on the branch, self-hosted runners, honest incidents274
275The base's build cache is the base itself: it is built with
276`BUILDKIT_INLINE_CACHE`, which records in the image how each layer was
277made, and `build-base` builds `--cache-from` the base it replaces. A
278machine with an empty cache, or one just pruned, pulls the unchanged
279layers from the registry instead of building them. (BuildKit's other
280registry cache, `--cache-to type=registry`, pushes a separate cache
281manifest that not every registry takes, and for a one-stage image adds
282nothing the inline cache lacks.)
283
284**The runner binary** (`crates/runner`) is built outside Docker by
285`scripts/build-runner.mjs` as one static binary for
286`x86_64-unknown-linux-musl`, so it runs on the base whatever its libc, and
287on a self-hosted runner's machine too. Where it is built:
288
289- on x86-64 Linux with the musl target and `musl-gcc`
290 (`rustup target add x86_64-unknown-linux-musl`, `apt-get install musl-tools`),
291 with the machine's own Cargo;
292- anywhere else (Windows, macOS) in a small builder container, Rust on
293 Alpine (whose own target is musl), with Docker volumes keeping Cargo's
294 registry and target directory between builds. Windows has no musl
295 cross-linker, and `ring` (under ureq's TLS) needs a C compiler for the
296 target, so a container is the dependable route.
297
298**The runner's image tag** is a hash of everything it is built from: its
299Dockerfile, `base.json`, the crates the binary is built from, the
300workspace's Cargo files and the build script. The same source always names
301the same image, so:
302
3031. A deploy computes the tag and asks the registry whether it is there
304 (a `HEAD` of its manifest, with credentials from Wrangler; no Docker).
3052. If it is, nothing is built: the deploy uses it.
Merge branch 'main' into actions-toolkit-oidc-artifacts3063. If not, it builds the binary and the image (seconds on a warm machine)
307 and pushes it. In `deploy.yml` that is the `runner-image` job, on
308 `g1t-4core`, with the job's own Docker Engine (see
309 [Docker in workflow jobs](#docker-in-workflow-jobs)): it adds the musl
310 target, builds the binary natively (the base has `musl-gcc`), pulls the
311 base from Cloudflare's registry, builds, and pushes one layer.
3124. If Docker does not answer (a machine without it, or jobs with Docker
313 turned off), the unit fails saying to run `node scripts/deploy.mjs
314 image` on a machine with Docker; then re-run the workflow.
Fast pages, required checks on the branch, self-hosted runners, honest incidents315
316Then `wrangler deploy` is given the image by reference, from a generated
317config (`services/runner/wrangler.deploy.json`, deleted after, ignored by
318git), so Wrangler builds nothing. When nothing the image is built from
319changed since the runner's live commit, the deploy also passes
320`--containers-rollout none`, which leaves running sandboxes alone.
321
322To get a new base out:
323
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow324```sh
Fast pages, required checks on the branch, self-hosted runners, honest incidents325node scripts/deploy.mjs build-base # build, push, write base.json (needs Docker)
326git commit services/runner/base.json -m "A new base image for g1t's sandboxes"
327node scripts/deploy.mjs image # optional: push the runner's image now, so CI finds it
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow328```
329
Fast pages, required checks on the branch, self-hosted runners, honest incidents330`.g1t/workflows/runner-base.yml` does the same weekly (and when the
331base's folder changes on `main`), and opens a pull request with
Merge branch 'main' into actions-toolkit-oidc-artifacts332`base.json`. It runs on a self-hosted runner with the `docker` label
333(`runs-on: [self-hosted, docker]`): g1t's own machines have Docker now,
334but the base's build downloads from Docker's apt repository over HTTPS,
335which does not trust a guarded job's egress certificate, so it stays on an
Runner base workflow: skipped until a self-hosted runner exists336open network. Its job is skipped until the repository variable
337`RUNNER_BASE_SELF_HOSTED` is `true` (set it once a runner is registered;
338without one, runs waited in the queue for a day and a half and then
339failed); until then run `build-base` by hand.
Fast pages, required checks on the branch, self-hosted runners, honest incidents340
341**Sandboxes start from the image.** Cloudflare pulls an image to a
342machine the first time a sandbox lands there, and keeps it. A smaller base
343pulls sooner, and a change to the runner alone sends machines one 5 MB
344layer instead of the whole image.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow345
Fast pages, required checks on the branch, self-hosted runners, honest incidents346#### Larger machines
347
348The same image runs on three instance types, each a Durable Object class
349of its own in `services/runner/wrangler.jsonc`: `AttemptSandbox`
350(`standard-1`), `Sandbox2Core` (`standard-3`) and `Sandbox4Core`
351(`standard-4`). Workflow jobs choose with `runs-on: g1t-2core` or
352`g1t-4core` (`g1t_contracts::actions::INSTANCE_TYPES`); the actions
353service passes the label to the runner, which starts the job in that
354class. Billing prices the larger ones from their memory and disk, and
355their CPU (see the public billing guide). The account's Containers limits
356must allow `standard-4`; Wrangler refuses the deploy otherwise.
357
Merge branch 'main' into actions-toolkit-oidc-artifacts358#### Docker in workflow jobs
359
360Workflow jobs on g1t's machines have a Docker Engine of their own
361(`crates/runner/src/docker/`; the public guide is
362`apps/docs/src/content/docs/guides/actions.md`, "Docker"). What Cloudflare
363Containers allow decides how it runs (findings in `docs/PLAN.md`,
364"Docker in workflow jobs"):
365
366| Piece | What it does |
367| --- | --- |
368| `dockerd` | Started as root with `sudo`, only when the job first uses Docker or has `services:` or `container:`. Flags: `--iptables=false --ip6tables=false --ip-forward=false` (Containers allow neither), the containerd image store, Docker Hub through `mirror.gcr.io`. Its config, socket and log are in `/run/g1t-docker` (`dockerd.log` is the place to look); its data in `/var/lib/docker`, on overlays when the disk takes them, plain copies (`native`) when not. |
369| `/var/run/docker.sock` | The runner's API proxy (`docker/api.rs`), which starts the Engine on the first connection. Containers that ask for a bridge network get the job's own (`host`), the names they would have had resolve to 127.0.0.1, and ports published under another number are forwarded. |
370| `runc` | The Engine finds the runner binary first on its `PATH` as `runc` (`docker/oci.rs`): BuildKit's `RUN` steps join the job's network, and in a guarded job every container gets the egress certificate at `/dev/g1t-egress`. Then the real `/usr/bin/runc` runs. |
371| cgroups | Before the Engine starts, the sandbox's processes move to a cgroup of their own and every controller is handed down, as Docker's own Docker-in-Docker image does, so `--cpus` and `--memory` work. |
372
373- **Turning it off:** set the runner Worker's `DOCKER` var to `off`
374 (`services/runner/wrangler.jsonc`) and deploy the runner: new jobs get
375 no Engine (`G1T_DOCKER=off`), and jobs that need one fail saying Docker
376 does not answer. Anything else is on.
377- **Network:** containers share the job's network, so the guardrails, the
378 workflow-only domains and the egress Worker apply to them unchanged. The
379 public registries are in `BUILD_HOSTS` (`services/runner/src/egress.ts`).
380- **Isolation:** one Engine per job, inside the job's sandbox (its own
381 VM), gone with it. The job already had root through `sudo`; Docker adds
382 no reach beyond the sandbox, and no host socket is ever mounted into one.
383- **Checked locally** (2026-10-08, Docker Desktop, the base and runner
384 image built from this tree, a privileged container standing in for a
385 sandbox, a pretend API): services with health checks, `localhost` and
386 names, port forwarding, `docker build` with a networked `RUN`, Compose
387 with a healthy dependency, `docker://` steps, a Dockerfile action, a
388 `container:` job with a JavaScript action, an Alpine job container, the
389 egress certificate in `run`, `exec` and build steps and in no layer,
390 plain-copy storage, and the deploy's own build and push to a registry.
391 Not yet seen on Cloudflare itself: watch the first runs' logs for the
392 `Docker: started` line, and `dockerd.log` if it does not come.
393
Merge remote-tracking branch 'origin/main' into workspace-chat394#### Sandbox errors
395
396Each sandbox is a Durable Object of `g1t-runner`, and Cloudflare counts
397each of its invocations: the `run` call that starts it, `halt`,
398`noteBlocked` and `flagAbuse`, and its alarms. The containers library keeps
399an alarm going for as long as the container runs (each one waits up to
400three minutes), and runs `onStop` from the alarm after the container
401exits, so most of a sandbox's invocations are alarms.
402
403To see what the errors are, by namespace and status, and what was thrown:
404
405```sh
406CLOUDFLARE_API_TOKEN=<token> node scripts/ops/runner-errors.mjs # the last 7 days
407CLOUDFLARE_API_TOKEN=<token> node scripts/ops/runner-errors.mjs --days 30 --json
408```
409
410The token needs Account Analytics: Read and Workers Observability: Read
411(Workers Scripts: Read adds namespace names). The report puts each message
412in a bucket and says whether it is expected:
413
414| Bucket | Expected | What it is |
415| --- | --- | --- |
416| `deploy_reset` | Yes | A runner deploy resets every sandbox's object, failing the alarm or call in flight. The container keeps running and the next alarm picks it up. |
417| `no_capacity` | Yes | No container instance was free (`max_instances`). The work fails to start and says so. |
418| `container_exited`, `caller_gone` | Yes | A container that stopped, or a caller that went away first. |
419| `stop_not_reported` | No | `onStop` could not tell a service that the sandbox stopped after two tries. The five-minute sweep catches the work up. |
420| `alarm_failed`, `not_started`, `storage`, `limits`, `other` | No | Read the message. |
421
422The runner logs these at error level, each with a fixed prefix you can
423search for in Workers Logs:
424
425- `sandbox not started`
426- `sandbox stop not reported`
427- `sandbox alarm failed`
428- `sandbox container error`
429- `sandbox not destroyed`
430
431`onStop` never throws. A throw would fail the alarm, which Cloudflare
432retries and counts as an error each time, running the whole stop again.
433A sandbox's run inside its time cap is not stopped for inactivity. The
434library's `sleepAfter` (100 minutes) would otherwise stop a run whose
435guardrails allow longer, up to 240 minutes, because the runner never
436fetches the container, so to the library every sandbox looks idle.
437
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow438## Build speed
439
440Measured on the development machine (Windows, 32 cores, warm Cargo cache),
441building `events`, `search` and `repos` after a change to `crates/kit`, as a
442deploy does but without uploading (`wrangler deploy --dry-run`):
443
444| | Time |
445| --- | --- |
446| Before: one after another, `cargo install worker-build` each time, wasm-opt `-O` | 81 s, 84 s |
447| Concurrent builds, wasm-opt `-O` | 50 s, 68 s |
448| Concurrent builds, wasm-opt `-O1` (now) | 28 s, 33 s |
449
450Where the time went, and what changed:
451
452- **wasm-opt** was most of it: `-O` took 38 s on `repos` and 19 s on `api`;
453 `-O1` takes 1 to 3 s. With worker-build's flags (it keeps the names
454 section) the `.wasm` is 24 to 28% larger raw but only 1 to 7% larger
455 gzipped, and Workers' size limit is on the compressed upload. `-Os` and
456 `-Oz` were no faster than `-O`. Set per crate in
457 `[package.metadata.wasm-pack.profile.release]`; a test keeps every Rust
458 unit on the same level.
459- **worker-build** is installed only when missing or another version
460 (`scripts/build-rust-worker.mjs`, which pins it). Locally `cargo install`
461 on an installed version cost under a second; on a fresh CI sandbox it is
462 a full compile, which the workflow caches instead.
463- **One Cargo target**: every Rust unit is a member of the workspace, so
464 they already share `target/`. Concurrent builds take turns on Cargo's
465 lock for the compile, and their wasm-bindgen, wasm-opt and uploads
466 overlap.
467- **No joint `cargo build -p a -p b`**: tried, and it is slower. Cargo
468 unifies features across the packages of one build (`serde_json`'s
469 `preserve_order` from `api` and `actions`, `digest` features from
470 `secrets`), so each worker-build afterwards compiled its own variant
471 again.
The runner's base image, pushed as Wrangler builds images, with pushes that fail loudly472- **The runner's images**, measured on the same machine on 2026-10-06
473 (Docker Desktop, 8 vCPUs; its disk was busy with other containers, so
474 the cold figures are slow and noisy):
475
476 | | Before (one image) | Now |
477 | --- | --- | --- |
478 | Size, unpacked / compressed (what a machine pulls) | 3.08 GB / 819 MB | 2.68 GB / 686 MB |
479 | A change to the runner | Docker rebuilds the image's Rust stage and pushes the image (1198 s in the first deploy after a prune) | binary 46–53 s (6 s unchanged), image 7 s, push one 5 MB layer (1.4 s to a local registry) |
480 | The base from nothing | 431 s (whole image, cold) | 758 s cold, rarely: weekly or when its folder changes |
481 | The base after `docker builder prune` | as from nothing | 68 s, its layers pulled from the registry it was pushed to |
482 | The runner binary, cold (builder container) | | 99 s |
483 | A Rust CI job's build (events, search, repos; 4 vCPUs) | | 51 s cold, 13 s with the Cargo target restored (107 MB zstd entry) |
484
Merge remote-tracking branch 'origin/main' into workspace-chat485- **sccache** (`scripts/sccache.sh`, pinned by version and sha256): a
486 checkout gives every source a new mtime, so Cargo calls rustc again for
487 every workspace crate a unit uses, however much of `target/` was
488 restored. With `RUSTC_WRAPPER=sccache` and its GitHub Actions backend,
489 each call that makes a library is looked up by its inputs in the
490 repository's Actions cache (the same cache as `actions/cache`, scoped to
491 `main`), and an unchanged crate comes back from it. Measured on the
492 development machine on 2026-10-08, `repos` and `events` for wasm32,
493 release, `-j 4`, the dependencies already built, every workspace source
494 touched as a checkout does (sccache's local disk cache; the Actions
495 cache adds a download per hit):
496
497 | | Time |
498 | --- | --- |
499 | Without sccache (before) | 44.3 s, 44.6 s |
500 | sccache, empty cache (its first run, filling it) | 48.8 s |
501 | sccache, filled (6 hits: `contracts`, `kit`, `rules`, `scan`, `secrets`, `blobstore`) | 24.2 s |
502
503 CI's build the same way (`cargo test --workspace --no-run`, debug,
504 `-j 4`, `CARGO_INCREMENTAL=0`): 70.2 s and 74.6 s without sccache, 66.5 s
505 filling it, 40.6 s filled (7 library hits; 23 test harnesses and
506 cdylibs compiled).
507
508 What is left is each Worker's own crate: a `cdylib` is linked, and
509 sccache does not cache what rustc links (binaries, cdylibs, proc
510 macros, build scripts, test harnesses). That compile is real work
511 anyway: a unit deploys because it, or a crate it uses, changed. The
512 same holds for CI's `cargo test`: the workspace's libraries come back
513 from the cache, the test harnesses are compiled. Every Rust job of
514 Deploy and CI's Rust job end with sccache's hits and misses in the
515 run's summary. If the cache cannot be reached, or the download fails
516 its checksum, the job warns and builds as before. On g1t, sccache
517 speaks the toolkit cache's older protocol (`GITHUB_SERVER_URL` is not
518 github.com): a lookup, a reservation, one `PATCH` of the whole entry
519 (`Content-Range: bytes 0-N/*`) and a commit per miss; a lookup and a
520 blob `GET` per hit. Its storage check saves `sccache/.sccache_check`
521 once, and takes the `409` on later runs as "already there".
522- **Not restoring mtimes:** setting each file's mtime from git (so Cargo
523 would trust the restored `target/`) was considered and left out. A
524 restored `target/` may come from another commit (the cache's
525 `restore-keys` take the nearest earlier entry, of any group), and a
526 file changed by an older commit than that build would look unchanged:
527 a stale crate deployed. sccache looks at what is compiled, not when.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow528- **Only what changed** is the largest saving: a change to one service
Merge remote-tracking branch 'origin/main' into workspace-chat529 deploys one service, and a change only to a crate's tests deploys
530 nothing.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow531
532## The workflow
533
534`.g1t/workflows/deploy.yml` runs on every push to `main`, and by hand
535(**Actions → Deploy → Run workflow**) with `units` (deploy these, changed or
536not), `all` and `dry_run` (plan only).
537
538| Job | Does | Needs |
539| --- | --- | --- |
540| `check` | `manifest --check` and `npm run test:deploy` | — |
Merge a faster deploy plan: Cloudflare's API read directly and in parallel (2 s against 4 min), and the plan runs beside the check541| `plan` | `plan --github-output`: outputs per stage, the plan in the run's summary. Reads Cloudflare's API itself, so it installs nothing. | — (runs beside `check`) |
542| `migrate` | `migrate --only <units with pending migrations>` | `check` and `plan`; skipped when none are pending |
543| `core`, `edge`, `front` | `deploy --only <units> --force --no-migrations`, one job per build group | `check`, `plan` and the stages before; skipped when empty |
544| `smoke` | `node scripts/ops/smoke.mjs`: the landing page, sign-in, sign-up and pricing load, and the waitlist form reaches identity (sent an address identity refuses before keeping or counting anything, so the real waitlist is never touched) | `check`, `plan` and every stage; skipped when nothing deployed |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow545
546- **One at a time:** `concurrency: deploy-production`, never cancelled in
547 progress; a second push waits.
Merge a faster deploy plan: Cloudflare's API read directly and in parallel (2 s against 4 min), and the plan runs beside the check548- **Check and plan side by side:** neither waits for the other, so the
549 plan's few seconds overlap the check's tests; nothing migrates or deploys
550 until both have succeeded. The plan job keeps `fetch-depth: 0`: it diffs
551 from each Worker's live commit, which may be any commit, and tells a
552 rollback by ancestry, which a shallow clone cannot answer.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow553- **Build groups:** a stage's units are split so each job shares a build:
Fast pages, required checks on the branch, self-hosted runners, honest incidents554 Rust workers at most four to a job (each a 4-vCPU `g1t-4core` machine), the
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow555 TypeScript Workers together, each site alone, and a unit whose image must
Merge branch 'main' into actions-toolkit-oidc-artifacts556 be rebuilt alone (`image: true` in the matrix, also on `g1t-4core`, where
557 it builds and pushes the image with the job's own Docker Engine). `fail-fast: false`, so one failed job does not cut
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow558 another off mid-upload; the next stage then does not start.
Merge remote-tracking branch 'origin/main' into workspace-chat559- **Tests:** `.g1t/workflows/ci.yml` tests every pull request into `main`
560 (and, on pushes to `main`, runs only its Rust job, to keep the caches
561 pull requests restore from current). `check` runs the deploy tool's own
562 tests. Deploy does not wait for CI.
Merge branch 'main' into actions-toolkit-oidc-artifacts563- **Machines:** Rust jobs and the runner's image run on `g1t-4core` (4 vCPUs,
564 12 GiB, 20 GB), the others on the standard machine
565 (`runs-on: ${{ (matrix.rust || matrix.image) && 'g1t-4core' || 'ubuntu-latest' }}`).
566 The image job needs the room: the base it builds on is about 3.2 GB
567 unpacked.
Fast pages, required checks on the branch, self-hosted runners, honest incidents568- **Caching** (`actions/cache`: up to 2 GB an entry, 10 GB a repository,
569 kept until unused for 7 days): the worker-build binary, worker-build's
570 downloaded tools, `~/.cargo/registry/cache`, and the Cargo target's
571 release dependencies (`target/release` and
572 `target/wasm32-unknown-unknown/release`, without `incremental` or
Merge remote-tracking branch 'origin/main' into workspace-chat573 `.wasm`), keyed by the build group, `Cargo.lock` and `base.json`. Cargo
574 calls rustc again for the workspace's own crates on every run (a
575 checkout's sources are newer than any cache); **sccache** answers those
576 calls from the repository's Actions cache when a crate's inputs did not
577 change (see [Build speed](#build-speed)). npm's cache is not kept: every
578 job runs `npm ci` of only what its units need (`deploy.mjs install`:
579 Wrangler alone for Rust jobs).
Fast pages, required checks on the branch, self-hosted runners, honest incidents580- **Conditions:** each stage runs with `!failure() && !cancelled()`, which
581 on g1t (as on GitHub) is true when no job before it failed, however far
582 back: a `migrate` job skipped for having nothing to apply does not stop
Merge a faster deploy plan: Cloudflare's API read directly and in parallel (2 s against 4 min), and the plan runs beside the check583 the stages after it, and a failed `check` or `plan` stops all of them
584 (`migrate`'s own condition needs both to have succeeded).
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow585- `crates/actions/tests/repository_workflows.rs` reads the workflow with
586 g1t's own parser and expressions, and checks the jobs start, wait and
587 stop as above (`cargo test -p g1t-actions --test repository_workflows`).
588
589### What the sandbox has
590
Fast pages, required checks on the branch, self-hosted runners, honest incidents591The base image (`services/runner/base/Dockerfile`) has Node 24, npm, git,
Merge runner image: Java 21, .NET 8, Ruby 3.3592Go, zstd, Docker, musl-tools, Rust stable for the `node` user with
593rustfmt, clippy and the `wasm32-unknown-unknown` target, Java (Temurin 21),
594the .NET 8 SDK and Ruby 3.3, but not worker-build or `gh`. Java and Ruby
595sit in the tool cache (`/home/runner/_tool/Java_Temurin-Hotspot_jdk/<semver
596with + as ->/x64` and `/home/runner/_tool/Ruby/<version>/x64`, each with
597an `x64.complete` marker), which is where `actions/setup-java` and
598`ruby/setup-ruby` look; .NET is in `/usr/share/dotnet`, owned by `node`,
599which is where `actions/setup-dotnet` installs. Bumping one means changing
600its version and checksum `ARG`s together (Java's tool cache name is
601Adoptium's `version_data.semver`; .NET's SHA-512 is in its release
602metadata; Ruby's SHA-256 is in `cache.ruby-lang.org/pub/ruby/index.txt`).
603The user docs list what jobs get in
604`apps/docs/src/content/docs/guides/actions.md` (The runner). The workflow's `rustup target add wasm32-unknown-unknown` is
Merge branch 'main' into actions-toolkit-oidc-artifacts605then a no-op, and worker-build is restored from the cache, installed on a
606miss. The image job adds `x86_64-unknown-linux-musl` (about 30 MB from
607`static.rust-lang.org`) and keeps its Cargo target in the cache. worker-build
Fast pages, required checks on the branch, self-hosted runners, honest incidents608fetches wasm-bindgen and wasm-opt from GitHub releases and esbuild from
Merge remote-tracking branch 'origin/main' into workspace-chat609npm. Each Rust job downloads sccache (about 10 MB) from its GitHub release
610too (`scripts/sccache.sh`, which checks its sha256); it is not in the base
611image, and putting it there means a pinned download in the base's
612Dockerfile and a base rebuild (`build-base`). All of those hosts are on
613the list every workflow job may reach.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow614
Fast pages, required checks on the branch, self-hosted runners, honest incidents615### Network
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow616
Fast pages, required checks on the branch, self-hosted runners, honest incidents617A workflow job reaches its project's allowed domains, g1t, what builds
618need (`services/runner/src/egress.ts`, `BUILD_HOSTS`), and the project's
619**workflow-only domains** that name its workflow and environment. Those are
620never reached by agents, checks, the merge queue, deploy builds or runs of
621pull requests from forks (`Guardrails::workflow_hosts`, the runner's
622`jobHosts`). Under flagon-io/g1t's **Settings → Guardrails**
623(Maintain role or higher), **Workflow-only domains**:
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow624
The runner's base image, pushed as Wrangler builds images, with pushes that fail loudly625```text
Fast pages, required checks on the branch, self-hosted runners, honest incidents626api.cloudflare.com | deploy.yml | production
627registry.cloudflare.com | deploy.yml, runner-base.yml | production
Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders628status.g1t.sh | deploy.yml | production
Fast pages, required checks on the branch, self-hosted runners, honest incidents629```
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow630
Merge a faster deploy plan: Cloudflare's API read directly and in parallel (2 s against 4 min), and the plan runs beside the check631`api.cloudflare.com` is Cloudflare's API, which Wrangler and the plan call; `registry.cloudflare.com` is where
Merge branch 'main' into actions-toolkit-oidc-artifacts632the deploy asks whether the runner's image is already built, and where the
633`runner-image` job pulls the base from and pushes the runner's image to
Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders634(as `runner-base.yml` pushes the base); `status.g1t.sh` hears the deploy
635start and finish (see "Telling the status page about a deploy"). The job's Docker Engine shares the
Merge branch 'main' into actions-toolkit-oidc-artifacts636job's network, so these lines are what let it reach the registry. If a pull
637is refused with `g1t guardrails: <host> is not on this project's allowed
638domains`, the registry sent the layers from another host: add that host on
639the same line. Only `deploy.yml`'s and `runner-base.yml`'s
Fast pages, required checks on the branch, self-hosted runners, honest incidents640jobs with `environment: production` reach them, which are also the only
641jobs that can read `CLOUDFLARE_API_TOKEN`. Each change to the list is in
642the workspace's audit log as `update_guardrails`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow643
644### The API token
645
646Create it at **dash.cloudflare.com → My Profile → API Tokens → Create
647Token → Custom token**, named `g1t deploys (CI)`:
648
649| Scope | Permission | Why |
650| --- | --- | --- |
Merge a faster deploy plan: Cloudflare's API read directly and in parallel (2 s against 4 min), and the plan runs beside the check651| Account | Workers Scripts: Edit | Upload, versions, deployments (the plan reads both), crons, bindings, `secret list` (doctor) |
652| Account | D1: Edit | The plan's query of each database's `d1_migrations`, and `d1 migrations apply` |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow653| Account | Queues: Edit | Attaching each unit's queue consumers on deploy |
654| Account | Workers R2 Storage: Read | Wrangler checks `og`'s bucket binding |
655| Account | Account Settings: Read | Wrangler reads the account |
Merge branch 'main' into actions-toolkit-oidc-artifacts656| Account | Containers: Edit | The runner's deploy updates its applications (the image reference, the three classes), and gets registry credentials (`wrangler containers registries credentials --push`, one hour) to look for, pull and push its image. `runner-base.yml` pushes images with it. |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow657| Zone (`g1t.sh`, `g1t.page`) | Workers Routes: Edit | `pages`' zone routes, and custom domains |
658| Zone (`g1t.sh`, `g1t.page`) | DNS: Edit | Custom domains (`api`, `mcp`, `og`, `models`, `status`, `sudo`, `docs`, `g1t.sh`, `g1t.page`) keep their DNS records |
659| Zone (`g1t.sh`, `g1t.page`) | Zone: Read | Finding the zone a route names |
660
661Restrict it to account `syntaqx` (`1e6f2cffa3f445920836e8ebe446bb58`) and
662the two zones. Not needed for deploys: KV (bindings are by id; creating a
663namespace is a one-time setup), Vectorize, Workers for Platforms beyond
664Workers Scripts, Cloudflare for SaaS custom hostnames (the deployments
665service does that at runtime with its own token), SSL and Certificates.
666Workers KV Storage: Edit and Vectorize: Edit are only for first-time setup,
667which stays a person's job.
668
669The list follows what our configs use; Cloudflare does not publish exactly
670what `wrangler deploy` checks for each binding. Bindings with no listed
671permission (Browser Rendering, Workers AI, Vectorize, Email Sending,
672Artifacts, dispatch namespaces) are assumed to need none beyond Workers
673Scripts. Verify on the first run with `workflow_dispatch` and `units:
674pages` (small, no secrets), then `units: og` (R2) and `units: runner`
675(Containers); a missing permission fails with `Authentication error [code:
67610000]` and the route it was refused.
677
678Add it to the repository:
679
6801. On g1t.sh, open **flagon-io/g1t → Settings → Secrets and variables**.
6812. **Add**: key `CLOUDFLARE_API_TOKEN`, type **Secret**, available to
682 **Workflows**, environment **Production**. Only jobs with
683 `environment: production` (the deploy jobs) can read it.
6843. **Add**: key `CLOUDFLARE_ACCOUNT_ID`, type **Variable**, value
685 `1e6f2cffa3f445920836e8ebe446bb58`, available to **Workflows**, all
686 environments.
687
688Or through the API:
689
690```sh
691curl -X PUT https://api.g1t.sh/repos/flagon-io/g1t/actions/secrets/CLOUDFLARE_API_TOKEN \
692 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
693 -d '{"value":"<the token>","environments":["production"],"available_to":["workflows"]}'
694
695curl -X POST https://api.g1t.sh/repos/flagon-io/g1t/actions/variables \
696 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
697 -d '{"name":"CLOUDFLARE_ACCOUNT_ID","value":"1e6f2cffa3f445920836e8ebe446bb58","available_to":["workflows"]}'
698```
699
700## Turning it on
701
7021. Create the token and add the secret and variable (above).
Fast pages, required checks on the branch, self-hosted runners, honest incidents7032. Add `api.cloudflare.com` and `registry.cloudflare.com` to flagon-io/g1t's
704 workflow-only domains, for `deploy.yml` in `production` (above).
7053. Build and push the base once, and commit `services/runner/base.json`:
706 `node scripts/deploy.mjs build-base`. Create the cache bucket:
Actions: OIDC tokens, the toolkit's cache and artifact services, and artifacts in R2707 `npx wrangler r2 bucket create g1t-actions-cache`, with lifecycle rules
708 deleting cache entries (`c/`) 30 days after upload and artifacts (`a/`)
709 after 91, a day past the longest they are kept, and unfinished uploads
710 after a day (the actions service deletes both sooner; the rules catch
711 what it misses):
712 `npx wrangler r2 bucket lifecycle add g1t-actions-cache expire-cache c/ --expire-days 30 --abort-multipart-days 1`
713 and `npx wrangler r2 bucket lifecycle add g1t-actions-cache expire-artifacts a/ --expire-days 91 --abort-multipart-days 1`.
714 A bucket made before artifacts moved there has one rule for everything,
715 `expire`, which would delete artifacts kept longer than 30 days: remove
716 it (`npx wrangler r2 bucket lifecycle remove g1t-actions-cache --id expire`)
717 and add the two above.
Fast pages, required checks on the branch, self-hosted runners, honest incidents7184. Adopt the live Workers once, from a laptop: deploy everything with the
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow719 tool so each version records its commit (`scripts/deploy.sh`, or
720 `node scripts/deploy.mjs deploy --all`). Until then every plan says "no
721 known commit" and deploys every unit. To see what has changed since a
722 commit you know production runs, without deploying:
723 `node scripts/deploy.mjs plan --since <sha>`.
Fast pages, required checks on the branch, self-hosted runners, honest incidents7245. Run the workflow by hand with `dry_run`, then with `units: pages`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow725
726## First deploy of a new account
727
728What the configs refer to must exist first. `node scripts/deploy.mjs
729manifest --json` lists, per unit, its queues, KV, R2, Vectorize and
730dispatch namespaces; each unit's `setup` and `secrets` say the rest.
731
732- D1: `npx wrangler d1 create <database>`, then put its id in the unit's
733 `wrangler.jsonc`.
734- KV: `npx wrangler kv namespace create <name>` for each name under
735 `resources.kv`, then the ids in the configs.
736- Queues: `npx wrangler queues create <queue>` for each queue in the
737 manifest: `g1t-events`, `g1t-events-<service>` for every subscriber,
Merge branch 'worktree-agent-ad8a36dfcd4176015' into spend-guardrails738 `g1t-search-jobs`, `g1t-context-jobs`, and the dead-letter queue
739 `g1t-events-dlq` (below).
Merge branch 'worktree-agent-a1b995daa94e4e1b7'740- R2: `npx wrangler r2 bucket create g1t-screenshots`,
Actions: OIDC tokens, the toolkit's cache and artifact services, and artifacts in R2741 `npx wrangler r2 bucket create g1t-actions-cache` (with its two
742 lifecycle rules, above), and `npx wrangler r2 bucket create g1t-git-packs`,
Merge branch 'worktree-agent-a1b995daa94e4e1b7'743 the clone pack cache (`services/repos/src/pack_cache.rs`), with a rule
744 deleting packs 7 days after they were written and unfinished uploads
745 after a day:
746 `npx wrangler r2 bucket lifecycle add g1t-git-packs expire-packs packs/ --expire-days 7 --abort-multipart-days 1`.
747 A Worker bound to a bucket that does not exist fails to deploy, so make
748 it before the first deploy of `g1t-repos` that binds it.
Fast pages, required checks on the branch, self-hosted runners, honest incidents749- The runner's base image: `node scripts/deploy.mjs build-base`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow750- Vectorize, dispatch namespace, DNS, Access, Email Sending, Artifacts: each
751 unit's `setup`.
752- Secrets: `npx wrangler secret put <NAME>` in the unit's folder;
753 `node scripts/deploy.mjs doctor` lists what is missing.
754
755Then `node scripts/deploy.mjs deploy --all`. A Worker bound to a service
756that does not exist yet may be refused; deploy that service first with
757`--only`.
758
Merge branch 'worktree-agent-ad8a36dfcd4176015' into spend-guardrails759## The dead-letter queue
760
761Every queue consumer (`g1t-events`, each `g1t-events-<service>`,
762`g1t-search-jobs`, `g1t-context-jobs`) names `max_retries` and the
763dead-letter queue `g1t-events-dlq`, so a message that keeps failing stops
764after its retries instead of being retried for ever. A consumer bound to a
765queue that does not exist fails to deploy, so create it once per account,
766before the first deploy that names it:
767
768```sh
769npx wrangler queues create g1t-events-dlq
770```
771
772Nothing consumes it: read what landed there with `npx wrangler queues
773info g1t-events-dlq`, fix the cause, and replay by hand if needed. The
774manifest lists each unit's dead-letter queues under `queues.dead_letter`.
775
Actions: OIDC tokens, the toolkit's cache and artifact services, and artifacts in R2776## OIDC tokens for workflow jobs
777
778The API is the issuer of workflow jobs' OIDC tokens,
779`https://api.g1t.sh/actions/oidc` (`apps/api/src/oidc.rs`): no host or DNS
780of its own. It signs with an RSA key kept as the API's secret
781`ACTIONS_OIDC_KEY`. Without it, the issuer's addresses answer 404 and jobs
782are not told where to ask for a token.
783
784To turn it on, make a key on a trusted machine and store it, then delete
785the file:
786
787```sh
788openssl genpkey -algorithm RSA -pkeyopt rsa_keygen_bits:2048 -out oidc.pem
789cd apps/api && npx wrangler secret put ACTIONS_OIDC_KEY < ../../oidc.pem
790rm ../../oidc.pem
791```
792
793Check `https://api.g1t.sh/actions/oidc/.well-known/jwks` lists one key.
794Its `kid` is the key's RFC 7638 thumbprint.
795
796To rotate it, keep the old key published while tokens it signed can still
797be presented (they last 5 minutes; relying parties cache keys for longer):
798
7991. Store the current key as `ACTIONS_OIDC_KEY_PREVIOUS` (the same PEM).
8002. Make a new key and store it as `ACTIONS_OIDC_KEY`. The JWKS now lists
801 both; new tokens are signed with the new one.
8023. A day later, `npx wrangler secret delete ACTIONS_OIDC_KEY_PREVIOUS`.
803
804A key that may have leaked is rotated the same way, skipping the first
805step, so that tokens it signed stop verifying at once.
806
Deploying: a deploy by hand shows on the repository's Deployments page807## Deployments on g1t
808
809Every deploy shows on the repository's Deployments page, so its production
810card says which commit runs.
811
812- **From the workflow**, the deploy jobs name `environment: {name:
813 production, url: https://g1t.sh}`, and g1t Actions records one production
814 deployment per run.
815- **By hand**, `deploy` reports one itself: in progress once migrations are
816 in, then success or failure. It needs a g1t token with `deployments:write`
817 in `G1T_DEPLOY_TOKEN` or the file `.credentials/g1t-deploy-token`; without
818 one, or from a dirty tree or a dry run, it sends nothing. It finds the
819 repository from the git remote on g1t.sh. A report that fails is one line
820 in the log and never fails the deploy.
821
Merge branch 'main' into actions-toolkit-oidc-artifacts822When CI cannot finish a deploy, for example a runner image that will not
823build there, deploy from a machine with Docker.
Deploying: a deploy by hand shows on the repository's Deployments page824The report records it, and production shows the commit that really runs.
825
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow826## Rolling back
827
828- **One unit, at once:** `npx wrangler rollback` in its folder (or
829 `npx wrangler rollback <version-id>`; `npx wrangler versions list` shows
830 each version's commit in its message). Code only: D1 migrations are not
831 undone. The next plan sees the older commit and deploys what changed since,
832 so revert the commit on `main` too, or the next push brings it back.
833- **To a commit:** check it out and `node scripts/deploy.mjs deploy --only
Deploying: never roll production back by accident834 <units> --rollback`. Without `--rollback` the tool refuses any unit whose
835 live commit is newer than the one checked out, even with `--force`, so a
836 re-run of an old workflow run (or an old checkout) never rolls production
837 back by accident. Migrations never run backwards: a migration that needs
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow838 undoing is a new migration.
839- **The runner's image:** a rollback of the Worker does not roll back the
Deploying: never roll production back by accident840 container image. Redeploy the older commit (`--only runner --rollback`): its
Fast pages, required checks on the branch, self-hosted runners, honest incidents841 image's tag is the hash of that commit's source, which is still in the
842 registry, so nothing is built. A bad base is undone by reverting the
843 commit that changed `services/runner/base.json`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow844
845## Adding a unit
846
8471. Make its folder with a `wrangler.jsonc` (and its D1 migrations, if any).
848 A Rust Worker is a workspace member in the root `Cargo.toml` with
849 `"build": { "command": "node ../../scripts/build-rust-worker.mjs" }` and
850 the `wasm-opt = ["-O1"]` metadata; a TypeScript one is an npm workspace.
8512. Add it to `deploy/stack.jsonc`: path, kind, worker, `d1`, stage (the
852 earliest stage after everything it binds to), secrets, setup, self_host.
853 Name any new KV id under `resources.kv`.
8543. If its sources import a file outside its folder that is not a workspace
855 crate or package, list it under `inputs`.
8564. Add a row to the service table in `docs/SELF_HOSTING.md`.
8575. `npm run test:deploy` and `node scripts/deploy.mjs manifest --check`
858 say what is missing. Then create its resources and secrets, and
859 `node scripts/deploy.mjs deploy --only <unit>`.
Fast pages, required checks on the branch, self-hosted runners, honest incidents860
861## The self-hosted runner
862
863`g1t-runner` (crates/runner) is also what customers run on their own
864machines (guide: `apps/docs/src/content/docs/guides/self-hosted-runners.md`).
865It is not deployed with the stack: it is released, and runners already out
866there update themselves to each release.
867
868| Piece | Where |
869| --- | --- |
870| The tool | `scripts/runner-release.mjs` (`keygen`, `build`, `sign`, `verify`, `publish`) |
871| The workflow | `.g1t/workflows/runner-release.yml`, on a tag `runner-v<version>` or by hand |
872| Where it is published | The R2 bucket `g1t-downloads`, served by the site at `g1t.sh/downloads/runner/<version>/<file>` and `/latest/<file>` (`apps/web/app/routes/downloads-runner.ts`) |
The runner's image is g1t.sh/flagon-io/g1t-runner, on g1t's own registry, public873| Its image | `deploy/runner/Dockerfile`, pushed to `g1t.sh/flagon-io/g1t-runner` (public) for amd64 and arm64 |
Fast pages, required checks on the branch, self-hosted runners, honest incidents874
875A release is five binaries (Linux x64 and arm64, both static musl; macOS
876x64 and arm64; Windows x64), `SHA256SUMS`, and `manifest.json`;
877`latest.json` and its Ed25519 signature `latest.json.sig` name the newest.
878A runner updates only to a release whose signature checks out against the
879public key built into it and whose download matches the manifest's SHA-256.
880
881The first time:
882
8831. `node scripts/runner-release.mjs keygen`. Put `RUNNER_RELEASE_KEY` in the
884 repository's secrets (production environment) and keep a copy offline;
g1t-runner 0.1.0 is released: signed binaries for five platforms at g1t.sh/downloads/runner; the Artifacts checks' results885 put the public key in its variables as `RUNNER_RELEASE_PUBLIC_KEY` (names
886 starting `G1T_` are reserved; the workflow hands it to the build as
887 `G1T_RUNNER_RELEASE_KEY`). A build made without the public key never
888 updates itself.
Fast pages, required checks on the branch, self-hosted runners, honest incidents8892. `npx wrangler r2 bucket create g1t-downloads`, and deploy the site so it
890 has the `DOWNLOADS` binding.
The runner's image is g1t.sh/flagon-io/g1t-runner, on g1t's own registry, public8913. The image job pushes to g1t's own registry with the run's `G1T_TOKEN`;
892 the `g1t-runner` package in flagon-io is public. Set the variable
893 `RUNNER_AGENT_IMAGE` (a public copy of `g1t-runner-base`, the image agent
894 work runs in on customers' runners) once there is one.
Fast pages, required checks on the branch, self-hosted runners, honest incidents8954. Register a self-hosted runner with the `docker` label for the image job.
896
897Each release:
898
8991. Bump `version` in `crates/runner/Cargo.toml` and merge it.
9002. Tag the commit `runner-v<version>` and push the tag. The workflow builds
901 every platform with cargo-zigbuild, signs, verifies, publishes the files
902 (the version's first, `latest.json` last), and pushes the image.
9033. By hand, the same is `node scripts/runner-release.mjs build`, then `sign`,
904 `verify` and `publish`, with the keys in the environment.
905
906Rolling back a release: copy the older version's `manifest.json` over
907`runner/latest.json` and sign it again (`sign` after checking the older
908version out). Runners never move to an older version on their own; a
909runner on a bad release is fixed by the next good one, or by downloading
910the older binary over it.

This file's history is long; its oldest lines are credited to the oldest commit read.