Skip to content
768 linesCodeBlameRaw
1# Deploying g1t
2
3How g1t.sh gets to Cloudflare: one manifest that lists every deployable
4unit, one tool that deploys only what changed, and a g1t Actions workflow
5that runs that tool on every push to `main`. Internal: the public
6self-hosting guide is `apps/docs/src/content/docs/guides/self-hosting.md`.
7
8| Piece | Where |
9| --- | --- |
10| The manifest | `deploy/stack.jsonc` |
11| The tool | `scripts/deploy.mjs` (library and tests in `scripts/deploy/`) |
12| Rust Worker builds | `scripts/build-rust-worker.mjs`, every Rust unit's build command |
13| The runner's images | `services/runner/base/Dockerfile`, `services/runner/Dockerfile`, `services/runner/base.json`, `scripts/build-runner.mjs`, `scripts/deploy/image.mjs` |
14| The workflows | `.g1t/workflows/deploy.yml`, `.g1t/workflows/runner-base.yml` |
15| The old entry point | `scripts/deploy.sh`, now a wrapper |
16
17## The manifest
18
19`deploy/stack.jsonc` names every unit: each folder with a `wrangler.jsonc`.
20It is JSONC rather than TOML so that Node reads it with no dependency,
21with the same parser as the Wrangler configs, and comments stay possible.
22
23| Field | |
24| --- | --- |
25| `path` | The unit's folder. |
26| `kind` | `rust-worker` (built by worker-build), `ts-worker` (Wrangler bundles it), `react-router` (`vite build` first), `astro` (`astro build` first). |
27| `worker` | The Worker's name. Must match its `wrangler.jsonc`. |
28| `d1` | `{ database, migrations }`, when it has a database. Must match its `wrangler.jsonc`. |
29| `stage` | `core`, `edge` or `front` (below). |
30| `secrets` | The Wrangler secrets it needs, by name. `node scripts/deploy.mjs doctor` checks they are set. |
31| `setup` | One-time steps no config can say, for a first deploy. |
32| `inputs` | Files outside its folder it is built from that no workspace metadata names. A test finds such imports. |
33| `image` | A Containers image: `dockerfile` (the image deployed), `crate` (the binary it adds), `base` (`{ context, lock }`: the base image's folder and the file recording the base that was pushed) and `repository` (where both are pushed). See [the runner's images](#the-runners-images). |
34| `self_host` | `run`, `off`, `separate` or `none`: what `deploy/self-host/configs.mjs` does with it. |
35
36What a unit is **built from** is never listed by hand. The tool reads it:
37Rust path dependencies from `cargo metadata`, workspace packages from each
38`package.json`, closed transitively. `node scripts/deploy.mjs manifest`
39prints the result:
40
41```
42unit stage kind worker d1 built from (besides its folder)
43events core rust-worker g1t-events g1t-events crates/contracts crates/kit
44repos core rust-worker g1t-repos g1t-repos crates/contracts crates/kit crates/scan crates/secrets
45runner core ts-worker g1t-runner crates/actions crates/runner packages/contracts
46og core ts-worker g1t-og packages/contracts
47web front react-router g1t packages/contracts packages/theme
48...
49```
50
51Root files count too: `Cargo.toml`, `Cargo.lock` and
52`scripts/build-rust-worker.mjs` for Rust units; `package-lock.json` and
53`tsconfig.base.json` for the others; `package.json` for all. A lockfile
54change counts for a unit only if a package in that unit's graph changed,
55read from the lockfile itself (`scripts/deploy/lockfiles.mjs`), so bumping
56`sharp` for the docs does not redeploy the Rust services.
57
58### Stages
59
60| Stage | What | Units |
61| --- | --- | --- |
62| `migrations` | Every pending D1 migration, in parallel, before any code | each unit's `d1` |
63| `core` | Services reached through bindings | the services, `og` |
64| `edge` | Public endpoints other than the site | `api`, `models`, `pages`, `status` |
65| `front` | The site, sudo, the docs | `web`, `sudo`, `docs` |
66
67A stage starts only when the one before it succeeded. Inside a stage units
68deploy in parallel. The rule the tests enforce: **a unit binds only to units
69in its own stage or an earlier one**, so new code never calls a service
70that has not shipped. (Services in `core` bind to each other in cycles,
71which is why they share a stage.)
72
73### What reads the manifest
74
75- `scripts/deploy.mjs`: everything below.
76- `deploy/self-host/configs.mjs`: which Workers a self-hosted installation
77 runs (`self_host: "run"`, the site first) and which are bound to the off
78 Worker (`"off"`). Its output is identical to before, apart from the order
79 of `workers.txt` after the site.
80- Tests (`npm run test:deploy`) check that: every `wrangler.jsonc` in the
81 repository has a unit; `worker`, `d1` and `image` match the configs;
82 every KV id is named under `resources.kv`; stages follow bindings; the
83 derived dependencies agree with Cargo's own resolved graph; every import
84 that leaves a unit's folder is covered; `deploy/self-host/Dockerfile`
85 builds exactly the Rust units self-hosting runs; and the service table in
86 `docs/SELF_HOSTING.md` names every unit.
87
88## The tool
89
90```sh
91node scripts/deploy.mjs plan # what would deploy, and why (read-only)
92node scripts/deploy.mjs deploy # migrations, then every changed unit
93node scripts/deploy.mjs deploy --only web,api # just these, if they changed
94node scripts/deploy.mjs deploy --only web --force # just this, changed or not
95node scripts/deploy.mjs deploy --all # everything
96node scripts/deploy.mjs build --only events # build as a deploy would; upload nothing
97node scripts/deploy.mjs migrate # pending migrations only
98node scripts/deploy.mjs manifest [--check|--json]
99node scripts/deploy.mjs doctor # secrets each unit lacks
100node scripts/deploy.mjs build-base # build and push the runner's base image (Docker)
101node scripts/deploy.mjs image # build and push the runner's image for this checkout (Docker)
102scripts/deploy.sh [units...] # the old entry point: all, or those named, always
103```
104
105| Flag | |
106| --- | --- |
107| `--only a,b` / `--skip a,b` | Units by short name, folder or Worker name. |
108| `--all` | Every unit, changed or not. |
109| `--force` | Deploy the selected units even if unchanged. |
110| `--concurrency N` | Units at once inside a stage, and migrations at once (default 4). |
111| `--stage core` | One stage only. |
112| `--no-migrations` | Skip the migrations step (the workflow runs it as its own job). |
113| `--allow-dirty` | Deploy with uncommitted changes in what deploys. The version records no commit, so the next plan deploys it again. |
114| `--rebuild-image` | Build the runner's image even if one for this source is already in the registry. |
115| `--rebuild-base` | Build and push a new base first (needs Docker), then the runner's image on it. Writes `services/runner/base.json`: commit it. |
116| `--no-push` | `build-base` and `image`: build locally, push nothing. |
117| `--no-cache` | `build-base`: build every layer again. |
118| `--since REV` | Treat Workers with no recorded commit as running `REV`. Used once to adopt Workers deployed before this tool. |
119| `--json`, `--out FILE`, `--github-output` | The plan as data, for the workflow. |
120
121### Where the deployed commit is kept
122
123On the Worker itself. Every deploy runs `wrangler deploy --message
124"g1t-deploy <40-char sha> <subject>" --tag g1t-<12-char sha>`, which
125Cloudflare keeps as the version's `workers/message` and `workers/tag`
126annotations. `plan` reads them back with `wrangler deployments status`
127(the live version) and `wrangler versions list` (its annotations): two
128read-only calls per unit, in parallel; a plan of all 22 units takes about
12910 seconds. No KV namespace or other infrastructure is needed.
130
131- A version made by `wrangler secret put` keeps the code of the one before
132 it, so the tool looks through those to the deploy before.
133- A version deployed any other way (by hand, from the dashboard) has no
134 commit, and the unit is deployed again.
135- During a gradual rollout the version with the most traffic counts.
136- After `wrangler rollback`, the plan sees the older commit and deploys
137 what changed since.
138
139### What a deploy does
140
1411. Plans: for each unit, the live commit; `git diff` from it to `HEAD`;
142 whether the changed files touch the unit (its folder, the crates and
143 packages it is built from, its inputs, lockfile changes that reach it).
1442. Refuses if uncommitted changes touch what would deploy (`--allow-dirty`).
1453. Applies every pending migration (`wrangler d1 migrations apply --remote`),
146 in parallel. Any failure stops the deploy before code.
1474. Installs worker-build once if any Rust unit is deploying.
1485. Each stage in turn: units in parallel (`--concurrency`), each `npm run
149 build` first for React Router and Astro, then `wrangler deploy` in the
150 unit's folder with the annotation. If any unit fails, later stages are
151 not started.
1526. Prints a table: unit, stage, result, version id, time. Each unit's full
153 output is kept in `$TMPDIR/g1t-deploy/<unit>.log`.
154
155#### Migrations run while the old code is still live
156
157Migrations apply before any code, and a stage takes a minute or more, so
158for that long the code in production is the old code reading the new
159schema. A migration must keep the old code right:
160
161- **Add, never change meaning.** New tables and columns (with defaults) are
162 safe. Rewriting what existing rows mean is not: on 2026-10-08,
163 `billing/0039` turned comped accounts into `custom` terms at 100%, and
164 for about 40 seconds the old billing code, which knew only `comped`, saw
165 flagon-io as a free workspace and refused its workflows.
166- **Change meaning in two deploys.** First ship code that reads both the old
167 and the new form (and keeps writing the old one); then, in a later
168 deploy, the migration that rewrites the rows; then, if you like, code
169 that drops the old form.
170- **Never drop or rename** a column or table the live code reads in the same
171 deploy that stops reading it.
172
173#### Telling the status page about a deploy
174
175Restarts during a deploy can make a part slow for a minute, which the
176status page's checks would otherwise draft as an incident. So
177`scripts/deploy.mjs deploy` says when it starts its stages and when they
178end (`withDeployWindow` in `scripts/deploy/status-window.mjs`; never in a
179dry run, and only once there is something to ship), by hand and in g1t
180Actions alike. It posts to `POST https://status.g1t.sh/deploys` with
181`Authorization: Bearer $STATUS_DEPLOY_TOKEN`, the same value as the status
182Worker's `STATUS_DEPLOY_TOKEN` secret, and `{"phase": "started" |
183"finished", "id": "<commit>"}`. During the deploy and for 3 minutes after
184it, detection keeps counting failed and slow checks but makes no new
185draft; trouble that outlasts that is drafted with its true start. Deploys
186that overlap (the jobs of one stage run at once, each announcing itself)
187are one window: the status Worker counts the starts, and the window closes
188when the last one finishes. A start with no finish stops counting 30
189minutes after the latest start. Without the token nothing is sent, and an
190announcement never fails a deploy: a refusal or network error is one
191warning line. By hand: `node scripts/deploy/status-window.mjs
192started|finished [id]`.
193
194To turn it on (once; until then deploys are not announced):
195
1961. Make a token and set it as the status Worker's secret:
197 `npx wrangler secret put STATUS_DEPLOY_TOKEN` in `apps/status` (as of
198 2026-10-08 the Worker has only `STATUS_SECRET`). Without it,
199 `POST /deploys` answers 404.
2002. Set the same value as the **`STATUS_DEPLOY_TOKEN`** Actions secret on
201 flagon-io/g1t (Settings, Secrets and variables, or
202 `PUT /repos/flagon-io/g1t/actions/secrets/STATUS_DEPLOY_TOKEN`);
203 `.g1t/workflows/deploy.yml` passes it to every deploy job.
2043. Add `status.g1t.sh | deploy.yml | production` to the project's
205 **Workflow-only domains** (see Network below), or the job's request is
206 refused by the guardrails (the deploy still goes on, with a warning).
2074. For deploys by hand, set `STATUS_DEPLOY_TOKEN` in your shell's
208 environment (the tool does not read `.env`).
209
210On a laptop the tool uses your `wrangler login` (or `CLOUDFLARE_DEPLOY_TOKEN`
211if set), as `scripts/deploy.sh` always did: a `CLOUDFLARE_API_TOKEN` or
212global API key in your shell, or in the repository's `.env`, is ignored.
213With `CI=true` it uses `CLOUDFLARE_API_TOKEN`.
214
215### The runner's images
216
217`services/runner` runs every sandbox (agents, checks, the merge queue,
218workflow jobs, g1t.page builds) from one Containers image, made in two
219parts:
220
221| Image | Built from | Holds | Rebuilt |
222| --- | --- | --- | --- |
223| **Base**, `g1t-runner:base-<date>-<inputs>` | `services/runner/base/Dockerfile` | Debian bookworm, Node 24, Python 3.11, Docker (Engine, Buildx, Compose, from Docker's apt repository), Go (from go.dev), Rust stable for the `node` user with rustfmt, clippy and the `wasm32-unknown-unknown` target, build-essential, musl-tools, git, ripgrep, jq, zstd, sudo, and the pinned Claude Code CLI on top | When its folder changes, weekly, or by hand (`build-base`) |
224| **Runner**, `g1t-runner:<content hash>` | `services/runner/Dockerfile`: `FROM` the base, plus one file | The g1t runner, a static binary | When the binary or the base changes |
225
226Both are pushed to one repository of Cloudflare's registry,
227`registry.cloudflare.com/<account>/g1t-runner`, so pushing the runner's
228image uploads only its own layer (about 5 MB): the base's layers are
229already there.
230
231Both are built as Wrangler builds images (`--platform linux/amd64
232--provenance=false --sbom=false`): one manifest, not an OCI index with a
233BuildKit attestation beside it. A push is tried up to three times and counts
234only when Docker reports the digest; the first push of the base once failed
235with `blob unknown to registry` and went through when run again (see
236`docs/CLOUDFLARE_FEEDBACK.md`, C3). `build-base` and `image` exit non-zero
237when a push fails, and `base.json` is written only after the push.
238
239**The base** is recorded in `services/runner/base.json`: its reference, its
240digest, a hash of its folder (`inputs`), when it was built, its size and
241each toolchain's version. `node scripts/deploy.mjs build-base` builds it,
242pushes it and rewrites the file; commit the file, and the next deploy
243builds the runner's image on it. `npm run test:deploy` fails while the
244folder and the file disagree, so a change to the base's Dockerfile cannot
245merge without the base it describes. Layers go from what changes least to
246most (system packages, Go, Rust, Java, .NET and Ruby, the Claude Code
247CLI), and the apt and npm caches stay in BuildKit's cache, out of the
248image. Ruby is compiled from source in a single step (a few minutes on a
249cold cache) that removes its source tree before the layer is written.
250
251The base's build cache is the base itself: it is built with
252`BUILDKIT_INLINE_CACHE`, which records in the image how each layer was
253made, and `build-base` builds `--cache-from` the base it replaces. A
254machine with an empty cache, or one just pruned, pulls the unchanged
255layers from the registry instead of building them. (BuildKit's other
256registry cache, `--cache-to type=registry`, pushes a separate cache
257manifest that not every registry takes, and for a one-stage image adds
258nothing the inline cache lacks.)
259
260**The runner binary** (`crates/runner`) is built outside Docker by
261`scripts/build-runner.mjs` as one static binary for
262`x86_64-unknown-linux-musl`, so it runs on the base whatever its libc, and
263on a self-hosted runner's machine too. Where it is built:
264
265- on x86-64 Linux with the musl target and `musl-gcc`
266 (`rustup target add x86_64-unknown-linux-musl`, `apt-get install musl-tools`),
267 with the machine's own Cargo;
268- anywhere else (Windows, macOS) in a small builder container, Rust on
269 Alpine (whose own target is musl), with Docker volumes keeping Cargo's
270 registry and target directory between builds. Windows has no musl
271 cross-linker, and `ring` (under ureq's TLS) needs a C compiler for the
272 target, so a container is the dependable route.
273
274**The runner's image tag** is a hash of everything it is built from: its
275Dockerfile, `base.json`, the crates the binary is built from, the
276workspace's Cargo files and the build script. The same source always names
277the same image, so:
278
2791. A deploy computes the tag and asks the registry whether it is there
280 (a `HEAD` of its manifest, with credentials from Wrangler; no Docker).
2812. If it is, nothing is built: the deploy uses it.
2823. If not, it builds the binary and the image (seconds on a warm machine)
283 and pushes it. In `deploy.yml` that is the `runner-image` job, on
284 `g1t-4core`, with the job's own Docker Engine (see
285 [Docker in workflow jobs](#docker-in-workflow-jobs)): it adds the musl
286 target, builds the binary natively (the base has `musl-gcc`), pulls the
287 base from Cloudflare's registry, builds, and pushes one layer.
2884. If Docker does not answer (a machine without it, or jobs with Docker
289 turned off), the unit fails saying to run `node scripts/deploy.mjs
290 image` on a machine with Docker; then re-run the workflow.
291
292Then `wrangler deploy` is given the image by reference, from a generated
293config (`services/runner/wrangler.deploy.json`, deleted after, ignored by
294git), so Wrangler builds nothing. When nothing the image is built from
295changed since the runner's live commit, the deploy also passes
296`--containers-rollout none`, which leaves running sandboxes alone.
297
298To get a new base out:
299
300```sh
301node scripts/deploy.mjs build-base # build, push, write base.json (needs Docker)
302git commit services/runner/base.json -m "A new base image for g1t's sandboxes"
303node scripts/deploy.mjs image # optional: push the runner's image now, so CI finds it
304```
305
306`.g1t/workflows/runner-base.yml` does the same weekly (and when the
307base's folder changes on `main`), and opens a pull request with
308`base.json`. It runs on a self-hosted runner with the `docker` label
309(`runs-on: [self-hosted, docker]`): g1t's own machines have Docker now,
310but the base's build downloads from Docker's apt repository over HTTPS,
311which does not trust a guarded job's egress certificate, so it stays on an
312open network. Its job is skipped until the repository variable
313`RUNNER_BASE_SELF_HOSTED` is `true` (set it once a runner is registered;
314without one, runs waited in the queue for a day and a half and then
315failed); until then run `build-base` by hand.
316
317**Sandboxes start from the image.** Cloudflare pulls an image to a
318machine the first time a sandbox lands there, and keeps it. A smaller base
319pulls sooner, and a change to the runner alone sends machines one 5 MB
320layer instead of the whole image.
321
322#### Larger machines
323
324The same image runs on three instance types, each a Durable Object class
325of its own in `services/runner/wrangler.jsonc`: `AttemptSandbox`
326(`standard-1`), `Sandbox2Core` (`standard-3`) and `Sandbox4Core`
327(`standard-4`). Workflow jobs choose with `runs-on: g1t-2core` or
328`g1t-4core` (`g1t_contracts::actions::INSTANCE_TYPES`); the actions
329service passes the label to the runner, which starts the job in that
330class. Billing prices the larger ones from their memory and disk, and
331their CPU (see the public billing guide). The account's Containers limits
332must allow `standard-4`; Wrangler refuses the deploy otherwise.
333
334#### Docker in workflow jobs
335
336Workflow jobs on g1t's machines have a Docker Engine of their own
337(`crates/runner/src/docker/`; the public guide is
338`apps/docs/src/content/docs/guides/actions.md`, "Docker"). What Cloudflare
339Containers allow decides how it runs (findings in `docs/PLAN.md`,
340"Docker in workflow jobs"):
341
342| Piece | What it does |
343| --- | --- |
344| `dockerd` | Started as root with `sudo`, only when the job first uses Docker or has `services:` or `container:`. Flags: `--iptables=false --ip6tables=false --ip-forward=false` (Containers allow neither), the containerd image store, Docker Hub through `mirror.gcr.io`. Its config, socket and log are in `/run/g1t-docker` (`dockerd.log` is the place to look); its data in `/var/lib/docker`, on overlays when the disk takes them, plain copies (`native`) when not. |
345| `/var/run/docker.sock` | The runner's API proxy (`docker/api.rs`), which starts the Engine on the first connection. Containers that ask for a bridge network get the job's own (`host`), the names they would have had resolve to 127.0.0.1, and ports published under another number are forwarded. |
346| `runc` | The Engine finds the runner binary first on its `PATH` as `runc` (`docker/oci.rs`): BuildKit's `RUN` steps join the job's network, and in a guarded job every container gets the egress certificate at `/dev/g1t-egress`. Then the real `/usr/bin/runc` runs. |
347| cgroups | Before the Engine starts, the sandbox's processes move to a cgroup of their own and every controller is handed down, as Docker's own Docker-in-Docker image does, so `--cpus` and `--memory` work. |
348
349- **Turning it off:** set the runner Worker's `DOCKER` var to `off`
350 (`services/runner/wrangler.jsonc`) and deploy the runner: new jobs get
351 no Engine (`G1T_DOCKER=off`), and jobs that need one fail saying Docker
352 does not answer. Anything else is on.
353- **Network:** containers share the job's network, so the guardrails, the
354 workflow-only domains and the egress Worker apply to them unchanged. The
355 public registries are in `BUILD_HOSTS` (`services/runner/src/egress.ts`).
356- **Isolation:** one Engine per job, inside the job's sandbox (its own
357 VM), gone with it. The job already had root through `sudo`; Docker adds
358 no reach beyond the sandbox, and no host socket is ever mounted into one.
359- **Checked locally** (2026-10-08, Docker Desktop, the base and runner
360 image built from this tree, a privileged container standing in for a
361 sandbox, a pretend API): services with health checks, `localhost` and
362 names, port forwarding, `docker build` with a networked `RUN`, Compose
363 with a healthy dependency, `docker://` steps, a Dockerfile action, a
364 `container:` job with a JavaScript action, an Alpine job container, the
365 egress certificate in `run`, `exec` and build steps and in no layer,
366 plain-copy storage, and the deploy's own build and push to a registry.
367 Not yet seen on Cloudflare itself: watch the first runs' logs for the
368 `Docker: started` line, and `dockerd.log` if it does not come.
369
370## Build speed
371
372Measured on the development machine (Windows, 32 cores, warm Cargo cache),
373building `events`, `search` and `repos` after a change to `crates/kit`, as a
374deploy does but without uploading (`wrangler deploy --dry-run`):
375
376| | Time |
377| --- | --- |
378| Before: one after another, `cargo install worker-build` each time, wasm-opt `-O` | 81 s, 84 s |
379| Concurrent builds, wasm-opt `-O` | 50 s, 68 s |
380| Concurrent builds, wasm-opt `-O1` (now) | 28 s, 33 s |
381
382Where the time went, and what changed:
383
384- **wasm-opt** was most of it: `-O` took 38 s on `repos` and 19 s on `api`;
385 `-O1` takes 1 to 3 s. With worker-build's flags (it keeps the names
386 section) the `.wasm` is 24 to 28% larger raw but only 1 to 7% larger
387 gzipped, and Workers' size limit is on the compressed upload. `-Os` and
388 `-Oz` were no faster than `-O`. Set per crate in
389 `[package.metadata.wasm-pack.profile.release]`; a test keeps every Rust
390 unit on the same level.
391- **worker-build** is installed only when missing or another version
392 (`scripts/build-rust-worker.mjs`, which pins it). Locally `cargo install`
393 on an installed version cost under a second; on a fresh CI sandbox it is
394 a full compile, which the workflow caches instead.
395- **One Cargo target**: every Rust unit is a member of the workspace, so
396 they already share `target/`. Concurrent builds take turns on Cargo's
397 lock for the compile, and their wasm-bindgen, wasm-opt and uploads
398 overlap.
399- **No joint `cargo build -p a -p b`**: tried, and it is slower. Cargo
400 unifies features across the packages of one build (`serde_json`'s
401 `preserve_order` from `api` and `actions`, `digest` features from
402 `secrets`), so each worker-build afterwards compiled its own variant
403 again.
404- **The runner's images**, measured on the same machine on 2026-10-06
405 (Docker Desktop, 8 vCPUs; its disk was busy with other containers, so
406 the cold figures are slow and noisy):
407
408 | | Before (one image) | Now |
409 | --- | --- | --- |
410 | Size, unpacked / compressed (what a machine pulls) | 3.08 GB / 819 MB | 2.68 GB / 686 MB |
411 | A change to the runner | Docker rebuilds the image's Rust stage and pushes the image (1198 s in the first deploy after a prune) | binary 46–53 s (6 s unchanged), image 7 s, push one 5 MB layer (1.4 s to a local registry) |
412 | The base from nothing | 431 s (whole image, cold) | 758 s cold, rarely: weekly or when its folder changes |
413 | The base after `docker builder prune` | as from nothing | 68 s, its layers pulled from the registry it was pushed to |
414 | The runner binary, cold (builder container) | | 99 s |
415 | A Rust CI job's build (events, search, repos; 4 vCPUs) | | 51 s cold, 13 s with the Cargo target restored (107 MB zstd entry) |
416
417- **Only what changed** is the largest saving: a change to one service
418 deploys one service.
419
420## The workflow
421
422`.g1t/workflows/deploy.yml` runs on every push to `main`, and by hand
423(**Actions → Deploy → Run workflow**) with `units` (deploy these, changed or
424not), `all` and `dry_run` (plan only).
425
426| Job | Does | Needs |
427| --- | --- | --- |
428| `check` | `manifest --check` and `npm run test:deploy` | — |
429| `plan` | `plan --github-output`: outputs per stage, the plan in the run's summary | `check` |
430| `migrate` | `migrate --only <units with pending migrations>` | `plan`; skipped when none are pending |
431| `core`, `edge`, `front` | `deploy --only <units> --force --no-migrations`, one job per build group | the stages before; skipped when empty |
432| `smoke` | `node scripts/ops/smoke.mjs`: the landing page, sign-in, sign-up and pricing load, and the waitlist form reaches identity (sent an address identity refuses before keeping or counting anything, so the real waitlist is never touched) | every stage; skipped when nothing deployed |
433
434- **One at a time:** `concurrency: deploy-production`, never cancelled in
435 progress; a second push waits.
436- **Build groups:** a stage's units are split so each job shares a build:
437 Rust workers at most four to a job (each a 4-vCPU `g1t-4core` machine), the
438 TypeScript Workers together, each site alone, and a unit whose image must
439 be rebuilt alone (`image: true` in the matrix, also on `g1t-4core`, where
440 it builds and pushes the image with the job's own Docker Engine). `fail-fast: false`, so one failed job does not cut
441 another off mid-upload; the next stage then does not start.
442- **Tests:** there is no CI workflow on g1t yet; `main` is kept passing by
443 the merge queue's checks. `check` runs the deploy tool's own tests. When a
444 CI workflow is added, make `plan` wait for it (`workflow_run`, or a job in
445 this file).
446- **Machines:** Rust jobs and the runner's image run on `g1t-4core` (4 vCPUs,
447 12 GiB, 20 GB), the others on the standard machine
448 (`runs-on: ${{ (matrix.rust || matrix.image) && 'g1t-4core' || 'ubuntu-latest' }}`).
449 The image job needs the room: the base it builds on is about 3.2 GB
450 unpacked.
451- **Caching** (`actions/cache`: up to 2 GB an entry, 10 GB a repository,
452 kept until unused for 7 days): the worker-build binary, worker-build's
453 downloaded tools, `~/.cargo/registry/cache`, and the Cargo target's
454 release dependencies (`target/release` and
455 `target/wasm32-unknown-unknown/release`, without `incremental` or
456 `.wasm`), keyed by the build group, `Cargo.lock` and `base.json`. The
457 workspace's own crates are compiled again on every run (a checkout's
458 sources are newer than any cache); the crates.io dependencies are not.
459 npm's cache is not kept: every job runs `npm ci` of only what its units
460 need (`deploy.mjs install`: Wrangler alone for Rust jobs).
461- **Conditions:** each stage runs with `!failure() && !cancelled()`, which
462 on g1t (as on GitHub) is true when no job before it failed, however far
463 back: a `migrate` job skipped for having nothing to apply does not stop
464 the stages after it, and a failed `check` stops all of them.
465- `crates/actions/tests/repository_workflows.rs` reads the workflow with
466 g1t's own parser and expressions, and checks the jobs start, wait and
467 stop as above (`cargo test -p g1t-actions --test repository_workflows`).
468
469### What the sandbox has
470
471The base image (`services/runner/base/Dockerfile`) has Node 24, npm, git,
472Go, zstd, Docker, musl-tools, Rust stable for the `node` user with
473rustfmt, clippy and the `wasm32-unknown-unknown` target, Java (Temurin 21),
474the .NET 8 SDK and Ruby 3.3, but not worker-build or `gh`. Java and Ruby
475sit in the tool cache (`/home/runner/_tool/Java_Temurin-Hotspot_jdk/<semver
476with + as ->/x64` and `/home/runner/_tool/Ruby/<version>/x64`, each with
477an `x64.complete` marker), which is where `actions/setup-java` and
478`ruby/setup-ruby` look; .NET is in `/usr/share/dotnet`, owned by `node`,
479which is where `actions/setup-dotnet` installs. Bumping one means changing
480its version and checksum `ARG`s together (Java's tool cache name is
481Adoptium's `version_data.semver`; .NET's SHA-512 is in its release
482metadata; Ruby's SHA-256 is in `cache.ruby-lang.org/pub/ruby/index.txt`).
483The user docs list what jobs get in
484`apps/docs/src/content/docs/guides/actions.md` (The runner). The workflow's `rustup target add wasm32-unknown-unknown` is
485then a no-op, and worker-build is restored from the cache, installed on a
486miss. The image job adds `x86_64-unknown-linux-musl` (about 30 MB from
487`static.rust-lang.org`) and keeps its Cargo target in the cache. worker-build
488fetches wasm-bindgen and wasm-opt from GitHub releases and esbuild from
489npm. All of those hosts are on the list every workflow job may reach.
490
491### Network
492
493A workflow job reaches its project's allowed domains, g1t, what builds
494need (`services/runner/src/egress.ts`, `BUILD_HOSTS`), and the project's
495**workflow-only domains** that name its workflow and environment. Those are
496never reached by agents, checks, the merge queue, deploy builds or runs of
497pull requests from forks (`Guardrails::workflow_hosts`, the runner's
498`jobHosts`). Under flagon-io/g1t's **Settings → Guardrails**
499(Maintain role or higher), **Workflow-only domains**:
500
501```text
502api.cloudflare.com | deploy.yml | production
503registry.cloudflare.com | deploy.yml, runner-base.yml | production
504status.g1t.sh | deploy.yml | production
505```
506
507`api.cloudflare.com` is Wrangler's API; `registry.cloudflare.com` is where
508the deploy asks whether the runner's image is already built, and where the
509`runner-image` job pulls the base from and pushes the runner's image to
510(as `runner-base.yml` pushes the base); `status.g1t.sh` hears the deploy
511start and finish (see "Telling the status page about a deploy"). The job's Docker Engine shares the
512job's network, so these lines are what let it reach the registry. If a pull
513is refused with `g1t guardrails: <host> is not on this project's allowed
514domains`, the registry sent the layers from another host: add that host on
515the same line. Only `deploy.yml`'s and `runner-base.yml`'s
516jobs with `environment: production` reach them, which are also the only
517jobs that can read `CLOUDFLARE_API_TOKEN`. Each change to the list is in
518the workspace's audit log as `update_guardrails`.
519
520### The API token
521
522Create it at **dash.cloudflare.com → My Profile → API Tokens → Create
523Token → Custom token**, named `g1t deploys (CI)`:
524
525| Scope | Permission | Why |
526| --- | --- | --- |
527| Account | Workers Scripts: Edit | Upload, versions, deployments, crons, bindings, `secret list` (doctor) |
528| Account | D1: Edit | `d1 migrations list` and `apply` |
529| Account | Queues: Edit | Attaching each unit's queue consumers on deploy |
530| Account | Workers R2 Storage: Read | Wrangler checks `og`'s bucket binding |
531| Account | Account Settings: Read | Wrangler reads the account |
532| Account | Containers: Edit | The runner's deploy updates its applications (the image reference, the three classes), and gets registry credentials (`wrangler containers registries credentials --push`, one hour) to look for, pull and push its image. `runner-base.yml` pushes images with it. |
533| Zone (`g1t.sh`, `g1t.page`) | Workers Routes: Edit | `pages`' zone routes, and custom domains |
534| Zone (`g1t.sh`, `g1t.page`) | DNS: Edit | Custom domains (`api`, `mcp`, `og`, `models`, `status`, `sudo`, `docs`, `g1t.sh`, `g1t.page`) keep their DNS records |
535| Zone (`g1t.sh`, `g1t.page`) | Zone: Read | Finding the zone a route names |
536
537Restrict it to account `syntaqx` (`1e6f2cffa3f445920836e8ebe446bb58`) and
538the two zones. Not needed for deploys: KV (bindings are by id; creating a
539namespace is a one-time setup), Vectorize, Workers for Platforms beyond
540Workers Scripts, Cloudflare for SaaS custom hostnames (the deployments
541service does that at runtime with its own token), SSL and Certificates.
542Workers KV Storage: Edit and Vectorize: Edit are only for first-time setup,
543which stays a person's job.
544
545The list follows what our configs use; Cloudflare does not publish exactly
546what `wrangler deploy` checks for each binding. Bindings with no listed
547permission (Browser Rendering, Workers AI, Vectorize, Email Sending,
548Artifacts, dispatch namespaces) are assumed to need none beyond Workers
549Scripts. Verify on the first run with `workflow_dispatch` and `units:
550pages` (small, no secrets), then `units: og` (R2) and `units: runner`
551(Containers); a missing permission fails with `Authentication error [code:
55210000]` and the route it was refused.
553
554Add it to the repository:
555
5561. On g1t.sh, open **flagon-io/g1t → Settings → Secrets and variables**.
5572. **Add**: key `CLOUDFLARE_API_TOKEN`, type **Secret**, available to
558 **Workflows**, environment **Production**. Only jobs with
559 `environment: production` (the deploy jobs) can read it.
5603. **Add**: key `CLOUDFLARE_ACCOUNT_ID`, type **Variable**, value
561 `1e6f2cffa3f445920836e8ebe446bb58`, available to **Workflows**, all
562 environments.
563
564Or through the API:
565
566```sh
567curl -X PUT https://api.g1t.sh/repos/flagon-io/g1t/actions/secrets/CLOUDFLARE_API_TOKEN \
568 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
569 -d '{"value":"<the token>","environments":["production"],"available_to":["workflows"]}'
570
571curl -X POST https://api.g1t.sh/repos/flagon-io/g1t/actions/variables \
572 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
573 -d '{"name":"CLOUDFLARE_ACCOUNT_ID","value":"1e6f2cffa3f445920836e8ebe446bb58","available_to":["workflows"]}'
574```
575
576## Turning it on
577
5781. Create the token and add the secret and variable (above).
5792. Add `api.cloudflare.com` and `registry.cloudflare.com` to flagon-io/g1t's
580 workflow-only domains, for `deploy.yml` in `production` (above).
5813. Build and push the base once, and commit `services/runner/base.json`:
582 `node scripts/deploy.mjs build-base`. Create the cache bucket:
583 `npx wrangler r2 bucket create g1t-actions-cache`, with lifecycle rules
584 deleting cache entries (`c/`) 30 days after upload and artifacts (`a/`)
585 after 91, a day past the longest they are kept, and unfinished uploads
586 after a day (the actions service deletes both sooner; the rules catch
587 what it misses):
588 `npx wrangler r2 bucket lifecycle add g1t-actions-cache expire-cache c/ --expire-days 30 --abort-multipart-days 1`
589 and `npx wrangler r2 bucket lifecycle add g1t-actions-cache expire-artifacts a/ --expire-days 91 --abort-multipart-days 1`.
590 A bucket made before artifacts moved there has one rule for everything,
591 `expire`, which would delete artifacts kept longer than 30 days: remove
592 it (`npx wrangler r2 bucket lifecycle remove g1t-actions-cache --id expire`)
593 and add the two above.
5944. Adopt the live Workers once, from a laptop: deploy everything with the
595 tool so each version records its commit (`scripts/deploy.sh`, or
596 `node scripts/deploy.mjs deploy --all`). Until then every plan says "no
597 known commit" and deploys every unit. To see what has changed since a
598 commit you know production runs, without deploying:
599 `node scripts/deploy.mjs plan --since <sha>`.
6005. Run the workflow by hand with `dry_run`, then with `units: pages`.
601
602## First deploy of a new account
603
604What the configs refer to must exist first. `node scripts/deploy.mjs
605manifest --json` lists, per unit, its queues, KV, R2, Vectorize and
606dispatch namespaces; each unit's `setup` and `secrets` say the rest.
607
608- D1: `npx wrangler d1 create <database>`, then put its id in the unit's
609 `wrangler.jsonc`.
610- KV: `npx wrangler kv namespace create <name>` for each name under
611 `resources.kv`, then the ids in the configs.
612- Queues: `npx wrangler queues create <queue>` for each queue in the
613 manifest: `g1t-events`, `g1t-events-<service>` for every subscriber,
614 `g1t-search-jobs`, `g1t-context-jobs`.
615- R2: `npx wrangler r2 bucket create g1t-screenshots`,
616 `npx wrangler r2 bucket create g1t-actions-cache` (with its two
617 lifecycle rules, above), and `npx wrangler r2 bucket create g1t-git-packs`,
618 the clone pack cache (`services/repos/src/pack_cache.rs`), with a rule
619 deleting packs 7 days after they were written and unfinished uploads
620 after a day:
621 `npx wrangler r2 bucket lifecycle add g1t-git-packs expire-packs packs/ --expire-days 7 --abort-multipart-days 1`.
622 A Worker bound to a bucket that does not exist fails to deploy, so make
623 it before the first deploy of `g1t-repos` that binds it.
624- The runner's base image: `node scripts/deploy.mjs build-base`.
625- Vectorize, dispatch namespace, DNS, Access, Email Sending, Artifacts: each
626 unit's `setup`.
627- Secrets: `npx wrangler secret put <NAME>` in the unit's folder;
628 `node scripts/deploy.mjs doctor` lists what is missing.
629
630Then `node scripts/deploy.mjs deploy --all`. A Worker bound to a service
631that does not exist yet may be refused; deploy that service first with
632`--only`.
633
634## OIDC tokens for workflow jobs
635
636The API is the issuer of workflow jobs' OIDC tokens,
637`https://api.g1t.sh/actions/oidc` (`apps/api/src/oidc.rs`): no host or DNS
638of its own. It signs with an RSA key kept as the API's secret
639`ACTIONS_OIDC_KEY`. Without it, the issuer's addresses answer 404 and jobs
640are not told where to ask for a token.
641
642To turn it on, make a key on a trusted machine and store it, then delete
643the file:
644
645```sh
646openssl genpkey -algorithm RSA -pkeyopt rsa_keygen_bits:2048 -out oidc.pem
647cd apps/api && npx wrangler secret put ACTIONS_OIDC_KEY < ../../oidc.pem
648rm ../../oidc.pem
649```
650
651Check `https://api.g1t.sh/actions/oidc/.well-known/jwks` lists one key.
652Its `kid` is the key's RFC 7638 thumbprint.
653
654To rotate it, keep the old key published while tokens it signed can still
655be presented (they last 5 minutes; relying parties cache keys for longer):
656
6571. Store the current key as `ACTIONS_OIDC_KEY_PREVIOUS` (the same PEM).
6582. Make a new key and store it as `ACTIONS_OIDC_KEY`. The JWKS now lists
659 both; new tokens are signed with the new one.
6603. A day later, `npx wrangler secret delete ACTIONS_OIDC_KEY_PREVIOUS`.
661
662A key that may have leaked is rotated the same way, skipping the first
663step, so that tokens it signed stop verifying at once.
664
665## Deployments on g1t
666
667Every deploy shows on the repository's Deployments page, so its production
668card says which commit runs.
669
670- **From the workflow**, the deploy jobs name `environment: {name:
671 production, url: https://g1t.sh}`, and g1t Actions records one production
672 deployment per run.
673- **By hand**, `deploy` reports one itself: in progress once migrations are
674 in, then success or failure. It needs a g1t token with `deployments:write`
675 in `G1T_DEPLOY_TOKEN` or the file `.credentials/g1t-deploy-token`; without
676 one, or from a dirty tree or a dry run, it sends nothing. It finds the
677 repository from the git remote on g1t.sh. A report that fails is one line
678 in the log and never fails the deploy.
679
680When CI cannot finish a deploy, for example a runner image that will not
681build there, deploy from a machine with Docker.
682The report records it, and production shows the commit that really runs.
683
684## Rolling back
685
686- **One unit, at once:** `npx wrangler rollback` in its folder (or
687 `npx wrangler rollback <version-id>`; `npx wrangler versions list` shows
688 each version's commit in its message). Code only: D1 migrations are not
689 undone. The next plan sees the older commit and deploys what changed since,
690 so revert the commit on `main` too, or the next push brings it back.
691- **To a commit:** check it out and `node scripts/deploy.mjs deploy --only
692 <units> --rollback`. Without `--rollback` the tool refuses any unit whose
693 live commit is newer than the one checked out, even with `--force`, so a
694 re-run of an old workflow run (or an old checkout) never rolls production
695 back by accident. Migrations never run backwards: a migration that needs
696 undoing is a new migration.
697- **The runner's image:** a rollback of the Worker does not roll back the
698 container image. Redeploy the older commit (`--only runner --rollback`): its
699 image's tag is the hash of that commit's source, which is still in the
700 registry, so nothing is built. A bad base is undone by reverting the
701 commit that changed `services/runner/base.json`.
702
703## Adding a unit
704
7051. Make its folder with a `wrangler.jsonc` (and its D1 migrations, if any).
706 A Rust Worker is a workspace member in the root `Cargo.toml` with
707 `"build": { "command": "node ../../scripts/build-rust-worker.mjs" }` and
708 the `wasm-opt = ["-O1"]` metadata; a TypeScript one is an npm workspace.
7092. Add it to `deploy/stack.jsonc`: path, kind, worker, `d1`, stage (the
710 earliest stage after everything it binds to), secrets, setup, self_host.
711 Name any new KV id under `resources.kv`.
7123. If its sources import a file outside its folder that is not a workspace
713 crate or package, list it under `inputs`.
7144. Add a row to the service table in `docs/SELF_HOSTING.md`.
7155. `npm run test:deploy` and `node scripts/deploy.mjs manifest --check`
716 say what is missing. Then create its resources and secrets, and
717 `node scripts/deploy.mjs deploy --only <unit>`.
718
719## The self-hosted runner
720
721`g1t-runner` (crates/runner) is also what customers run on their own
722machines (guide: `apps/docs/src/content/docs/guides/self-hosted-runners.md`).
723It is not deployed with the stack: it is released, and runners already out
724there update themselves to each release.
725
726| Piece | Where |
727| --- | --- |
728| The tool | `scripts/runner-release.mjs` (`keygen`, `build`, `sign`, `verify`, `publish`) |
729| The workflow | `.g1t/workflows/runner-release.yml`, on a tag `runner-v<version>` or by hand |
730| Where it is published | The R2 bucket `g1t-downloads`, served by the site at `g1t.sh/downloads/runner/<version>/<file>` and `/latest/<file>` (`apps/web/app/routes/downloads-runner.ts`) |
731| Its image | `deploy/runner/Dockerfile`, pushed to `g1t.sh/flagon-io/g1t-runner` (public) for amd64 and arm64 |
732
733A release is five binaries (Linux x64 and arm64, both static musl; macOS
734x64 and arm64; Windows x64), `SHA256SUMS`, and `manifest.json`;
735`latest.json` and its Ed25519 signature `latest.json.sig` name the newest.
736A runner updates only to a release whose signature checks out against the
737public key built into it and whose download matches the manifest's SHA-256.
738
739The first time:
740
7411. `node scripts/runner-release.mjs keygen`. Put `RUNNER_RELEASE_KEY` in the
742 repository's secrets (production environment) and keep a copy offline;
743 put the public key in its variables as `RUNNER_RELEASE_PUBLIC_KEY` (names
744 starting `G1T_` are reserved; the workflow hands it to the build as
745 `G1T_RUNNER_RELEASE_KEY`). A build made without the public key never
746 updates itself.
7472. `npx wrangler r2 bucket create g1t-downloads`, and deploy the site so it
748 has the `DOWNLOADS` binding.
7493. The image job pushes to g1t's own registry with the run's `G1T_TOKEN`;
750 the `g1t-runner` package in flagon-io is public. Set the variable
751 `RUNNER_AGENT_IMAGE` (a public copy of `g1t-runner-base`, the image agent
752 work runs in on customers' runners) once there is one.
7534. Register a self-hosted runner with the `docker` label for the image job.
754
755Each release:
756
7571. Bump `version` in `crates/runner/Cargo.toml` and merge it.
7582. Tag the commit `runner-v<version>` and push the tag. The workflow builds
759 every platform with cargo-zigbuild, signs, verifies, publishes the files
760 (the version's first, `latest.json` last), and pushes the image.
7613. By hand, the same is `node scripts/runner-release.mjs build`, then `sign`,
762 `verify` and `publish`, with the keys in the environment.
763
764Rolling back a release: copy the older version's `manifest.json` over
765`runner/latest.json` and sign it again (`sign` after checking the older
766version out). Runners never move to an older version on their own; a
767runner on a bad release is fixed by the next good one, or by downloading
768the older binary over it.