g1t/docs/DEPLOYING.md

590 lines33,298 bytesCodeBlame

Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.

Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow1# Deploying g1t
2
3How g1t.sh gets to Cloudflare: one manifest that lists every deployable
4unit, one tool that deploys only what changed, and a g1t Actions workflow
5that runs that tool on every push to `main`. Internal: the public
6self-hosting guide is `apps/docs/src/content/docs/guides/self-hosting.md`.
7
8| Piece | Where |
9| --- | --- |
10| The manifest | `deploy/stack.jsonc` |
11| The tool | `scripts/deploy.mjs` (library and tests in `scripts/deploy/`) |
12| Rust Worker builds | `scripts/build-rust-worker.mjs`, every Rust unit's build command |
Fast pages, required checks on the branch, self-hosted runners, honest incidents13| The runner's images | `services/runner/base/Dockerfile`, `services/runner/Dockerfile`, `services/runner/base.json`, `scripts/build-runner.mjs`, `scripts/deploy/image.mjs` |
14| The workflows | `.g1t/workflows/deploy.yml`, `.g1t/workflows/runner-base.yml` |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow15| The old entry point | `scripts/deploy.sh`, now a wrapper |
16
17## The manifest
18
19`deploy/stack.jsonc` names every unit: each folder with a `wrangler.jsonc`.
20It is JSONC rather than TOML so that Node reads it with no dependency,
21with the same parser as the Wrangler configs, and comments stay possible.
22
23| Field | |
24| --- | --- |
25| `path` | The unit's folder. |
26| `kind` | `rust-worker` (built by worker-build), `ts-worker` (Wrangler bundles it), `react-router` (`vite build` first), `astro` (`astro build` first). |
27| `worker` | The Worker's name. Must match its `wrangler.jsonc`. |
28| `d1` | `{ database, migrations }`, when it has a database. Must match its `wrangler.jsonc`. |
29| `stage` | `core`, `edge` or `front` (below). |
30| `secrets` | The Wrangler secrets it needs, by name. `node scripts/deploy.mjs doctor` checks they are set. |
31| `setup` | One-time steps no config can say, for a first deploy. |
32| `inputs` | Files outside its folder it is built from that no workspace metadata names. A test finds such imports. |
Fast pages, required checks on the branch, self-hosted runners, honest incidents33| `image` | A Containers image: `dockerfile` (the image deployed), `crate` (the binary it adds), `base` (`{ context, lock }`: the base image's folder and the file recording the base that was pushed) and `repository` (where both are pushed). See [the runner's images](#the-runners-images). |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow34| `self_host` | `run`, `off`, `separate` or `none`: what `deploy/self-host/configs.mjs` does with it. |
35
36What a unit is **built from** is never listed by hand. The tool reads it:
37Rust path dependencies from `cargo metadata`, workspace packages from each
38`package.json`, closed transitively. `node scripts/deploy.mjs manifest`
39prints the result:
40
41```
42unit stage kind worker d1 built from (besides its folder)
43events core rust-worker g1t-events g1t-events crates/contracts crates/kit
44repos core rust-worker g1t-repos g1t-repos crates/contracts crates/kit crates/scan crates/secrets
45runner core ts-worker g1t-runner crates/actions crates/runner packages/contracts
46og core ts-worker g1t-og packages/contracts
47web front react-router g1t packages/contracts packages/theme
48...
49```
50
51Root files count too: `Cargo.toml`, `Cargo.lock` and
52`scripts/build-rust-worker.mjs` for Rust units; `package-lock.json` and
53`tsconfig.base.json` for the others; `package.json` for all. A lockfile
54change counts for a unit only if a package in that unit's graph changed,
55read from the lockfile itself (`scripts/deploy/lockfiles.mjs`), so bumping
56`sharp` for the docs does not redeploy the Rust services.
57
58### Stages
59
60| Stage | What | Units |
61| --- | --- | --- |
62| `migrations` | Every pending D1 migration, in parallel, before any code | each unit's `d1` |
63| `core` | Services reached through bindings | the services, `og` |
64| `edge` | Public endpoints other than the site | `api`, `models`, `pages`, `status` |
65| `front` | The site, sudo, the docs | `web`, `sudo`, `docs` |
66
67A stage starts only when the one before it succeeded. Inside a stage units
68deploy in parallel. The rule the tests enforce: **a unit binds only to units
69in its own stage or an earlier one**, so new code never calls a service
70that has not shipped. (Services in `core` bind to each other in cycles,
71which is why they share a stage.)
72
73### What reads the manifest
74
75- `scripts/deploy.mjs`: everything below.
76- `deploy/self-host/configs.mjs`: which Workers a self-hosted installation
77 runs (`self_host: "run"`, the site first) and which are bound to the off
78 Worker (`"off"`). Its output is identical to before, apart from the order
79 of `workers.txt` after the site.
80- Tests (`npm run test:deploy`) check that: every `wrangler.jsonc` in the
81 repository has a unit; `worker`, `d1` and `image` match the configs;
82 every KV id is named under `resources.kv`; stages follow bindings; the
83 derived dependencies agree with Cargo's own resolved graph; every import
84 that leaves a unit's folder is covered; `deploy/self-host/Dockerfile`
85 builds exactly the Rust units self-hosting runs; and the service table in
86 `docs/SELF_HOSTING.md` names every unit.
87
88## The tool
89
90```sh
91node scripts/deploy.mjs plan # what would deploy, and why (read-only)
92node scripts/deploy.mjs deploy # migrations, then every changed unit
93node scripts/deploy.mjs deploy --only web,api # just these, if they changed
94node scripts/deploy.mjs deploy --only web --force # just this, changed or not
95node scripts/deploy.mjs deploy --all # everything
96node scripts/deploy.mjs build --only events # build as a deploy would; upload nothing
97node scripts/deploy.mjs migrate # pending migrations only
98node scripts/deploy.mjs manifest [--check|--json]
99node scripts/deploy.mjs doctor # secrets each unit lacks
Fast pages, required checks on the branch, self-hosted runners, honest incidents100node scripts/deploy.mjs build-base # build and push the runner's base image (Docker)
101node scripts/deploy.mjs image # build and push the runner's image for this checkout (Docker)
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow102scripts/deploy.sh [units...] # the old entry point: all, or those named, always
103```
104
105| Flag | |
106| --- | --- |
107| `--only a,b` / `--skip a,b` | Units by short name, folder or Worker name. |
108| `--all` | Every unit, changed or not. |
109| `--force` | Deploy the selected units even if unchanged. |
110| `--concurrency N` | Units at once inside a stage, and migrations at once (default 4). |
111| `--stage core` | One stage only. |
112| `--no-migrations` | Skip the migrations step (the workflow runs it as its own job). |
113| `--allow-dirty` | Deploy with uncommitted changes in what deploys. The version records no commit, so the next plan deploys it again. |
Fast pages, required checks on the branch, self-hosted runners, honest incidents114| `--rebuild-image` | Build the runner's image even if one for this source is already in the registry. |
115| `--rebuild-base` | Build and push a new base first (needs Docker), then the runner's image on it. Writes `services/runner/base.json`: commit it. |
116| `--no-push` | `build-base` and `image`: build locally, push nothing. |
117| `--no-cache` | `build-base`: build every layer again. |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow118| `--since REV` | Treat Workers with no recorded commit as running `REV`. Used once to adopt Workers deployed before this tool. |
119| `--json`, `--out FILE`, `--github-output` | The plan as data, for the workflow. |
120
121### Where the deployed commit is kept
122
123On the Worker itself. Every deploy runs `wrangler deploy --message
124"g1t-deploy <40-char sha> <subject>" --tag g1t-<12-char sha>`, which
125Cloudflare keeps as the version's `workers/message` and `workers/tag`
126annotations. `plan` reads them back with `wrangler deployments status`
127(the live version) and `wrangler versions list` (its annotations): two
128read-only calls per unit, in parallel; a plan of all 22 units takes about
12910 seconds. No KV namespace or other infrastructure is needed.
130
131- A version made by `wrangler secret put` keeps the code of the one before
132 it, so the tool looks through those to the deploy before.
133- A version deployed any other way (by hand, from the dashboard) has no
134 commit, and the unit is deployed again.
135- During a gradual rollout the version with the most traffic counts.
136- After `wrangler rollback`, the plan sees the older commit and deploys
137 what changed since.
138
139### What a deploy does
140
1411. Plans: for each unit, the live commit; `git diff` from it to `HEAD`;
142 whether the changed files touch the unit (its folder, the crates and
143 packages it is built from, its inputs, lockfile changes that reach it).
1442. Refuses if uncommitted changes touch what would deploy (`--allow-dirty`).
1453. Applies every pending migration (`wrangler d1 migrations apply --remote`),
146 in parallel. Any failure stops the deploy before code.
1474. Installs worker-build once if any Rust unit is deploying.
1485. Each stage in turn: units in parallel (`--concurrency`), each `npm run
149 build` first for React Router and Astro, then `wrangler deploy` in the
150 unit's folder with the annotation. If any unit fails, later stages are
151 not started.
1526. Prints a table: unit, stage, result, version id, time. Each unit's full
153 output is kept in `$TMPDIR/g1t-deploy/<unit>.log`.
154
Fast pages, required checks on the branch, self-hosted runners, honest incidents155#### Telling the status page about a deploy
156
157Restarts during a deploy can make a part slow for a minute, which the
158status page's checks would otherwise draft as an incident. Before the
159first stage and after the last, a deploy can say so with
160`scripts/deploy/status-window.mjs` (`announceDeploy("started" | "finished",
161{ id })`, or `node scripts/deploy/status-window.mjs started|finished [id]`).
162It posts to `POST https://status.g1t.sh/deploys` with
163`Authorization: Bearer $STATUS_DEPLOY_TOKEN`, the same value as the status
164Worker's `STATUS_DEPLOY_TOKEN` secret. During the deploy and for 3 minutes
165after it, detection keeps counting failed and slow checks but makes no new
166draft; trouble that outlasts that is drafted with its true start. A
167start with no finish stops counting after 30 minutes. Without the token
168the helper does nothing, and it never fails a deploy.
169
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow170On a laptop the tool uses your `wrangler login` (or `CLOUDFLARE_DEPLOY_TOKEN`
171if set), as `scripts/deploy.sh` always did: a `CLOUDFLARE_API_TOKEN` or
172global API key in your shell, or in the repository's `.env`, is ignored.
173With `CI=true` it uses `CLOUDFLARE_API_TOKEN`.
174
Fast pages, required checks on the branch, self-hosted runners, honest incidents175### The runner's images
176
177`services/runner` runs every sandbox (agents, checks, the merge queue,
178workflow jobs, g1t.page builds) from one Containers image, made in two
179parts:
180
181| Image | Built from | Holds | Rebuilt |
182| --- | --- | --- | --- |
183| **Base**, `g1t-runner:base-<date>-<inputs>` | `services/runner/base/Dockerfile` | Debian bookworm, Node 24, Python 3.11, Go (from go.dev), Rust stable for the `node` user with rustfmt, clippy and the `wasm32-unknown-unknown` target, build-essential, git, ripgrep, jq, zstd, sudo, and the pinned Claude Code CLI on top | When its folder changes, weekly, or by hand (`build-base`) |
184| **Runner**, `g1t-runner:<content hash>` | `services/runner/Dockerfile`: `FROM` the base, plus one file | The g1t runner, a static binary | When the binary or the base changes |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow185
Fast pages, required checks on the branch, self-hosted runners, honest incidents186Both are pushed to one repository of Cloudflare's registry,
187`registry.cloudflare.com/<account>/g1t-runner`, so pushing the runner's
188image uploads only its own layer (about 5 MB): the base's layers are
189already there.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow190
The runner's base image, pushed as Wrangler builds images, with pushes that fail loudly191Both are built as Wrangler builds images (`--platform linux/amd64
192--provenance=false --sbom=false`): one manifest, not an OCI index with a
193BuildKit attestation beside it. A push is tried up to three times and counts
194only when Docker reports the digest; the first push of the base once failed
195with `blob unknown to registry` and went through when run again (see
196`docs/CLOUDFLARE_FEEDBACK.md`, C3). `build-base` and `image` exit non-zero
197when a push fails, and `base.json` is written only after the push.
198
Fast pages, required checks on the branch, self-hosted runners, honest incidents199**The base** is recorded in `services/runner/base.json`: its reference, its
200digest, a hash of its folder (`inputs`), when it was built, its size and
201each toolchain's version. `node scripts/deploy.mjs build-base` builds it,
202pushes it and rewrites the file; commit the file, and the next deploy
203builds the runner's image on it. `npm run test:deploy` fails while the
204folder and the file disagree, so a change to the base's Dockerfile cannot
205merge without the base it describes. Layers go from what changes least to
206most (system packages, Go, Rust, the Claude Code CLI), and the apt and npm
207caches stay in BuildKit's cache, out of the image.
208
209The base's build cache is the base itself: it is built with
210`BUILDKIT_INLINE_CACHE`, which records in the image how each layer was
211made, and `build-base` builds `--cache-from` the base it replaces. A
212machine with an empty cache, or one just pruned, pulls the unchanged
213layers from the registry instead of building them. (BuildKit's other
214registry cache, `--cache-to type=registry`, pushes a separate cache
215manifest that not every registry takes, and for a one-stage image adds
216nothing the inline cache lacks.)
217
218**The runner binary** (`crates/runner`) is built outside Docker by
219`scripts/build-runner.mjs` as one static binary for
220`x86_64-unknown-linux-musl`, so it runs on the base whatever its libc, and
221on a self-hosted runner's machine too. Where it is built:
222
223- on x86-64 Linux with the musl target and `musl-gcc`
224 (`rustup target add x86_64-unknown-linux-musl`, `apt-get install musl-tools`),
225 with the machine's own Cargo;
226- anywhere else (Windows, macOS) in a small builder container, Rust on
227 Alpine (whose own target is musl), with Docker volumes keeping Cargo's
228 registry and target directory between builds. Windows has no musl
229 cross-linker, and `ring` (under ureq's TLS) needs a C compiler for the
230 target, so a container is the dependable route.
231
232**The runner's image tag** is a hash of everything it is built from: its
233Dockerfile, `base.json`, the crates the binary is built from, the
234workspace's Cargo files and the build script. The same source always names
235the same image, so:
236
2371. A deploy computes the tag and asks the registry whether it is there
238 (a `HEAD` of its manifest, with credentials from Wrangler; no Docker).
2392. If it is, nothing is built: the deploy uses it.
2403. If not, and Docker is here, it builds the binary and the image (seconds
241 on a warm machine) and pushes it.
2424. If not, and Docker is not here (a g1t Actions sandbox), the unit fails
243 saying to run `node scripts/deploy.mjs image` on a machine with Docker;
244 then re-run the workflow.
245
246Then `wrangler deploy` is given the image by reference, from a generated
247config (`services/runner/wrangler.deploy.json`, deleted after, ignored by
248git), so Wrangler builds nothing. When nothing the image is built from
249changed since the runner's live commit, the deploy also passes
250`--containers-rollout none`, which leaves running sandboxes alone.
251
252To get a new base out:
253
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow254```sh
Fast pages, required checks on the branch, self-hosted runners, honest incidents255node scripts/deploy.mjs build-base # build, push, write base.json (needs Docker)
256git commit services/runner/base.json -m "A new base image for g1t's sandboxes"
257node scripts/deploy.mjs image # optional: push the runner's image now, so CI finds it
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow258```
259
Fast pages, required checks on the branch, self-hosted runners, honest incidents260`.g1t/workflows/runner-base.yml` does the same weekly (and when the
261base's folder changes on `main`), and opens a pull request with
262`base.json`. It needs Docker, so it runs on a self-hosted runner with the
263`docker` label (`runs-on: [self-hosted, docker]`). Until one is
264registered, its runs wait for one; run `build-base` by hand instead.
265
266**Sandboxes start from the image.** Cloudflare pulls an image to a
267machine the first time a sandbox lands there, and keeps it. A smaller base
268pulls sooner, and a change to the runner alone sends machines one 5 MB
269layer instead of the whole image.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow270
Fast pages, required checks on the branch, self-hosted runners, honest incidents271#### Larger machines
272
273The same image runs on three instance types, each a Durable Object class
274of its own in `services/runner/wrangler.jsonc`: `AttemptSandbox`
275(`standard-1`), `Sandbox2Core` (`standard-3`) and `Sandbox4Core`
276(`standard-4`). Workflow jobs choose with `runs-on: g1t-2core` or
277`g1t-4core` (`g1t_contracts::actions::INSTANCE_TYPES`); the actions
278service passes the label to the runner, which starts the job in that
279class. Billing prices the larger ones from their memory and disk, and
280their CPU (see the public billing guide). The account's Containers limits
281must allow `standard-4`; Wrangler refuses the deploy otherwise.
282
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow283## Build speed
284
285Measured on the development machine (Windows, 32 cores, warm Cargo cache),
286building `events`, `search` and `repos` after a change to `crates/kit`, as a
287deploy does but without uploading (`wrangler deploy --dry-run`):
288
289| | Time |
290| --- | --- |
291| Before: one after another, `cargo install worker-build` each time, wasm-opt `-O` | 81 s, 84 s |
292| Concurrent builds, wasm-opt `-O` | 50 s, 68 s |
293| Concurrent builds, wasm-opt `-O1` (now) | 28 s, 33 s |
294
295Where the time went, and what changed:
296
297- **wasm-opt** was most of it: `-O` took 38 s on `repos` and 19 s on `api`;
298 `-O1` takes 1 to 3 s. With worker-build's flags (it keeps the names
299 section) the `.wasm` is 24 to 28% larger raw but only 1 to 7% larger
300 gzipped, and Workers' size limit is on the compressed upload. `-Os` and
301 `-Oz` were no faster than `-O`. Set per crate in
302 `[package.metadata.wasm-pack.profile.release]`; a test keeps every Rust
303 unit on the same level.
304- **worker-build** is installed only when missing or another version
305 (`scripts/build-rust-worker.mjs`, which pins it). Locally `cargo install`
306 on an installed version cost under a second; on a fresh CI sandbox it is
307 a full compile, which the workflow caches instead.
308- **One Cargo target**: every Rust unit is a member of the workspace, so
309 they already share `target/`. Concurrent builds take turns on Cargo's
310 lock for the compile, and their wasm-bindgen, wasm-opt and uploads
311 overlap.
312- **No joint `cargo build -p a -p b`**: tried, and it is slower. Cargo
313 unifies features across the packages of one build (`serde_json`'s
314 `preserve_order` from `api` and `actions`, `digest` features from
315 `secrets`), so each worker-build afterwards compiled its own variant
316 again.
The runner's base image, pushed as Wrangler builds images, with pushes that fail loudly317- **The runner's images**, measured on the same machine on 2026-10-06
318 (Docker Desktop, 8 vCPUs; its disk was busy with other containers, so
319 the cold figures are slow and noisy):
320
321 | | Before (one image) | Now |
322 | --- | --- | --- |
323 | Size, unpacked / compressed (what a machine pulls) | 3.08 GB / 819 MB | 2.68 GB / 686 MB |
324 | A change to the runner | Docker rebuilds the image's Rust stage and pushes the image (1198 s in the first deploy after a prune) | binary 46–53 s (6 s unchanged), image 7 s, push one 5 MB layer (1.4 s to a local registry) |
325 | The base from nothing | 431 s (whole image, cold) | 758 s cold, rarely: weekly or when its folder changes |
326 | The base after `docker builder prune` | as from nothing | 68 s, its layers pulled from the registry it was pushed to |
327 | The runner binary, cold (builder container) | | 99 s |
328 | A Rust CI job's build (events, search, repos; 4 vCPUs) | | 51 s cold, 13 s with the Cargo target restored (107 MB zstd entry) |
329
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow330- **Only what changed** is the largest saving: a change to one service
331 deploys one service.
332
333## The workflow
334
335`.g1t/workflows/deploy.yml` runs on every push to `main`, and by hand
336(**Actions → Deploy → Run workflow**) with `units` (deploy these, changed or
337not), `all` and `dry_run` (plan only).
338
339| Job | Does | Needs |
340| --- | --- | --- |
341| `check` | `manifest --check` and `npm run test:deploy` | — |
342| `plan` | `plan --github-output`: outputs per stage, the plan in the run's summary | `check` |
343| `migrate` | `migrate --only <units with pending migrations>` | `plan`; skipped when none are pending |
344| `core`, `edge`, `front` | `deploy --only <units> --force --no-migrations`, one job per build group | the stages before; skipped when empty |
345
346- **One at a time:** `concurrency: deploy-production`, never cancelled in
347 progress; a second push waits.
348- **Build groups:** a stage's units are split so each job shares a build:
Fast pages, required checks on the branch, self-hosted runners, honest incidents349 Rust workers at most four to a job (each a 4-vCPU `g1t-4core` machine), the
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow350 TypeScript Workers together, each site alone, and a unit whose image must
351 be rebuilt alone. `fail-fast: false`, so one failed job does not cut
352 another off mid-upload; the next stage then does not start.
353- **Tests:** there is no CI workflow on g1t yet; `main` is kept passing by
354 the merge queue's checks. `check` runs the deploy tool's own tests. When a
355 CI workflow is added, make `plan` wait for it (`workflow_run`, or a job in
356 this file).
Fast pages, required checks on the branch, self-hosted runners, honest incidents357- **Machines:** Rust jobs run on `g1t-4core` (4 vCPUs, 12 GiB), the
358 others on the standard machine (`runs-on: ${{ matrix.rust && 'g1t-4core' || 'ubuntu-latest' }}`).
359- **Caching** (`actions/cache`: up to 2 GB an entry, 10 GB a repository,
360 kept until unused for 7 days): the worker-build binary, worker-build's
361 downloaded tools, `~/.cargo/registry/cache`, and the Cargo target's
362 release dependencies (`target/release` and
363 `target/wasm32-unknown-unknown/release`, without `incremental` or
364 `.wasm`), keyed by the build group, `Cargo.lock` and `base.json`. The
365 workspace's own crates are compiled again on every run (a checkout's
366 sources are newer than any cache); the crates.io dependencies are not.
367 npm's cache is not kept: every job runs `npm ci` of only what its units
368 need (`deploy.mjs install`: Wrangler alone for Rust jobs).
369- **Conditions:** each stage runs with `!failure() && !cancelled()`, which
370 on g1t (as on GitHub) is true when no job before it failed, however far
371 back: a `migrate` job skipped for having nothing to apply does not stop
372 the stages after it, and a failed `check` stops all of them.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow373- `crates/actions/tests/repository_workflows.rs` reads the workflow with
374 g1t's own parser and expressions, and checks the jobs start, wait and
375 stop as above (`cargo test -p g1t-actions --test repository_workflows`).
376
377### What the sandbox has
378
Fast pages, required checks on the branch, self-hosted runners, honest incidents379The base image (`services/runner/base/Dockerfile`) has Node 24, npm, git,
380Go, zstd, and Rust stable for the `node` user with rustfmt, clippy and the
381`wasm32-unknown-unknown` target, but not worker-build or Docker. The
382workflow's `rustup target add wasm32-unknown-unknown` is then a no-op, and
383worker-build is restored from the cache, installed on a miss. worker-build
384fetches wasm-bindgen and wasm-opt from GitHub releases and esbuild from
385npm. All of those hosts are on the list every workflow job may reach.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow386
Fast pages, required checks on the branch, self-hosted runners, honest incidents387### Network
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow388
Fast pages, required checks on the branch, self-hosted runners, honest incidents389A workflow job reaches its project's allowed domains, g1t, what builds
390need (`services/runner/src/egress.ts`, `BUILD_HOSTS`), and the project's
391**workflow-only domains** that name its workflow and environment. Those are
392never reached by agents, checks, the merge queue, deploy builds or runs of
393pull requests from forks (`Guardrails::workflow_hosts`, the runner's
394`jobHosts`). Under flagon-io/g1t's **Settings → Guardrails**
395(Maintain role or higher), **Workflow-only domains**:
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow396
The runner's base image, pushed as Wrangler builds images, with pushes that fail loudly397```text
Fast pages, required checks on the branch, self-hosted runners, honest incidents398api.cloudflare.com | deploy.yml | production
399registry.cloudflare.com | deploy.yml, runner-base.yml | production
400```
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow401
Fast pages, required checks on the branch, self-hosted runners, honest incidents402`api.cloudflare.com` is Wrangler's API; `registry.cloudflare.com` is where
403the deploy asks whether the runner's image is already built (and where
404`runner-base.yml` pushes). Only `deploy.yml`'s and `runner-base.yml`'s
405jobs with `environment: production` reach them, which are also the only
406jobs that can read `CLOUDFLARE_API_TOKEN`. Each change to the list is in
407the workspace's audit log as `update_guardrails`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow408
409### The API token
410
411Create it at **dash.cloudflare.com → My Profile → API Tokens → Create
412Token → Custom token**, named `g1t deploys (CI)`:
413
414| Scope | Permission | Why |
415| --- | --- | --- |
416| Account | Workers Scripts: Edit | Upload, versions, deployments, crons, bindings, `secret list` (doctor) |
417| Account | D1: Edit | `d1 migrations list` and `apply` |
418| Account | Queues: Edit | Attaching each unit's queue consumers on deploy |
419| Account | Workers R2 Storage: Read | Wrangler checks `og`'s bucket binding |
420| Account | Account Settings: Read | Wrangler reads the account |
Fast pages, required checks on the branch, self-hosted runners, honest incidents421| Account | Containers: Edit | The runner's deploy updates its applications (the image reference, the three classes), and gets registry credentials to look for its image. `runner-base.yml` pushes images with it. |
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow422| Zone (`g1t.sh`, `g1t.page`) | Workers Routes: Edit | `pages`' zone routes, and custom domains |
423| Zone (`g1t.sh`, `g1t.page`) | DNS: Edit | Custom domains (`api`, `mcp`, `og`, `models`, `status`, `sudo`, `docs`, `g1t.sh`, `g1t.page`) keep their DNS records |
424| Zone (`g1t.sh`, `g1t.page`) | Zone: Read | Finding the zone a route names |
425
426Restrict it to account `syntaqx` (`1e6f2cffa3f445920836e8ebe446bb58`) and
427the two zones. Not needed for deploys: KV (bindings are by id; creating a
428namespace is a one-time setup), Vectorize, Workers for Platforms beyond
429Workers Scripts, Cloudflare for SaaS custom hostnames (the deployments
430service does that at runtime with its own token), SSL and Certificates.
431Workers KV Storage: Edit and Vectorize: Edit are only for first-time setup,
432which stays a person's job.
433
434The list follows what our configs use; Cloudflare does not publish exactly
435what `wrangler deploy` checks for each binding. Bindings with no listed
436permission (Browser Rendering, Workers AI, Vectorize, Email Sending,
437Artifacts, dispatch namespaces) are assumed to need none beyond Workers
438Scripts. Verify on the first run with `workflow_dispatch` and `units:
439pages` (small, no secrets), then `units: og` (R2) and `units: runner`
440(Containers); a missing permission fails with `Authentication error [code:
44110000]` and the route it was refused.
442
443Add it to the repository:
444
4451. On g1t.sh, open **flagon-io/g1t → Settings → Secrets and variables**.
4462. **Add**: key `CLOUDFLARE_API_TOKEN`, type **Secret**, available to
447 **Workflows**, environment **Production**. Only jobs with
448 `environment: production` (the deploy jobs) can read it.
4493. **Add**: key `CLOUDFLARE_ACCOUNT_ID`, type **Variable**, value
450 `1e6f2cffa3f445920836e8ebe446bb58`, available to **Workflows**, all
451 environments.
452
453Or through the API:
454
455```sh
456curl -X PUT https://api.g1t.sh/repos/flagon-io/g1t/actions/secrets/CLOUDFLARE_API_TOKEN \
457 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
458 -d '{"value":"<the token>","environments":["production"],"available_to":["workflows"]}'
459
460curl -X POST https://api.g1t.sh/repos/flagon-io/g1t/actions/variables \
461 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
462 -d '{"name":"CLOUDFLARE_ACCOUNT_ID","value":"1e6f2cffa3f445920836e8ebe446bb58","available_to":["workflows"]}'
463```
464
465## Turning it on
466
4671. Create the token and add the secret and variable (above).
Fast pages, required checks on the branch, self-hosted runners, honest incidents4682. Add `api.cloudflare.com` and `registry.cloudflare.com` to flagon-io/g1t's
469 workflow-only domains, for `deploy.yml` in `production` (above).
4703. Build and push the base once, and commit `services/runner/base.json`:
471 `node scripts/deploy.mjs build-base`. Create the cache bucket:
472 `npx wrangler r2 bucket create g1t-actions-cache`, with a lifecycle rule
473 deleting objects 30 days after upload
474 (`npx wrangler r2 bucket lifecycle add g1t-actions-cache expire --expire-days 30 --abort-multipart-days 1`).
4754. Adopt the live Workers once, from a laptop: deploy everything with the
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow476 tool so each version records its commit (`scripts/deploy.sh`, or
477 `node scripts/deploy.mjs deploy --all`). Until then every plan says "no
478 known commit" and deploys every unit. To see what has changed since a
479 commit you know production runs, without deploying:
480 `node scripts/deploy.mjs plan --since <sha>`.
Fast pages, required checks on the branch, self-hosted runners, honest incidents4815. Run the workflow by hand with `dry_run`, then with `units: pages`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow482
483## First deploy of a new account
484
485What the configs refer to must exist first. `node scripts/deploy.mjs
486manifest --json` lists, per unit, its queues, KV, R2, Vectorize and
487dispatch namespaces; each unit's `setup` and `secrets` say the rest.
488
489- D1: `npx wrangler d1 create <database>`, then put its id in the unit's
490 `wrangler.jsonc`.
491- KV: `npx wrangler kv namespace create <name>` for each name under
492 `resources.kv`, then the ids in the configs.
493- Queues: `npx wrangler queues create <queue>` for each queue in the
494 manifest: `g1t-events`, `g1t-events-<service>` for every subscriber,
495 `g1t-search-jobs`, `g1t-context-jobs`.
Fast pages, required checks on the branch, self-hosted runners, honest incidents496- R2: `npx wrangler r2 bucket create g1t-screenshots` and
497 `npx wrangler r2 bucket create g1t-actions-cache` (with its 30-day
498 lifecycle rule, above).
499- The runner's base image: `node scripts/deploy.mjs build-base`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow500- Vectorize, dispatch namespace, DNS, Access, Email Sending, Artifacts: each
501 unit's `setup`.
502- Secrets: `npx wrangler secret put <NAME>` in the unit's folder;
503 `node scripts/deploy.mjs doctor` lists what is missing.
504
505Then `node scripts/deploy.mjs deploy --all`. A Worker bound to a service
506that does not exist yet may be refused; deploy that service first with
507`--only`.
508
509## Rolling back
510
511- **One unit, at once:** `npx wrangler rollback` in its folder (or
512 `npx wrangler rollback <version-id>`; `npx wrangler versions list` shows
513 each version's commit in its message). Code only: D1 migrations are not
514 undone. The next plan sees the older commit and deploys what changed since,
515 so revert the commit on `main` too, or the next push brings it back.
516- **To a commit:** check it out and `node scripts/deploy.mjs deploy --only
517 <units> --force`. Migrations never run backwards: a migration that needs
518 undoing is a new migration.
519- **The runner's image:** a rollback of the Worker does not roll back the
Fast pages, required checks on the branch, self-hosted runners, honest incidents520 container image. Redeploy the older commit (`--only runner --force`): its
521 image's tag is the hash of that commit's source, which is still in the
522 registry, so nothing is built. A bad base is undone by reverting the
523 commit that changed `services/runner/base.json`.
Deploys as code: a manifest of every Worker, a deploy tool that ships only what changed in parallel stages, and a g1t Actions workflow524
525## Adding a unit
526
5271. Make its folder with a `wrangler.jsonc` (and its D1 migrations, if any).
528 A Rust Worker is a workspace member in the root `Cargo.toml` with
529 `"build": { "command": "node ../../scripts/build-rust-worker.mjs" }` and
530 the `wasm-opt = ["-O1"]` metadata; a TypeScript one is an npm workspace.
5312. Add it to `deploy/stack.jsonc`: path, kind, worker, `d1`, stage (the
532 earliest stage after everything it binds to), secrets, setup, self_host.
533 Name any new KV id under `resources.kv`.
5343. If its sources import a file outside its folder that is not a workspace
535 crate or package, list it under `inputs`.
5364. Add a row to the service table in `docs/SELF_HOSTING.md`.
5375. `npm run test:deploy` and `node scripts/deploy.mjs manifest --check`
538 say what is missing. Then create its resources and secrets, and
539 `node scripts/deploy.mjs deploy --only <unit>`.
Fast pages, required checks on the branch, self-hosted runners, honest incidents540
541## The self-hosted runner
542
543`g1t-runner` (crates/runner) is also what customers run on their own
544machines (guide: `apps/docs/src/content/docs/guides/self-hosted-runners.md`).
545It is not deployed with the stack: it is released, and runners already out
546there update themselves to each release.
547
548| Piece | Where |
549| --- | --- |
550| The tool | `scripts/runner-release.mjs` (`keygen`, `build`, `sign`, `verify`, `publish`) |
551| The workflow | `.g1t/workflows/runner-release.yml`, on a tag `runner-v<version>` or by hand |
552| Where it is published | The R2 bucket `g1t-downloads`, served by the site at `g1t.sh/downloads/runner/<version>/<file>` and `/latest/<file>` (`apps/web/app/routes/downloads-runner.ts`) |
The runner's image is g1t.sh/flagon-io/g1t-runner, on g1t's own registry, public553| Its image | `deploy/runner/Dockerfile`, pushed to `g1t.sh/flagon-io/g1t-runner` (public) for amd64 and arm64 |
Fast pages, required checks on the branch, self-hosted runners, honest incidents554
555A release is five binaries (Linux x64 and arm64, both static musl; macOS
556x64 and arm64; Windows x64), `SHA256SUMS`, and `manifest.json`;
557`latest.json` and its Ed25519 signature `latest.json.sig` name the newest.
558A runner updates only to a release whose signature checks out against the
559public key built into it and whose download matches the manifest's SHA-256.
560
561The first time:
562
5631. `node scripts/runner-release.mjs keygen`. Put `RUNNER_RELEASE_KEY` in the
564 repository's secrets (production environment) and keep a copy offline;
g1t-runner 0.1.0 is released: signed binaries for five platforms at g1t.sh/downloads/runner; the Artifacts checks' results565 put the public key in its variables as `RUNNER_RELEASE_PUBLIC_KEY` (names
566 starting `G1T_` are reserved; the workflow hands it to the build as
567 `G1T_RUNNER_RELEASE_KEY`). A build made without the public key never
568 updates itself.
Fast pages, required checks on the branch, self-hosted runners, honest incidents5692. `npx wrangler r2 bucket create g1t-downloads`, and deploy the site so it
570 has the `DOWNLOADS` binding.
The runner's image is g1t.sh/flagon-io/g1t-runner, on g1t's own registry, public5713. The image job pushes to g1t's own registry with the run's `G1T_TOKEN`;
572 the `g1t-runner` package in flagon-io is public. Set the variable
573 `RUNNER_AGENT_IMAGE` (a public copy of `g1t-runner-base`, the image agent
574 work runs in on customers' runners) once there is one.
Fast pages, required checks on the branch, self-hosted runners, honest incidents5754. Register a self-hosted runner with the `docker` label for the image job.
576
577Each release:
578
5791. Bump `version` in `crates/runner/Cargo.toml` and merge it.
5802. Tag the commit `runner-v<version>` and push the tag. The workflow builds
581 every platform with cargo-zigbuild, signs, verifies, publishes the files
582 (the version's first, `latest.json` last), and pushes the image.
5833. By hand, the same is `node scripts/runner-release.mjs build`, then `sign`,
584 `verify` and `publish`, with the keys in the environment.
585
586Rolling back a release: copy the older version's `manifest.json` over
587`runner/latest.json` and sign it again (`sign` after checking the older
588version out). Runners never move to an older version on their own; a
589runner on a bad release is fixed by the next good one, or by downloading
590the older binary over it.