flagon-io/g1t

public

Where people and agents ship software together. The open-source git platform for the whole job: issues, agents, checks and deploys to the edge.

g1t/docs/DEPLOYING.md

394 lines21,033 bytesCodeBlame
1# Deploying g1t
2
3How g1t.sh gets to Cloudflare: one manifest that lists every deployable
4unit, one tool that deploys only what changed, and a g1t Actions workflow
5that runs that tool on every push to `main`. Internal: the public
6self-hosting guide is `apps/docs/src/content/docs/guides/self-hosting.md`.
7
8| Piece | Where |
9| --- | --- |
10| The manifest | `deploy/stack.jsonc` |
11| The tool | `scripts/deploy.mjs` (library and tests in `scripts/deploy/`) |
12| Rust Worker builds | `scripts/build-rust-worker.mjs`, every Rust unit's build command |
13| The workflow | `.g1t/workflows/deploy.yml` |
14| The old entry point | `scripts/deploy.sh`, now a wrapper |
15
16## The manifest
17
18`deploy/stack.jsonc` names every unit: each folder with a `wrangler.jsonc`.
19It is JSONC rather than TOML so that Node reads it with no dependency,
20with the same parser as the Wrangler configs, and comments stay possible.
21
22| Field | |
23| --- | --- |
24| `path` | The unit's folder. |
25| `kind` | `rust-worker` (built by worker-build), `ts-worker` (Wrangler bundles it), `react-router` (`vite build` first), `astro` (`astro build` first). |
26| `worker` | The Worker's name. Must match its `wrangler.jsonc`. |
27| `d1` | `{ database, migrations }`, when it has a database. Must match its `wrangler.jsonc`. |
28| `stage` | `core`, `edge` or `front` (below). |
29| `secrets` | The Wrangler secrets it needs, by name. `node scripts/deploy.mjs doctor` checks they are set. |
30| `setup` | One-time steps no config can say, for a first deploy. |
31| `inputs` | Files outside its folder it is built from that no workspace metadata names. A test finds such imports. |
32| `image` | A Containers image (`dockerfile`, and the `crate` it compiles), which needs Docker to build. |
33| `self_host` | `run`, `off`, `separate` or `none`: what `deploy/self-host/configs.mjs` does with it. |
34
35What a unit is **built from** is never listed by hand. The tool reads it:
36Rust path dependencies from `cargo metadata`, workspace packages from each
37`package.json`, closed transitively. `node scripts/deploy.mjs manifest`
38prints the result:
39
40```
41unit stage kind worker d1 built from (besides its folder)
42events core rust-worker g1t-events g1t-events crates/contracts crates/kit
43repos core rust-worker g1t-repos g1t-repos crates/contracts crates/kit crates/scan crates/secrets
44runner core ts-worker g1t-runner crates/actions crates/runner packages/contracts
45og core ts-worker g1t-og packages/contracts
46web front react-router g1t packages/contracts packages/theme
47...
48```
49
50Root files count too: `Cargo.toml`, `Cargo.lock` and
51`scripts/build-rust-worker.mjs` for Rust units; `package-lock.json` and
52`tsconfig.base.json` for the others; `package.json` for all. A lockfile
53change counts for a unit only if a package in that unit's graph changed,
54read from the lockfile itself (`scripts/deploy/lockfiles.mjs`), so bumping
55`sharp` for the docs does not redeploy the Rust services.
56
57### Stages
58
59| Stage | What | Units |
60| --- | --- | --- |
61| `migrations` | Every pending D1 migration, in parallel, before any code | each unit's `d1` |
62| `core` | Services reached through bindings | the services, `og` |
63| `edge` | Public endpoints other than the site | `api`, `models`, `pages`, `status` |
64| `front` | The site, sudo, the docs | `web`, `sudo`, `docs` |
65
66A stage starts only when the one before it succeeded. Inside a stage units
67deploy in parallel. The rule the tests enforce: **a unit binds only to units
68in its own stage or an earlier one**, so new code never calls a service
69that has not shipped. (Services in `core` bind to each other in cycles,
70which is why they share a stage.)
71
72### What reads the manifest
73
74- `scripts/deploy.mjs`: everything below.
75- `deploy/self-host/configs.mjs`: which Workers a self-hosted installation
76 runs (`self_host: "run"`, the site first) and which are bound to the off
77 Worker (`"off"`). Its output is identical to before, apart from the order
78 of `workers.txt` after the site.
79- Tests (`npm run test:deploy`) check that: every `wrangler.jsonc` in the
80 repository has a unit; `worker`, `d1` and `image` match the configs;
81 every KV id is named under `resources.kv`; stages follow bindings; the
82 derived dependencies agree with Cargo's own resolved graph; every import
83 that leaves a unit's folder is covered; `deploy/self-host/Dockerfile`
84 builds exactly the Rust units self-hosting runs; and the service table in
85 `docs/SELF_HOSTING.md` names every unit.
86
87## The tool
88
89```sh
90node scripts/deploy.mjs plan # what would deploy, and why (read-only)
91node scripts/deploy.mjs deploy # migrations, then every changed unit
92node scripts/deploy.mjs deploy --only web,api # just these, if they changed
93node scripts/deploy.mjs deploy --only web --force # just this, changed or not
94node scripts/deploy.mjs deploy --all # everything
95node scripts/deploy.mjs build --only events # build as a deploy would; upload nothing
96node scripts/deploy.mjs migrate # pending migrations only
97node scripts/deploy.mjs manifest [--check|--json]
98node scripts/deploy.mjs doctor # secrets each unit lacks
99scripts/deploy.sh [units...] # the old entry point: all, or those named, always
100```
101
102| Flag | |
103| --- | --- |
104| `--only a,b` / `--skip a,b` | Units by short name, folder or Worker name. |
105| `--all` | Every unit, changed or not. |
106| `--force` | Deploy the selected units even if unchanged. |
107| `--concurrency N` | Units at once inside a stage, and migrations at once (default 4). |
108| `--stage core` | One stage only. |
109| `--no-migrations` | Skip the migrations step (the workflow runs it as its own job). |
110| `--allow-dirty` | Deploy with uncommitted changes in what deploys. The version records no commit, so the next plan deploys it again. |
111| `--rebuild-image` | Build the runner's image even if nothing it is built from changed. |
112| `--since REV` | Treat Workers with no recorded commit as running `REV`. Used once to adopt Workers deployed before this tool. |
113| `--json`, `--out FILE`, `--github-output` | The plan as data, for the workflow. |
114
115### Where the deployed commit is kept
116
117On the Worker itself. Every deploy runs `wrangler deploy --message
118"g1t-deploy <40-char sha> <subject>" --tag g1t-<12-char sha>`, which
119Cloudflare keeps as the version's `workers/message` and `workers/tag`
120annotations. `plan` reads them back with `wrangler deployments status`
121(the live version) and `wrangler versions list` (its annotations): two
122read-only calls per unit, in parallel; a plan of all 22 units takes about
12310 seconds. No KV namespace or other infrastructure is needed.
124
125- A version made by `wrangler secret put` keeps the code of the one before
126 it, so the tool looks through those to the deploy before.
127- A version deployed any other way (by hand, from the dashboard) has no
128 commit, and the unit is deployed again.
129- During a gradual rollout the version with the most traffic counts.
130- After `wrangler rollback`, the plan sees the older commit and deploys
131 what changed since.
132
133### What a deploy does
134
1351. Plans: for each unit, the live commit; `git diff` from it to `HEAD`;
136 whether the changed files touch the unit (its folder, the crates and
137 packages it is built from, its inputs, lockfile changes that reach it).
1382. Refuses if uncommitted changes touch what would deploy (`--allow-dirty`).
1393. Applies every pending migration (`wrangler d1 migrations apply --remote`),
140 in parallel. Any failure stops the deploy before code.
1414. Installs worker-build once if any Rust unit is deploying.
1425. Each stage in turn: units in parallel (`--concurrency`), each `npm run
143 build` first for React Router and Astro, then `wrangler deploy` in the
144 unit's folder with the annotation. If any unit fails, later stages are
145 not started.
1466. Prints a table: unit, stage, result, version id, time. Each unit's full
147 output is kept in `$TMPDIR/g1t-deploy/<unit>.log`.
148
149On a laptop the tool uses your `wrangler login` (or `CLOUDFLARE_DEPLOY_TOKEN`
150if set), as `scripts/deploy.sh` always did: a `CLOUDFLARE_API_TOKEN` or
151global API key in your shell, or in the repository's `.env`, is ignored.
152With `CI=true` it uses `CLOUDFLARE_API_TOKEN`.
153
154### The runner's image
155
156`services/runner` deploys a Containers image built from
157`services/runner/Dockerfile`, which needs Docker. The tool rebuilds it only
158when something the image is built from changed since the runner's live
159commit (the Dockerfile, `crates/runner` and the crates it uses,
160`Cargo.toml`, `Cargo.lock`), or with `--rebuild-image`. Otherwise it deploys
161the Worker with `--containers-rollout none`, which leaves the running image
162alone. g1t Actions sandboxes have no Docker, so a CI deploy that needs a new
163image fails that unit with what to do, and later stages wait:
164
165```sh
166node scripts/deploy.mjs deploy --only runner # on a machine with Docker
167```
168
169then re-run the workflow; its plan now sees the runner up to date.
170
171## Build speed
172
173Measured on the development machine (Windows, 32 cores, warm Cargo cache),
174building `events`, `search` and `repos` after a change to `crates/kit`, as a
175deploy does but without uploading (`wrangler deploy --dry-run`):
176
177| | Time |
178| --- | --- |
179| Before: one after another, `cargo install worker-build` each time, wasm-opt `-O` | 81 s, 84 s |
180| Concurrent builds, wasm-opt `-O` | 50 s, 68 s |
181| Concurrent builds, wasm-opt `-O1` (now) | 28 s, 33 s |
182
183Where the time went, and what changed:
184
185- **wasm-opt** was most of it: `-O` took 38 s on `repos` and 19 s on `api`;
186 `-O1` takes 1 to 3 s. With worker-build's flags (it keeps the names
187 section) the `.wasm` is 24 to 28% larger raw but only 1 to 7% larger
188 gzipped, and Workers' size limit is on the compressed upload. `-Os` and
189 `-Oz` were no faster than `-O`. Set per crate in
190 `[package.metadata.wasm-pack.profile.release]`; a test keeps every Rust
191 unit on the same level.
192- **worker-build** is installed only when missing or another version
193 (`scripts/build-rust-worker.mjs`, which pins it). Locally `cargo install`
194 on an installed version cost under a second; on a fresh CI sandbox it is
195 a full compile, which the workflow caches instead.
196- **One Cargo target**: every Rust unit is a member of the workspace, so
197 they already share `target/`. Concurrent builds take turns on Cargo's
198 lock for the compile, and their wasm-bindgen, wasm-opt and uploads
199 overlap.
200- **No joint `cargo build -p a -p b`**: tried, and it is slower. Cargo
201 unifies features across the packages of one build (`serde_json`'s
202 `preserve_order` from `api` and `actions`, `digest` features from
203 `secrets`), so each worker-build afterwards compiled its own variant
204 again.
205- **Only what changed** is the largest saving: a change to one service
206 deploys one service.
207
208## The workflow
209
210`.g1t/workflows/deploy.yml` runs on every push to `main`, and by hand
211(**Actions → Deploy → Run workflow**) with `units` (deploy these, changed or
212not), `all` and `dry_run` (plan only).
213
214| Job | Does | Needs |
215| --- | --- | --- |
216| `check` | `manifest --check` and `npm run test:deploy` | — |
217| `plan` | `plan --github-output`: outputs per stage, the plan in the run's summary | `check` |
218| `migrate` | `migrate --only <units with pending migrations>` | `plan`; skipped when none are pending |
219| `core`, `edge`, `front` | `deploy --only <units> --force --no-migrations`, one job per build group | the stages before; skipped when empty |
220
221- **One at a time:** `concurrency: deploy-production`, never cancelled in
222 progress; a second push waits.
223- **Build groups:** a stage's units are split so each job shares a build:
224 Rust workers at most four to a job (a sandbox has half a CPU), the
225 TypeScript Workers together, each site alone, and a unit whose image must
226 be rebuilt alone. `fail-fast: false`, so one failed job does not cut
227 another off mid-upload; the next stage then does not start.
228- **Tests:** there is no CI workflow on g1t yet; `main` is kept passing by
229 the merge queue's checks. `check` runs the deploy tool's own tests. When a
230 CI workflow is added, make `plan` wait for it (`workflow_run`, or a job in
231 this file).
232- **Caching** (`actions/cache`, kept per repository for 7 days, at most
233 60 MB an entry): the worker-build binary, worker-build's downloaded tools,
234 and `~/.cargo/registry/cache`. Not cached: the Cargo target directory
235 (about 500 MB for these crates) and npm's cache (over 150 MB), so every
236 Rust job compiles its crates from scratch and every job runs `npm ci` of
237 only what its units need (`deploy.mjs install`: Wrangler alone for Rust
238 jobs).
239- `crates/actions/tests/repository_workflows.rs` reads the workflow with
240 g1t's own parser and expressions, and checks the jobs start, wait and
241 stop as above (`cargo test -p g1t-actions --test repository_workflows`).
242
243### What the sandbox has
244
245The runner image (`services/runner/Dockerfile`) has Node 24, npm, git, and
246Rust stable for the `node` user (rustfmt, clippy) but not the
247`wasm32-unknown-unknown` target, worker-build or Docker. The workflow adds
248the target (`rustup target add`, from `static.rust-lang.org`) and restores
249worker-build from the cache, installing it on a miss. worker-build fetches
250wasm-bindgen and wasm-opt from GitHub releases and esbuild from npm. All
251of those hosts are on the list every workflow job may reach.
252
253Adding the wasm target and worker-build to the image instead would save
254about a minute per Rust job, but every image change needs a Docker deploy of
255the runner and replaces every sandbox, so it is left for when the image
256next changes anyway.
257
258### Network
259
260A workflow job reaches only its project's allowed domains, g1t, and what
261builds need (`services/runner/src/egress.ts`, `BUILD_HOSTS`). Cloudflare's
262API is not among them for workflows (only for g1t.page deploy builds), and
263guardrails have no per-workflow list. The least that works today: add
264`api.cloudflare.com` to **flagon-io/g1t's allowed domains** (repository
265**Settings → Guardrails**, Maintain role or higher).
266
267That opens the host to every sandbox of that project, agents included.
268Agents never get the token (secrets go only to trusted workflow jobs), but
269any code there could talk to Cloudflare's API with credentials of its own.
270The better fix is a list of domains only workflow jobs may reach, set by a
271maintainer, or hosts a job asks for honored only for trusted jobs; that is
272a change to guardrails (`crates/contracts/src/guardrails.rs`, the work
273service, the runner's `buildGuardFor`).
274
275### The API token
276
277Create it at **dash.cloudflare.com → My Profile → API Tokens → Create
278Token → Custom token**, named `g1t deploys (CI)`:
279
280| Scope | Permission | Why |
281| --- | --- | --- |
282| Account | Workers Scripts: Edit | Upload, versions, deployments, crons, bindings, `secret list` (doctor) |
283| Account | D1: Edit | `d1 migrations list` and `apply` |
284| Account | Queues: Edit | Attaching each unit's queue consumers on deploy |
285| Account | Workers R2 Storage: Read | Wrangler checks `og`'s bucket binding |
286| Account | Account Settings: Read | Wrangler reads the account |
287| Account | Containers: Read | The runner's deploy with `--containers-rollout none` reads its application (Edit only if CI ever builds images) |
288| Zone (`g1t.sh`, `g1t.page`) | Workers Routes: Edit | `pages`' zone routes, and custom domains |
289| Zone (`g1t.sh`, `g1t.page`) | DNS: Edit | Custom domains (`api`, `mcp`, `og`, `models`, `status`, `sudo`, `docs`, `g1t.sh`, `g1t.page`) keep their DNS records |
290| Zone (`g1t.sh`, `g1t.page`) | Zone: Read | Finding the zone a route names |
291
292Restrict it to account `syntaqx` (`1e6f2cffa3f445920836e8ebe446bb58`) and
293the two zones. Not needed for deploys: KV (bindings are by id; creating a
294namespace is a one-time setup), Vectorize, Workers for Platforms beyond
295Workers Scripts, Cloudflare for SaaS custom hostnames (the deployments
296service does that at runtime with its own token), SSL and Certificates.
297Workers KV Storage: Edit and Vectorize: Edit are only for first-time setup,
298which stays a person's job.
299
300The list follows what our configs use; Cloudflare does not publish exactly
301what `wrangler deploy` checks for each binding. Bindings with no listed
302permission (Browser Rendering, Workers AI, Vectorize, Email Sending,
303Artifacts, dispatch namespaces) are assumed to need none beyond Workers
304Scripts. Verify on the first run with `workflow_dispatch` and `units:
305pages` (small, no secrets), then `units: og` (R2) and `units: runner`
306(Containers); a missing permission fails with `Authentication error [code:
30710000]` and the route it was refused.
308
309Add it to the repository:
310
3111. On g1t.sh, open **flagon-io/g1t → Settings → Secrets and variables**.
3122. **Add**: key `CLOUDFLARE_API_TOKEN`, type **Secret**, available to
313 **Workflows**, environment **Production**. Only jobs with
314 `environment: production` (the deploy jobs) can read it.
3153. **Add**: key `CLOUDFLARE_ACCOUNT_ID`, type **Variable**, value
316 `1e6f2cffa3f445920836e8ebe446bb58`, available to **Workflows**, all
317 environments.
318
319Or through the API:
320
321```sh
322curl -X PUT https://api.g1t.sh/repos/flagon-io/g1t/actions/secrets/CLOUDFLARE_API_TOKEN \
323 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
324 -d '{"value":"<the token>","environments":["production"],"available_to":["workflows"]}'
325
326curl -X POST https://api.g1t.sh/repos/flagon-io/g1t/actions/variables \
327 -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
328 -d '{"name":"CLOUDFLARE_ACCOUNT_ID","value":"1e6f2cffa3f445920836e8ebe446bb58","available_to":["workflows"]}'
329```
330
331## Turning it on
332
3331. Create the token and add the secret and variable (above).
3342. Add `api.cloudflare.com` to flagon-io/g1t's allowed domains.
3353. Adopt the live Workers once, from a laptop: deploy everything with the
336 tool so each version records its commit (`scripts/deploy.sh`, or
337 `node scripts/deploy.mjs deploy --all`). Until then every plan says "no
338 known commit" and deploys every unit. To see what has changed since a
339 commit you know production runs, without deploying:
340 `node scripts/deploy.mjs plan --since <sha>`.
3414. Run the workflow by hand with `dry_run`, then with `units: pages`.
342
343## First deploy of a new account
344
345What the configs refer to must exist first. `node scripts/deploy.mjs
346manifest --json` lists, per unit, its queues, KV, R2, Vectorize and
347dispatch namespaces; each unit's `setup` and `secrets` say the rest.
348
349- D1: `npx wrangler d1 create <database>`, then put its id in the unit's
350 `wrangler.jsonc`.
351- KV: `npx wrangler kv namespace create <name>` for each name under
352 `resources.kv`, then the ids in the configs.
353- Queues: `npx wrangler queues create <queue>` for each queue in the
354 manifest: `g1t-events`, `g1t-events-<service>` for every subscriber,
355 `g1t-search-jobs`, `g1t-context-jobs`.
356- R2: `npx wrangler r2 bucket create g1t-screenshots`.
357- Vectorize, dispatch namespace, DNS, Access, Email Sending, Artifacts: each
358 unit's `setup`.
359- Secrets: `npx wrangler secret put <NAME>` in the unit's folder;
360 `node scripts/deploy.mjs doctor` lists what is missing.
361
362Then `node scripts/deploy.mjs deploy --all`. A Worker bound to a service
363that does not exist yet may be refused; deploy that service first with
364`--only`.
365
366## Rolling back
367
368- **One unit, at once:** `npx wrangler rollback` in its folder (or
369 `npx wrangler rollback <version-id>`; `npx wrangler versions list` shows
370 each version's commit in its message). Code only: D1 migrations are not
371 undone. The next plan sees the older commit and deploys what changed since,
372 so revert the commit on `main` too, or the next push brings it back.
373- **To a commit:** check it out and `node scripts/deploy.mjs deploy --only
374 <units> --force`. Migrations never run backwards: a migration that needs
375 undoing is a new migration.
376- **The runner's image:** a rollback of the Worker does not roll back the
377 container image; redeploy the older commit with `--rebuild-image` on a
378 machine with Docker.
379
380## Adding a unit
381
3821. Make its folder with a `wrangler.jsonc` (and its D1 migrations, if any).
383 A Rust Worker is a workspace member in the root `Cargo.toml` with
384 `"build": { "command": "node ../../scripts/build-rust-worker.mjs" }` and
385 the `wasm-opt = ["-O1"]` metadata; a TypeScript one is an npm workspace.
3862. Add it to `deploy/stack.jsonc`: path, kind, worker, `d1`, stage (the
387 earliest stage after everything it binds to), secrets, setup, self_host.
388 Name any new KV id under `resources.kv`.
3893. If its sources import a file outside its folder that is not a workspace
390 crate or package, list it under `inputs`.
3914. Add a row to the service table in `docs/SELF_HOSTING.md`.
3925. `npm run test:deploy` and `node scripts/deploy.mjs manifest --check`
393 say what is missing. Then create its resources and secrets, and
394 `node scripts/deploy.mjs deploy --only <unit>`.