Skip to content
234 linesCodeBlameRaw
1# Spend guardrails
2
3Cloudflare has no hard spending cap. A loop in a Worker, a queue that
4retries forever or a cron that lists KV every second shows up on the bill
5weeks later unless something here catches it. This is what catches it,
6how to stop it, and what the owner has to set up by hand in Cloudflare's
7dashboard. Internal: the public side is the "When g1t pauses work for
8everyone" section of `apps/docs/src/content/docs/guides/usage-and-billing.md`
9and "How they are enforced" in `guides/guardrails.md`.
10
11## What catches what
12
13| Guardrail | Catches | Where |
14| --- | --- | --- |
15| A workspace's limits, caps and spike pause | One workspace's agents, sandboxes and builds running away | `services/billing/src/limits.rs`, `compute.rs`; reserved through `ComputeGate.admit` (`packages/contracts/src/compute.ts`) |
16| The comped budget and the daily breaker | g1t's own spend on agents: comped accounts, trials, pools; $75 a day pauses hosted-model agent runs g1t pays for | `services/billing/src/budget.rs`; [BILLING_OPERATIONS.md](BILLING_OPERATIONS.md) |
17| The daily reconciliation | What Cloudflare billed against what g1t counted, a day later | `services/billing/src/costs.rs`, `margin.rs` |
18| **The hourly platform watch** | Platform cost no workspace's limit covers: Workers requests and CPU, D1 rows, Queue operations, Durable Objects, KV, Artifacts. Within the hour. | `services/billing/src/platform.rs` |
19| **The platform pause** | A staff (or automatic) brake on whole kinds of work across g1t | `platform.rs`; read by every service through `g1t_kit::pause` (Rust) or `platformPaused` (`packages/contracts/src/platform.ts`) |
20| **The models proxy's run cap** | An agent spending past its run's cap by calling the model proxy itself (a prompt-injected `curl`), around the sandbox's `--max-budget-usd` | `services/models/src/spend.ts`, `run-spend.ts` |
21| Cloudflare's own notifications | Anything else, by product, as a last line | Set up by hand: [the checklist](#manual-steps-in-cloudflares-dashboard) |
22
23## The hourly platform watch
24
25At a quarter past each hour (billing's `*/15` cron, gated to the tick at
26`:15`), billing reads two windows from Cloudflare's GraphQL Analytics API:
27the hour before, and the month so far. It uses the costs reconciliation's
28token, `CLOUDFLARE_BILLING_TOKEN` (else `CLOUDFLARE_USAGE_TOKEN`), which
29needs **Account Analytics Read**. Without either it reads nothing and sudo
30says so; the pause still works.
31
32| Metric | Dataset and field | Named by |
33| --- | --- | --- |
34| `workers_requests` | `workersInvocationsAdaptive` `sum.requests` | script |
35| `workers_cpu_ms` | `workersInvocationsAdaptive` `sum.cpuTimeUs` / 1000 | script |
36| `d1_rows_read`, `d1_rows_written` | `d1AnalyticsAdaptiveGroups` `sum.rowsRead`, `sum.rowsWritten` | database id |
37| `queue_operations` | `queueMessageOperationsAdaptiveGroups` `sum.billableOperations` | queue id |
38| `do_requests` | `durableObjectsInvocationsAdaptiveGroups` `sum.requests` | script |
39| `do_active_seconds`, `do_storage_write_units` | `durableObjectsPeriodicGroups` `sum.activeTime` (µs), `sum.storageWriteUnits` | namespace id |
40| `do_rows_written` | `durableObjectsPeriodicGroups` `sum.rowsWritten` | namespace id |
41| `kv_reads`, `kv_writes`, `kv_deletes`, `kv_lists` | `kvOperationsAdaptiveGroups` `sum.requests` by `actionType` | namespace id |
42| `artifacts_events` | `artifactsEventsAdaptiveGroups` `count` | repository |
43
44Each metric whose field is less certain has a query of its own, so a
45dataset or field GraphQL refuses is skipped and the others still count.
46Workers Logs has no dataset here; its volume follows Workers requests, and
47log sampling is set per Worker.
48
49Each hour is kept in billing's D1 (`platform_usage`, migration 0051) with
50the script, queue, database or namespace that counted most; the month so
51far in `platform_usage_month`; breaches in `platform_alerts`.
52
53### When the watcher cannot see
54
55A skipped query leaves the watcher blind on its metrics, which is the
56failure this guards against, so it is never quiet:
57
58- Every run records the queries that failed, with Cloudflare's error, and
59 whether every dataset answered with no rows (`platform_watch_runs`).
60- sudo's **Platform pause** shows **The watcher can't see: …** with each
61 error whenever the latest run skipped anything, and a warning when every
62 dataset was empty.
63- A query failing **3 hourly runs in a row** emails staff through the same
64 alert path, at most once a day per query (`platform_watch_alerts`).
65- **Every dataset answering with no rows** for 3 runs in a row, the month so
66 far included, is emailed the same way (key `all_empty`): almost certainly
67 the wrong `CLOUDFLARE_ACCOUNT_ID` or a token without Account Analytics Read.
68
69These field names are from Cloudflare's documentation and were **not
70checked against the live schema** when the watcher was written:
71
72| Query | Unverified |
73| --- | --- |
74| `workers_cpu` | `workersInvocationsAdaptive` `sum.cpuTimeUs` |
75| `do_periodic`, `do_sql` | `durableObjectsPeriodicGroups` with a `datetime_geq` / `datetime_lt` filter, and its `sum.activeTime`, `sum.storageWriteUnits` and `sum.rowsWritten` |
76| `d1` | `d1AnalyticsAdaptiveGroups` with a `datetimeHour_geq` / `datetimeHour_lt` filter |
77
78**Check them with one command** after any deploy that touches the queries:
79
80```sh
81CLOUDFLARE_API_TOKEN=<Account Analytics Read> node scripts/ops/platform-usage.mjs
82```
83
84It runs billing's own queries (`WATCHER_QUERIES`, kept identical to
85`QUERIES` in `platform.rs` by its test) over the last full hour. Every
86dataset should report rows. It exits 1 and names every dataset that
87errored, fell back to fewer fields, or answered empty. A renamed field is
88fixed in `QUERIES` in `services/billing/src/platform.rs` and in
89`WATCHER_QUERIES` together.
90
91### Thresholds
92
93Hourly, in `services/billing/wrangler.jsonc` `vars`. Each is about a dollar
94to a few dollars an hour at list prices, far above a small alpha's normal
95hour. `0` turns a metric's threshold (and its spike rule) off.
96
97| Variable | Default | About |
98| --- | --- | --- |
99| `PLATFORM_HOURLY_WORKERS_REQUESTS` | 20,000,000 | $6/hour |
100| `PLATFORM_HOURLY_WORKERS_CPU_MS` | 100,000,000 | $2/hour |
101| `PLATFORM_HOURLY_D1_ROWS_READ` | 2,000,000,000 | $2/hour |
102| `PLATFORM_HOURLY_D1_ROWS_WRITTEN` | 5,000,000 | $5/hour |
103| `PLATFORM_HOURLY_QUEUE_OPERATIONS` | 5,000,000 | $2/hour |
104| `PLATFORM_HOURLY_DO_REQUESTS` | 20,000,000 | $3/hour |
105| `PLATFORM_HOURLY_DO_ROWS_WRITTEN` | 5,000,000 | $5/hour |
106| `PLATFORM_HOURLY_DO_STORAGE_WRITE_UNITS` | 5,000,000 | $5/hour |
107| `PLATFORM_HOURLY_DO_ACTIVE_SECONDS` | 3,000,000 | about $5/hour at 128 MB |
108| `PLATFORM_HOURLY_KV_READS` | 10,000,000 | $5/hour |
109| `PLATFORM_HOURLY_KV_WRITES` | 200,000 | $1/hour |
110| `PLATFORM_HOURLY_KV_DELETES` | 200,000 | $1/hour |
111| `PLATFORM_HOURLY_KV_LISTS` | 200,000 | $1/hour |
112| `PLATFORM_HOURLY_ARTIFACTS_EVENTS` | 1,000,000 | |
113
114The rules, in order:
115
1161. **Threshold**: an hour over its threshold is a breach.
1172. **Spike**: otherwise, an hour over `PLATFORM_SPIKE_FACTOR` (10) times the
118 median hour of the week before, and at least `PLATFORM_SPIKE_FLOOR_PERCENT`
119 (10%) of its threshold, is a breach. It needs a day of history first.
1203. **Severe**: a threshold breach at `PLATFORM_SEVERE_FACTOR` (5) times the
121 threshold or more pauses the levels that metric feeds, of those
122 `AUTO_PAUSE` names (default `schedules,indexing`; `compute` and `renders`
123 are off unless added). Spikes never pause.
124
125| Metric | Feeds |
126| --- | --- |
127| Workers | schedules, indexing, renders |
128| D1, Queues | schedules, indexing |
129| Durable Objects, Artifacts | compute, schedules |
130| KV | indexing, renders |
131
132On a breach, staff are emailed at `COSTS_ALERT_EMAIL` through billing's
133`EMAIL` binding, once per metric every 6 hours (a new automatic pause is
134always emailed), and sudo's Costs & margin shows it under **Platform
135pause**. The alert names the top script, queue, database or namespace.
136Ids are Cloudflare's; `node scripts/ops/platform-usage.mjs` names them.
137
138To change a threshold: edit the variable, push, and billing redeploys. The
139next hour reads with it.
140
141### Looking by hand
142
143```sh
144node scripts/ops/platform-usage.mjs # month so far and the last 24 hours
145node scripts/ops/platform-usage.mjs --json # the same, as JSON
146```
147
148It needs `CLOUDFLARE_API_TOKEN` (or `CLOUDFLARE_API_KEY` and `CLOUDFLARE_EMAIL`) with Account
149Analytics Read, as `scripts/deploy/cloudflare.mjs` `cloudflareAuth` reads them, and names ids
150with the D1, Queues, KV and Durable Objects listings when the token can
151read them.
152
153## The platform pause
154
155Four levels, each independent, kept in billing's `platform_pause` table.
156
157| Level | Stops | Where it is checked |
158| --- | --- | --- |
159| `compute` | Every reservation through billing's `reserve` except embeddings: agent runs, checks, the merge queue, workflow jobs on g1t's machines, deploy builds. For every workspace and plan, and where payments are not set up. Refused with `paused`. | `services/billing/src/compute.rs` `reserve` |
160| `schedules` | Actions' cron-triggered runs (skipped, not run late), and the runner's sweep that starts queued agents (they stay queued) | `services/actions/src/plan.rs` `on_minute`; `services/runner/src/index.ts` `scheduled` |
161| `indexing` | Embeddings (refused in `reserve`), context backfills (**Rebuild** refused, queued jobs closed with a note), search backfills (pages parked in search's `meta` and resumed where they were) | `reserve`; `services/context/src/index.ts`; `services/search/src/lib.rs` |
162| `renders` | Social cards: a cache miss gets the cached brand card or a redirect to `https://g1t.sh/brand/g1t-logo-on-dark.png`, kept a minute | `services/og/src/index.ts`, `paused.ts` |
163
164Work already running finishes at every level.
165
166**Reads are cheap.** Every caller keeps the flags 30 seconds in its isolate:
167billing itself (`pause_now`), Rust services through `g1t_kit::pause`, and
168TypeScript services through `platformPaused`. A change reaches everything
169within about 30 seconds, and there is never a D1 read per request.
170
171**When the flag cannot be read, nothing is paused** (fail open), and the
172failure is kept for the same 30 seconds. A pause is a brake someone pulls
173on purpose; failing closed would turn a billing outage into a platform
174outage. The other guardrails (workspace limits, the breaker, the model
175proxy's run cap) do not depend on it.
176
177### Pausing and resuming
178
1791. Open sudo, **Costs & margin**, **Platform pause** (`https://sudo.g1t.sh/costs#platform`).
1802. On the level's card, write why, and choose **Pause** or **Resume**.
181
182Every change is in the audit log (`platform_paused`, `platform_resumed`),
183with who and why. While any level is paused, every sudo page shows a red
184**Platform pause** bar. A level the usage watcher paused says so and stays
185paused until staff resume it: fix or understand the cause first.
186
187Without sudo (billing's RPC, through a service binding):
188`admin_set_pause` with `{ "level": "schedules", "paused": false, "note": "…", "by": "you@flagon.io" }`.
189
190## The models proxy's run cap
191
192A run's model cap was only enforced inside the sandbox
193(`--max-budget-usd`). Now the proxy holds it too. When a run starts, the
194runner gives its model session token (`g1tm_`) the run's cap
195(`cap_model_sessions` on integrations). The proxy counts each answer's cost
196from its token usage in one `RunSpend` Durable Object per session, so
197requests fanned out across isolates cannot each spend the cap. At the cap it
198answers `402` with `run_cap_reached`. At most 16 answers are counted in
199flight at once, which bounds the overshoot. A session without a cap gets a
200$100 backstop. A run token reaches only the message and model routes, and is
201closed when the run ends.
202
203## Manual steps in Cloudflare's dashboard
204
205These are the owner's, once per account. None can be set from code.
206
207- [ ] **Billing → Billable Usage notifications** (Manage Account → Billing →
208 Notifications, or Notifications → Add → "Usage Based Billing"). Add
209 one per product g1t uses: Workers (requests and CPU), Workers KV, D1,
210 Queues, Durable Objects, R2, Workers Logs, Containers, Browser
211 Rendering, Workers AI and Vectorize. Set each threshold near this
212 doc's hourly threshold times about 24 times 3 (a day at a third of the
213 watch's line), and send them to hey@flagon.io.
214- [ ] **Budget alerts** (Manage Account → Billing → Budget alerts, where the
215 account has them): one for the whole account's monthly usage, at the
216 month's expected bill and at twice it.
217- [ ] **Notifications → Destinations**: add hey@flagon.io (and a webhook,
218 if one is set up for paging) so the alerts above reach someone.
219- [ ] **Account API token for billing**: check `CLOUDFLARE_BILLING_TOKEN`
220 (or `CLOUDFLARE_USAGE_TOKEN`) has Account Analytics Read. sudo's
221 Platform pause says when the watch cannot read.
222- [ ] Where to see them: Notifications → History for what fired; Billing →
223 Billable Usage for the month so far by product.
224
225## Deploy order for these changes
226
2271. Billing (migration 0051 runs first, then the Worker): the pause, the
228 watch, `platform_pause`, `admin_platform_guard`, `admin_set_pause`.
2292. Search and og, with their new `BILLING` binding; actions, context and
230 runner. Before billing has `platform_pause`, they read nothing paused.
2313. Sudo.
2324. For the run cap: integrations (migration 0006), then models (its
233 `RunSpend` Durable Object migration), then runner. Any order works; a
234 session without a cap gets the $100 backstop.