Skip to content
196 linesCodeBlameRaw
1# Spend guardrails
2
3Cloudflare has no hard spending cap. A loop in a Worker, a queue that
4retries forever or a cron that lists KV every second shows up on the bill
5weeks later unless something here catches it. This is what catches it,
6how to stop it, and what the owner has to set up by hand in Cloudflare's
7dashboard. Internal: the public side is the "When g1t pauses work for
8everyone" section of `apps/docs/src/content/docs/guides/usage-and-billing.md`
9and "How they are enforced" in `guides/guardrails.md`.
10
11## What catches what
12
13| Guardrail | Catches | Where |
14| --- | --- | --- |
15| A workspace's limits, caps and spike pause | One workspace's agents, sandboxes and builds running away | `services/billing/src/limits.rs`, `compute.rs`; reserved through `ComputeGate.admit` (`packages/contracts/src/compute.ts`) |
16| The comped budget and the daily breaker | g1t's own spend on agents: comped accounts, trials, pools; $75 a day pauses hosted-model agent runs g1t pays for | `services/billing/src/budget.rs`; [BILLING_OPERATIONS.md](BILLING_OPERATIONS.md) |
17| The daily reconciliation | What Cloudflare billed against what g1t counted, a day later | `services/billing/src/costs.rs`, `margin.rs` |
18| **The hourly platform watch** | Platform cost no workspace's limit covers: Workers requests and CPU, D1 rows, Queue operations, Durable Objects, KV, Artifacts. Within the hour. | `services/billing/src/platform.rs` |
19| **The platform pause** | A staff (or automatic) brake on whole kinds of work across g1t | `platform.rs`; read by every service through `g1t_kit::pause` (Rust) or `platformPaused` (`packages/contracts/src/platform.ts`) |
20| **The models proxy's run cap** | An agent spending past its run's cap by calling the model proxy itself (a prompt-injected `curl`), around the sandbox's `--max-budget-usd` | `services/models/src/spend.ts`, `run-spend.ts` |
21| Cloudflare's own notifications | Anything else, by product, as a last line | Set up by hand: [the checklist](#manual-steps-in-cloudflares-dashboard) |
22
23## The hourly platform watch
24
25At a quarter past each hour (billing's `*/15` cron, gated to the tick at
26`:15`), billing reads two windows from Cloudflare's GraphQL Analytics API:
27the hour before, and the month so far. It uses the costs reconciliation's
28token, `CLOUDFLARE_BILLING_TOKEN` (else `CLOUDFLARE_USAGE_TOKEN`), which
29needs **Account Analytics Read**. Without either it reads nothing and sudo
30says so; the pause still works.
31
32| Metric | Dataset and field | Named by |
33| --- | --- | --- |
34| `workers_requests` | `workersInvocationsAdaptive` `sum.requests` | script |
35| `workers_cpu_ms` | `workersInvocationsAdaptive` `sum.cpuTimeUs` / 1000 | script |
36| `d1_rows_read`, `d1_rows_written` | `d1AnalyticsAdaptiveGroups` `sum.rowsRead`, `sum.rowsWritten` | database id |
37| `queue_operations` | `queueMessageOperationsAdaptiveGroups` `sum.billableOperations` | queue id |
38| `do_requests` | `durableObjectsInvocationsAdaptiveGroups` `sum.requests` | script |
39| `do_active_seconds`, `do_storage_write_units` | `durableObjectsPeriodicGroups` `sum.activeTime` (µs), `sum.storageWriteUnits` | namespace id |
40| `do_rows_written` | `durableObjectsPeriodicGroups` `sum.rowsWritten` | namespace id |
41| `kv_reads`, `kv_writes`, `kv_deletes`, `kv_lists` | `kvOperationsAdaptiveGroups` `sum.requests` by `actionType` | namespace id |
42| `artifacts_events` | `artifactsEventsAdaptiveGroups` `count` | repository |
43
44Each metric whose field is less certain has a query of its own, so a
45dataset or field GraphQL refuses is logged (`platform watch: skipped …`)
46and the others still count. Workers Logs has no dataset here; its volume
47follows Workers requests, and log sampling is set per Worker.
48
49Each hour is kept in billing's D1 (`platform_usage`, migration 0050) with
50the script, queue, database or namespace that counted most; the month so
51far in `platform_usage_month`; breaches in `platform_alerts`.
52
53### Thresholds
54
55Hourly, in `services/billing/wrangler.jsonc` `vars`. Each is about a dollar
56to a few dollars an hour at list prices, far above a small alpha's normal
57hour. `0` turns a metric's threshold (and its spike rule) off.
58
59| Variable | Default | About |
60| --- | --- | --- |
61| `PLATFORM_HOURLY_WORKERS_REQUESTS` | 20,000,000 | $6/hour |
62| `PLATFORM_HOURLY_WORKERS_CPU_MS` | 100,000,000 | $2/hour |
63| `PLATFORM_HOURLY_D1_ROWS_READ` | 2,000,000,000 | $2/hour |
64| `PLATFORM_HOURLY_D1_ROWS_WRITTEN` | 5,000,000 | $5/hour |
65| `PLATFORM_HOURLY_QUEUE_OPERATIONS` | 5,000,000 | $2/hour |
66| `PLATFORM_HOURLY_DO_REQUESTS` | 20,000,000 | $3/hour |
67| `PLATFORM_HOURLY_DO_ROWS_WRITTEN` | 5,000,000 | $5/hour |
68| `PLATFORM_HOURLY_DO_STORAGE_WRITE_UNITS` | 5,000,000 | $5/hour |
69| `PLATFORM_HOURLY_DO_ACTIVE_SECONDS` | 3,000,000 | about $5/hour at 128 MB |
70| `PLATFORM_HOURLY_KV_READS` | 10,000,000 | $5/hour |
71| `PLATFORM_HOURLY_KV_WRITES` | 200,000 | $1/hour |
72| `PLATFORM_HOURLY_KV_DELETES` | 200,000 | $1/hour |
73| `PLATFORM_HOURLY_KV_LISTS` | 200,000 | $1/hour |
74| `PLATFORM_HOURLY_ARTIFACTS_EVENTS` | 1,000,000 | |
75
76The rules, in order:
77
781. **Threshold**: an hour over its threshold is a breach.
792. **Spike**: otherwise, an hour over `PLATFORM_SPIKE_FACTOR` (10) times the
80 median hour of the week before, and at least `PLATFORM_SPIKE_FLOOR_PERCENT`
81 (10%) of its threshold, is a breach. It needs a day of history first.
823. **Severe**: a threshold breach at `PLATFORM_SEVERE_FACTOR` (5) times the
83 threshold or more pauses the levels that metric feeds, of those
84 `AUTO_PAUSE` names (default `schedules,indexing`; `compute` and `renders`
85 are off unless added). Spikes never pause.
86
87| Metric | Feeds |
88| --- | --- |
89| Workers | schedules, indexing, renders |
90| D1, Queues | schedules, indexing |
91| Durable Objects, Artifacts | compute, schedules |
92| KV | indexing, renders |
93
94On a breach, staff are emailed at `COSTS_ALERT_EMAIL` through billing's
95`EMAIL` binding, once per metric every 6 hours (a new automatic pause is
96always emailed), and sudo's Costs & margin shows it under **Platform
97pause**. The alert names the top script, queue, database or namespace.
98Ids are Cloudflare's; `node scripts/ops/platform-usage.mjs` names them.
99
100To change a threshold: edit the variable, push, and billing redeploys. The
101next hour reads with it.
102
103### Looking by hand
104
105```sh
106node scripts/ops/platform-usage.mjs # month so far and the last 24 hours
107node scripts/ops/platform-usage.mjs --json # the same, as JSON
108```
109
110It needs `CLOUDFLARE_API_TOKEN` (or `CLOUDFLARE_API_KEY` and `CLOUDFLARE_EMAIL`) with Account
111Analytics Read, as `scripts/deploy/cloudflare.mjs` `cloudflareAuth` reads them, and names ids
112with the D1, Queues, KV and Durable Objects listings when the token can
113read them.
114
115## The platform pause
116
117Four levels, each independent, kept in billing's `platform_pause` table.
118
119| Level | Stops | Where it is checked |
120| --- | --- | --- |
121| `compute` | Every reservation through billing's `reserve` except embeddings: agent runs, checks, the merge queue, workflow jobs on g1t's machines, deploy builds. For every workspace and plan, and where payments are not set up. Refused with `paused`. | `services/billing/src/compute.rs` `reserve` |
122| `schedules` | Actions' cron-triggered runs (skipped, not run late), and the runner's sweep that starts queued agents (they stay queued) | `services/actions/src/plan.rs` `on_minute`; `services/runner/src/index.ts` `scheduled` |
123| `indexing` | Embeddings (refused in `reserve`), context backfills (**Rebuild** refused, queued jobs closed with a note), search backfills (pages parked in search's `meta` and resumed where they were) | `reserve`; `services/context/src/index.ts`; `services/search/src/lib.rs` |
124| `renders` | Social cards: a cache miss gets the cached brand card or a redirect to `https://g1t.sh/brand/g1t-logo-on-dark.png`, kept a minute | `services/og/src/index.ts`, `paused.ts` |
125
126Work already running finishes at every level.
127
128**Reads are cheap.** Every caller keeps the flags 30 seconds in its isolate:
129billing itself (`pause_now`), Rust services through `g1t_kit::pause`, and
130TypeScript services through `platformPaused`. A change reaches everything
131within about 30 seconds, and there is never a D1 read per request.
132
133**When the flag cannot be read, nothing is paused** (fail open), and the
134failure is kept for the same 30 seconds. A pause is a brake someone pulls
135on purpose; failing closed would turn a billing outage into a platform
136outage. The other guardrails (workspace limits, the breaker, the model
137proxy's run cap) do not depend on it.
138
139### Pausing and resuming
140
1411. Open sudo, **Costs & margin**, **Platform pause** (`https://sudo.g1t.sh/costs#platform`).
1422. On the level's card, write why, and choose **Pause** or **Resume**.
143
144Every change is in the audit log (`platform_paused`, `platform_resumed`),
145with who and why. While any level is paused, every sudo page shows a red
146**Platform pause** bar. A level the usage watcher paused says so and stays
147paused until staff resume it: fix or understand the cause first.
148
149Without sudo (billing's RPC, through a service binding):
150`admin_set_pause` with `{ "level": "schedules", "paused": false, "note": "…", "by": "you@flagon.io" }`.
151
152## The models proxy's run cap
153
154A run's model cap was only enforced inside the sandbox
155(`--max-budget-usd`). Now the proxy holds it too. When a run starts, the
156runner gives its model session token (`g1tm_`) the run's cap
157(`cap_model_sessions` on integrations). The proxy counts each answer's cost
158from its token usage in one `RunSpend` Durable Object per session, so
159requests fanned out across isolates cannot each spend the cap. At the cap it
160answers `402` with `run_cap_reached`. At most 16 answers are counted in
161flight at once, which bounds the overshoot. A session without a cap gets a
162$100 backstop. A run token reaches only the message and model routes, and is
163closed when the run ends.
164
165## Manual steps in Cloudflare's dashboard
166
167These are the owner's, once per account. None can be set from code.
168
169- [ ] **Billing → Billable Usage notifications** (Manage Account → Billing →
170 Notifications, or Notifications → Add → "Usage Based Billing"). Add
171 one per product g1t uses: Workers (requests and CPU), Workers KV, D1,
172 Queues, Durable Objects, R2, Workers Logs, Containers, Browser
173 Rendering, Workers AI and Vectorize. Set each threshold near this
174 doc's hourly threshold times about 24 times 3 (a day at a third of the
175 watch's line), and send them to hey@flagon.io.
176- [ ] **Budget alerts** (Manage Account → Billing → Budget alerts, where the
177 account has them): one for the whole account's monthly usage, at the
178 month's expected bill and at twice it.
179- [ ] **Notifications → Destinations**: add hey@flagon.io (and a webhook,
180 if one is set up for paging) so the alerts above reach someone.
181- [ ] **Account API token for billing**: check `CLOUDFLARE_BILLING_TOKEN`
182 (or `CLOUDFLARE_USAGE_TOKEN`) has Account Analytics Read. sudo's
183 Platform pause says when the watch cannot read.
184- [ ] Where to see them: Notifications → History for what fired; Billing →
185 Billable Usage for the month so far by product.
186
187## Deploy order for these changes
188
1891. Billing (migration 0050 runs first, then the Worker): the pause, the
190 watch, `platform_pause`, `admin_platform_guard`, `admin_set_pause`.
1912. Search and og, with their new `BILLING` binding; actions, context and
192 runner. Before billing has `platform_pause`, they read nothing paused.
1933. Sudo.
1944. For the run cap: integrations (migration 0006), then models (its
195 `RunSpend` Durable Object migration), then runner. Any order works; a
196 session without a cap gets the $100 backstop.