Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.
| Merge platform pause and the hourly usage watcher: staff can pause compute, schedules, indexing or renders for everyone, the watcher emails on a breach and is never blind quietly, and the models proxy holds each run to its cap (billing 0051, integrations 0006) | 1 | # Spend guardrails |
| 2 | ||
| 3 | Cloudflare has no hard spending cap. A loop in a Worker, a queue that | |
| 4 | retries forever or a cron that lists KV every second shows up on the bill | |
| 5 | weeks later unless something here catches it. This is what catches it, | |
| 6 | how to stop it, and what the owner has to set up by hand in Cloudflare's | |
| 7 | dashboard. Internal: the public side is the "When g1t pauses work for | |
| 8 | everyone" section of `apps/docs/src/content/docs/guides/usage-and-billing.md` | |
| 9 | and "How they are enforced" in `guides/guardrails.md`. | |
| 10 | ||
| 11 | ## What catches what | |
| 12 | ||
| 13 | | Guardrail | Catches | Where | | |
| 14 | | --- | --- | --- | | |
| 15 | | A workspace's limits, caps and spike pause | One workspace's agents, sandboxes and builds running away | `services/billing/src/limits.rs`, `compute.rs`; reserved through `ComputeGate.admit` (`packages/contracts/src/compute.ts`) | | |
| 16 | | The comped budget and the daily breaker | g1t's own spend on agents: comped accounts, trials, pools; $75 a day pauses hosted-model agent runs g1t pays for | `services/billing/src/budget.rs`; [BILLING_OPERATIONS.md](BILLING_OPERATIONS.md) | | |
| 17 | | The daily reconciliation | What Cloudflare billed against what g1t counted, a day later | `services/billing/src/costs.rs`, `margin.rs` | | |
| 18 | | **The hourly platform watch** | Platform cost no workspace's limit covers: Workers requests and CPU, D1 rows, Queue operations, Durable Objects, KV, Artifacts. Within the hour. | `services/billing/src/platform.rs` | | |
| 19 | | **The platform pause** | A staff (or automatic) brake on whole kinds of work across g1t | `platform.rs`; read by every service through `g1t_kit::pause` (Rust) or `platformPaused` (`packages/contracts/src/platform.ts`) | | |
| 20 | | **The models proxy's run cap** | An agent spending past its run's cap by calling the model proxy itself (a prompt-injected `curl`), around the sandbox's `--max-budget-usd` | `services/models/src/spend.ts`, `run-spend.ts` | | |
| 21 | | Cloudflare's own notifications | Anything else, by product, as a last line | Set up by hand: [the checklist](#manual-steps-in-cloudflares-dashboard) | | |
| 22 | ||
| 23 | ## The hourly platform watch | |
| 24 | ||
| 25 | At a quarter past each hour (billing's `*/15` cron, gated to the tick at | |
| 26 | `:15`), billing reads two windows from Cloudflare's GraphQL Analytics API: | |
| 27 | the hour before, and the month so far. It uses the costs reconciliation's | |
| 28 | token, `CLOUDFLARE_BILLING_TOKEN` (else `CLOUDFLARE_USAGE_TOKEN`), which | |
| 29 | needs **Account Analytics Read**. Without either it reads nothing and sudo | |
| 30 | says so; the pause still works. | |
| 31 | ||
| 32 | | Metric | Dataset and field | Named by | | |
| 33 | | --- | --- | --- | | |
| 34 | | `workers_requests` | `workersInvocationsAdaptive` `sum.requests` | script | | |
| 35 | | `workers_cpu_ms` | `workersInvocationsAdaptive` `sum.cpuTimeUs` / 1000 | script | | |
| 36 | | `d1_rows_read`, `d1_rows_written` | `d1AnalyticsAdaptiveGroups` `sum.rowsRead`, `sum.rowsWritten` | database id | | |
| 37 | | `queue_operations` | `queueMessageOperationsAdaptiveGroups` `sum.billableOperations` | queue id | | |
| 38 | | `do_requests` | `durableObjectsInvocationsAdaptiveGroups` `sum.requests` | script | | |
| 39 | | `do_active_seconds`, `do_storage_write_units` | `durableObjectsPeriodicGroups` `sum.activeTime` (µs), `sum.storageWriteUnits` | namespace id | | |
| 40 | | `do_rows_written` | `durableObjectsPeriodicGroups` `sum.rowsWritten` | namespace id | | |
| 41 | | `kv_reads`, `kv_writes`, `kv_deletes`, `kv_lists` | `kvOperationsAdaptiveGroups` `sum.requests` by `actionType` | namespace id | | |
| 42 | | `artifacts_events` | `artifactsEventsAdaptiveGroups` `count` | repository | | |
| 43 | ||
| 44 | Each metric whose field is less certain has a query of its own, so a | |
| 45 | dataset or field GraphQL refuses is skipped and the others still count. | |
| 46 | Workers Logs has no dataset here; its volume follows Workers requests, and | |
| 47 | log sampling is set per Worker. | |
| 48 | ||
| 49 | Each hour is kept in billing's D1 (`platform_usage`, migration 0051) with | |
| 50 | the script, queue, database or namespace that counted most; the month so | |
| 51 | far in `platform_usage_month`; breaches in `platform_alerts`. | |
| 52 | ||
| 53 | ### When the watcher cannot see | |
| 54 | ||
| 55 | A skipped query leaves the watcher blind on its metrics, which is the | |
| 56 | failure this guards against, so it is never quiet: | |
| 57 | ||
| 58 | - Every run records the queries that failed, with Cloudflare's error, and | |
| 59 | whether every dataset answered with no rows (`platform_watch_runs`). | |
| 60 | - sudo's **Platform pause** shows **The watcher can't see: …** with each | |
| 61 | error whenever the latest run skipped anything, and a warning when every | |
| 62 | dataset was empty. | |
| 63 | - A query failing **3 hourly runs in a row** emails staff through the same | |
| 64 | alert path, at most once a day per query (`platform_watch_alerts`). | |
| 65 | - **Every dataset answering with no rows** for 3 runs in a row, the month so | |
| 66 | far included, is emailed the same way (key `all_empty`): almost certainly | |
| 67 | the wrong `CLOUDFLARE_ACCOUNT_ID` or a token without Account Analytics Read. | |
| 68 | ||
| 69 | These field names are from Cloudflare's documentation and were **not | |
| 70 | checked against the live schema** when the watcher was written: | |
| 71 | ||
| 72 | | Query | Unverified | | |
| 73 | | --- | --- | | |
| 74 | | `workers_cpu` | `workersInvocationsAdaptive` `sum.cpuTimeUs` | | |
| 75 | | `do_periodic`, `do_sql` | `durableObjectsPeriodicGroups` with a `datetime_geq` / `datetime_lt` filter, and its `sum.activeTime`, `sum.storageWriteUnits` and `sum.rowsWritten` | | |
| 76 | | `d1` | `d1AnalyticsAdaptiveGroups` with a `datetimeHour_geq` / `datetimeHour_lt` filter | | |
| 77 | ||
| 78 | **Check them with one command** after any deploy that touches the queries: | |
| 79 | ||
| 80 | ```sh | |
| 81 | CLOUDFLARE_API_TOKEN=<Account Analytics Read> node scripts/ops/platform-usage.mjs | |
| 82 | ``` | |
| 83 | ||
| 84 | It runs billing's own queries (`WATCHER_QUERIES`, kept identical to | |
| 85 | `QUERIES` in `platform.rs` by its test) over the last full hour. Every | |
| 86 | dataset should report rows. It exits 1 and names every dataset that | |
| 87 | errored, fell back to fewer fields, or answered empty. A renamed field is | |
| 88 | fixed in `QUERIES` in `services/billing/src/platform.rs` and in | |
| 89 | `WATCHER_QUERIES` together. | |
| 90 | ||
| 91 | ### Thresholds | |
| 92 | ||
| 93 | Hourly, in `services/billing/wrangler.jsonc` `vars`. Each is about a dollar | |
| 94 | to a few dollars an hour at list prices, far above a small alpha's normal | |
| 95 | hour. `0` turns a metric's threshold (and its spike rule) off. | |
| 96 | ||
| 97 | | Variable | Default | About | | |
| 98 | | --- | --- | --- | | |
| 99 | | `PLATFORM_HOURLY_WORKERS_REQUESTS` | 20,000,000 | $6/hour | | |
| 100 | | `PLATFORM_HOURLY_WORKERS_CPU_MS` | 100,000,000 | $2/hour | | |
| 101 | | `PLATFORM_HOURLY_D1_ROWS_READ` | 2,000,000,000 | $2/hour | | |
| 102 | | `PLATFORM_HOURLY_D1_ROWS_WRITTEN` | 5,000,000 | $5/hour | | |
| 103 | | `PLATFORM_HOURLY_QUEUE_OPERATIONS` | 5,000,000 | $2/hour | | |
| 104 | | `PLATFORM_HOURLY_DO_REQUESTS` | 20,000,000 | $3/hour | | |
| 105 | | `PLATFORM_HOURLY_DO_ROWS_WRITTEN` | 5,000,000 | $5/hour | | |
| 106 | | `PLATFORM_HOURLY_DO_STORAGE_WRITE_UNITS` | 5,000,000 | $5/hour | | |
| 107 | | `PLATFORM_HOURLY_DO_ACTIVE_SECONDS` | 3,000,000 | about $5/hour at 128 MB | | |
| 108 | | `PLATFORM_HOURLY_KV_READS` | 10,000,000 | $5/hour | | |
| 109 | | `PLATFORM_HOURLY_KV_WRITES` | 200,000 | $1/hour | | |
| 110 | | `PLATFORM_HOURLY_KV_DELETES` | 200,000 | $1/hour | | |
| 111 | | `PLATFORM_HOURLY_KV_LISTS` | 200,000 | $1/hour | | |
| 112 | | `PLATFORM_HOURLY_ARTIFACTS_EVENTS` | 1,000,000 | | | |
| 113 | ||
| 114 | The rules, in order: | |
| 115 | ||
| 116 | 1. **Threshold**: an hour over its threshold is a breach. | |
| 117 | 2. **Spike**: otherwise, an hour over `PLATFORM_SPIKE_FACTOR` (10) times the | |
| 118 | median hour of the week before, and at least `PLATFORM_SPIKE_FLOOR_PERCENT` | |
| 119 | (10%) of its threshold, is a breach. It needs a day of history first. | |
| 120 | 3. **Severe**: a threshold breach at `PLATFORM_SEVERE_FACTOR` (5) times the | |
| 121 | threshold or more pauses the levels that metric feeds, of those | |
| 122 | `AUTO_PAUSE` names (default `schedules,indexing`; `compute` and `renders` | |
| 123 | are off unless added). Spikes never pause. | |
| 124 | ||
| 125 | | Metric | Feeds | | |
| 126 | | --- | --- | | |
| 127 | | Workers | schedules, indexing, renders | | |
| 128 | | D1, Queues | schedules, indexing | | |
| 129 | | Durable Objects, Artifacts | compute, schedules | | |
| 130 | | KV | indexing, renders | | |
| 131 | ||
| 132 | On a breach, staff are emailed at `COSTS_ALERT_EMAIL` through billing's | |
| 133 | `EMAIL` binding, once per metric every 6 hours (a new automatic pause is | |
| 134 | always emailed), and sudo's Costs & margin shows it under **Platform | |
| 135 | pause**. The alert names the top script, queue, database or namespace. | |
| 136 | Ids are Cloudflare's; `node scripts/ops/platform-usage.mjs` names them. | |
| 137 | ||
| 138 | To change a threshold: edit the variable, push, and billing redeploys. The | |
| 139 | next hour reads with it. | |
| 140 | ||
| 141 | ### Looking by hand | |
| 142 | ||
| 143 | ```sh | |
| 144 | node scripts/ops/platform-usage.mjs # month so far and the last 24 hours | |
| 145 | node scripts/ops/platform-usage.mjs --json # the same, as JSON | |
| 146 | ``` | |
| 147 | ||
| 148 | It needs `CLOUDFLARE_API_TOKEN` (or `CLOUDFLARE_API_KEY` and `CLOUDFLARE_EMAIL`) with Account | |
| 149 | Analytics Read, as `scripts/deploy/cloudflare.mjs` `cloudflareAuth` reads them, and names ids | |
| 150 | with the D1, Queues, KV and Durable Objects listings when the token can | |
| 151 | read them. | |
| 152 | ||
| 153 | ## The platform pause | |
| 154 | ||
| 155 | Four levels, each independent, kept in billing's `platform_pause` table. | |
| 156 | ||
| 157 | | Level | Stops | Where it is checked | | |
| 158 | | --- | --- | --- | | |
| 159 | | `compute` | Every reservation through billing's `reserve` except embeddings: agent runs, checks, the merge queue, workflow jobs on g1t's machines, deploy builds. For every workspace and plan, and where payments are not set up. Refused with `paused`. | `services/billing/src/compute.rs` `reserve` | | |
| 160 | | `schedules` | Actions' cron-triggered runs (skipped, not run late), and the runner's sweep that starts queued agents (they stay queued) | `services/actions/src/plan.rs` `on_minute`; `services/runner/src/index.ts` `scheduled` | | |
| 161 | | `indexing` | Embeddings (refused in `reserve`), context backfills (**Rebuild** refused, queued jobs closed with a note), search backfills (pages parked in search's `meta` and resumed where they were) | `reserve`; `services/context/src/index.ts`; `services/search/src/lib.rs` | | |
| 162 | | `renders` | Social cards: a cache miss gets the cached brand card or a redirect to `https://g1t.sh/brand/g1t-logo-on-dark.png`, kept a minute | `services/og/src/index.ts`, `paused.ts` | | |
| 163 | ||
| 164 | Work already running finishes at every level. | |
| 165 | ||
| 166 | **Reads are cheap.** Every caller keeps the flags 30 seconds in its isolate: | |
| 167 | billing itself (`pause_now`), Rust services through `g1t_kit::pause`, and | |
| 168 | TypeScript services through `platformPaused`. A change reaches everything | |
| 169 | within about 30 seconds, and there is never a D1 read per request. | |
| 170 | ||
| 171 | **When the flag cannot be read, nothing is paused** (fail open), and the | |
| 172 | failure is kept for the same 30 seconds. A pause is a brake someone pulls | |
| 173 | on purpose; failing closed would turn a billing outage into a platform | |
| 174 | outage. The other guardrails (workspace limits, the breaker, the model | |
| 175 | proxy's run cap) do not depend on it. | |
| 176 | ||
| 177 | ### Pausing and resuming | |
| 178 | ||
| 179 | 1. Open sudo, **Costs & margin**, **Platform pause** (`https://sudo.g1t.sh/costs#platform`). | |
| 180 | 2. On the level's card, write why, and choose **Pause** or **Resume**. | |
| 181 | ||
| 182 | Every change is in the audit log (`platform_paused`, `platform_resumed`), | |
| 183 | with who and why. While any level is paused, every sudo page shows a red | |
| 184 | **Platform pause** bar. A level the usage watcher paused says so and stays | |
| 185 | paused until staff resume it: fix or understand the cause first. | |
| 186 | ||
| 187 | Without sudo (billing's RPC, through a service binding): | |
| 188 | `admin_set_pause` with `{ "level": "schedules", "paused": false, "note": "…", "by": "you@flagon.io" }`. | |
| 189 | ||
| 190 | ## The models proxy's run cap | |
| 191 | ||
| 192 | A run's model cap was only enforced inside the sandbox | |
| 193 | (`--max-budget-usd`). Now the proxy holds it too. When a run starts, the | |
| 194 | runner gives its model session token (`g1tm_`) the run's cap | |
| 195 | (`cap_model_sessions` on integrations). The proxy counts each answer's cost | |
| 196 | from its token usage in one `RunSpend` Durable Object per session, so | |
| 197 | requests fanned out across isolates cannot each spend the cap. At the cap it | |
| 198 | answers `402` with `run_cap_reached`. At most 16 answers are counted in | |
| 199 | flight at once, which bounds the overshoot. A session without a cap gets a | |
| 200 | $100 backstop. A run token reaches only the message and model routes, and is | |
| 201 | closed when the run ends. | |
| 202 | ||
| 203 | ## Manual steps in Cloudflare's dashboard | |
| 204 | ||
| 205 | These are the owner's, once per account. None can be set from code. | |
| 206 | ||
| 207 | - [ ] **Billing → Billable Usage notifications** (Manage Account → Billing → | |
| 208 | Notifications, or Notifications → Add → "Usage Based Billing"). Add | |
| 209 | one per product g1t uses: Workers (requests and CPU), Workers KV, D1, | |
| 210 | Queues, Durable Objects, R2, Workers Logs, Containers, Browser | |
| 211 | Rendering, Workers AI and Vectorize. Set each threshold near this | |
| 212 | doc's hourly threshold times about 24 times 3 (a day at a third of the | |
| 213 | watch's line), and send them to hey@flagon.io. | |
| 214 | - [ ] **Budget alerts** (Manage Account → Billing → Budget alerts, where the | |
| 215 | account has them): one for the whole account's monthly usage, at the | |
| 216 | month's expected bill and at twice it. | |
| 217 | - [ ] **Notifications → Destinations**: add hey@flagon.io (and a webhook, | |
| 218 | if one is set up for paging) so the alerts above reach someone. | |
| 219 | - [ ] **Account API token for billing**: check `CLOUDFLARE_BILLING_TOKEN` | |
| 220 | (or `CLOUDFLARE_USAGE_TOKEN`) has Account Analytics Read. sudo's | |
| 221 | Platform pause says when the watch cannot read. | |
| 222 | - [ ] Where to see them: Notifications → History for what fired; Billing → | |
| 223 | Billable Usage for the month so far by product. | |
| 224 | ||
| 225 | ## Deploy order for these changes | |
| 226 | ||
| 227 | 1. Billing (migration 0051 runs first, then the Worker): the pause, the | |
| 228 | watch, `platform_pause`, `admin_platform_guard`, `admin_set_pause`. | |
| 229 | 2. Search and og, with their new `BILLING` binding; actions, context and | |
| 230 | runner. Before billing has `platform_pause`, they read nothing paused. | |
| 231 | 3. Sudo. | |
| 232 | 4. For the run cap: integrations (migration 0006), then models (its | |
| 233 | `RunSpend` Durable Object migration), then runner. Any order works; a | |
| 234 | session without a cap gets the $100 backstop. |
This file's history is long; its oldest lines are credited to the oldest commit read.