Skip to content

Commit

Docs: Auto, choosing a model, and the agent rate on your own key

The models guide explains Auto, choosing a tier per kind of work, and what own-key runs are charged; Working with g1t has the full routing rules and the AGENT_ROUTING settings; Usage and billing gains the agent rate (how it is counted and weighted), its own-key line, and the statement's agent-rate rows. BILLING_OPERATIONS covers own-key metering, the cron that closes those runs, and how to set the token weights (cache reads at a tenth). PLAN records routing for cost, its measured savings, and the Workers AI route designed and left off. llms.txt follows.

syntaqxcommitted Parentf454eacBrowse files
7 files+306−850/7 viewed
+77−21
11 ---
22 title: Model providers
3−description: Connect Anthropic, OpenAI, Gemini or any compatible endpoint, choose which model does which work, and pay for it where you choose.
3+description: Let Auto choose the model for each job, or choose it yourself; connect Anthropic, OpenAI, Gemini or any compatible endpoint, and pay for models where you choose.
44 ---
55
66 Each workspace decides where its agents' model spend goes:
77
8−- **g1t's hosted models.** g1t chooses the model for each kind of work, pays
9− the provider, and charges your workspace what it cost plus 20%. The
10− plan's included usage and [the trial](/guides/usage-and-billing/#the-trial)
8+- **g1t's hosted models.** **Auto** chooses the model for each job, g1t pays
9+ the provider, and your workspace is charged the provider's price, with no
10+ markup, plus the [agent rate](/guides/usage-and-billing/#the-agent-rate).
11+ The plan's included usage and [the trial](/guides/usage-and-billing/#the-trial)
1112 pay for it first. While payments are in test mode, they are open only
1213 to a few invited workspaces, g1t's own among them; a card check or a
1314 trial does not open them. Every other workspace connects its own
1516 that says so. Once payments go live, they are open to all.
1617 - **Your own providers.** Connect as many as you use, then choose, for each
1718 kind of work, which provider and model it runs on. Each provider bills
18− you directly. Open to every workspace now.
19+ you for the model directly. Open to every workspace now.
1920
20−g1t's own routing is fixed; yours is not.
21+## Auto
2122
23+On g1t's models you do not have to pick a model. **Auto**, the default,
24+sends each job to the least costly model that can do it, from three tiers:
25+**fast** (Claude Haiku 4.5 today), **standard** (Claude Sonnet 5.5) and
26+**most capable** (Claude Opus 5.5). It decides by the kind of job, the
27+size of the change it reads, the issue's labels, whether the last attempt
28+at the same work failed, and what has worked in the repository before:
29+
30+- Catching up, answering a question and reviewing a small change that
31+ touches no sensitive path start on the fast model.
32+- Making and revising changes, planning, and most reviews start on the
33+ standard model.
34+- A review of a very large change, work on an issue labelled
35+ `architecture`, and work that failed twice in a row go to the most
36+ capable model. One failure moves the next attempt up one tier.
37+- When the cheaper model finished nearly all of a repository's recent runs
38+ of the same kind, Auto moves that work down a tier there; when a model
39+ keeps failing, up.
40+
41+Every run says which model it used and why, in one line on its run and in
42+its pull request's session, such as *Used a fast model (Claude Haiku 4.5):
43+small change, 3 files and 80 lines.* The full rules are in
44+[which model runs](/guides/working-with-g1t/#which-model-runs).
45+
2246 ## Providers
2347
2448 Labs and platforms need only a key; g1t knows where they are.
7599 | Catching up | Bringing a change up to date with `main`. |
76100
77101 Each can go to g1t's models, or to any of your providers on any of its
78−models. An Anthropic provider also offers **g1t's choice of Claude**,
79−which runs g1t's large-tier model, Claude Sonnet 5.5 today, on your key.
80−Routing between tiers by the size of the work is only for g1t's hosted
81−models; see [which model runs](/guides/working-with-g1t/#which-model-runs). For example: make
82−changes on Claude through your Anthropic key, review on GPT through your
83−OpenAI key, and catch up on a small model through OpenRouter.
102+models.
103+
104+- **On g1t's models**, choose **Auto** (the default), or a tier for every
105+ run of that kind: **Fast**, **Standard** or **Most capable**. A run on a
106+ chosen tier says the workspace chose it.
107+- **On an Anthropic provider**, leave the model empty for g1t's choice of
108+ Claude: Auto picks the tier's model for each job, as on g1t's models, on
109+ your key. Name a model to run that one every time.
110+- **On any other provider**, name the model.
111+
112+For example: make changes on Claude through your Anthropic key, review on
113+GPT through your OpenAI key, catch up on g1t's models on Fast, and plan on
114+Most capable.
84115
85116 **Save routing**, and the next runs use it. Without any routing, work goes
86117 to g1t's models where they are open to the workspace, and otherwise to the
90121
91122 ## What it costs
92123
93−Your providers bill you for the models. g1t charges only each run's
94−[sandbox time](/guides/usage-and-billing/#sandbox-time), at what it costs
95−g1t plus 20%, by the second. A change, a review, a revision, a catch-up
96−and a plan each run in a sandbox, and each sandbox is a line on the
97−statement. See [Usage and billing](/guides/usage-and-billing/).
124+On g1t's models, each run is charged the model at the provider's price
125+(what AI Gateway priced its requests at, with no markup), the agent rate
126+on its tokens, and its sandbox time. Auto keeps the first of those down:
127+the fast model costs about half what the standard one does, and the most
128+capable up to twice as much, so most of what it saves comes from sending small
129+jobs to the fast model and from finishing hard ones instead of retrying
130+them on the same model.
98131
99−That sandbox time counts toward the workspace's usage limit like any
100−other.
132+On your own providers, they bill you for the models. g1t charges each
133+run's [sandbox time](/guides/usage-and-billing/#sandbox-time), at what it
134+costs g1t plus 20%, by the second, and the
135+[agent rate](/guides/usage-and-billing/#the-agent-rate) on the tokens the
136+run used, from Oct 22, 2026: $0.25 per million, as on g1t's models. Tokens
137+are counted by g1t's model proxy as answers pass, and by the agent in the
138+sandbox; the more of the two is charged. On **Usage** it is the line
139+**Agent rate, your own model key**, with its tokens weighted as the
140+pricing page says.
141+
142+Both count toward the workspace's usage limit like any other charge.
143+**Usage** also shows the agent's tokens by model.
101144
102145 ## Your keys never reach a sandbox
103146
114157
115158 As each answer passes, the proxy reads how many tokens it used (input,
116159 output, and cache reads and writes) and counts them for the run, under the
117−person it was for. Those counts are for usage views; they never change what
118−a run is charged.
160+person it was for. They show on **Usage** by model, and the
161+[agent rate](/guides/usage-and-billing/#the-agent-rate) is charged on them.
162+On g1t's models, the model itself is charged at what AI Gateway priced it
163+at, never from these counts.
119164
120165 The token stops working within seconds of the run finishing, however it
121166 ends, and within seconds if you disconnect the provider. A run whose end
142187 `task` is `default`, `implement`, `review`, `plan` or `update`.
143188 `connection_id` is null for g1t's hosted models. `model` is null for the
144189 provider's default, or for an Anthropic provider, g1t's choice of Claude.
190+On g1t's hosted models, `model` is `small` (Fast), `large` (Standard) or
191+`frontier` (Most capable), or null for Auto:
192+
193+```sh
194+curl -X PUT https://api.g1t.sh/workspaces/acme/model-routes \
195+ -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \
196+ -d '{"routes": [
197+ {"task": "default", "connection_id": null, "model": null},
198+ {"task": "plan", "connection_id": null, "model": "frontier"}
199+ ]}'
200+```
+34−8
8181 | What | Unit | Costs g1t | You pay |
8282 | --- | --- | --- | --- |
8383 | Agent models | A run | What the provider charged | The provider's price, from [AI credit](#ai-credit) |
84−| g1t agent rate | Million tokens a run uses (input, output and cached) | — | $0.25, from Oct 22, 2026 |
84+| g1t agent rate | Million tokens a run uses (input, output and cached), [weighted by kind](#the-agent-rate) | — | $0.25, from Oct 22, 2026 |
85+| g1t agent rate, your own model key | The same, on runs that use [your own provider](/guides/models/) | — | $0.25, from Oct 22, 2026 |
8586 | AI Gateway | A request | What the provider charged | The provider's price: free of markup during beta |
8687 | Sandbox time (agents, workflows, the merge queue) | Second | About $0.001 a minute | About $0.0012 a minute |
8788 | [Larger machines](#workflow-jobs-on-larger-machines) for workflow jobs (`g1t-2core`, `g1t-4core`) | Second | About 2.8 and 5.1 times a sandbox second | Cost + 20% |
232233 are dated changes on the pricing page.
233234
234235 Work a workspace routes to [its own model providers](/guides/models/) is
235−paid for at those providers instead. Such a run is charged here only for
236−its [sandbox time](#sandbox-time), like any other sandbox.
236+paid for at those providers instead. Such a run is charged here for its
237+[sandbox time](#sandbox-time), like any other sandbox, and the agent rate
238+on the tokens it used, on a line of its own (*g1t agent rate, your own
239+model key: 980,000 tokens for work on acme/api#12*).
240+
241+### The agent rate
242+
243+The agent rate pays for what g1t adds around the model: context, memory,
244+routing and orchestration. It is charged per million tokens a run used,
245+on g1t's models and on your own model key alike:
246+
247+1. g1t's model proxy counts each answer's tokens as it passes: input,
248+ output, and prompt-cache reads and writes. The sandbox reports what its
249+ agent counted too, and the rate is charged on the more of the two.
250+2. Each kind of token counts at its weight. Today every token counts
251+ once: input ×1, output ×1, cache reads ×1, cache writes ×1. The weights
252+ are on [g1t.sh/pricing](https://g1t.sh/pricing) under the rate, and a
253+ change to them is a dated price change like any other.
254+3. The run is charged when it reports, and again for tokens counted after
255+ that, never twice for the same token.
237256
257+On the **Usage** page the agent rate's lines count weighted tokens and name
258+the weights: **Agent rate** for runs on g1t's models, and **Agent rate, your
259+own model key** for runs on your own provider.
260+
238261 The charge goes to the workspace that owns the repository, whoever
239262 assigned the issue. That is why putting g1t to work on a
240263 repository needs the Write [role](/guides/access-and-roles/) or higher on it.
808831 Hover or focus a column for each product's part; **Show as a table** has
809832 every number.
810833 - **The breakdown**: each product family with its meters (the agent's
811− model tokens, agent rate and sandbox time; sandbox time; builds; git
812− operations and private storage with what is free; and so on), each with a
813− trend line, how much was used and its charge at price. Open a meter for
814− its projects. The agent also shows its runs, reviews, plans and checks.
834+ model tokens, agent rate, agent rate on your own model key and sandbox
835+ time; sandbox time; builds; git operations and private storage with what
836+ is free; and so on), each with a trend line, how much was used and its
837+ charge at price. Open a meter for its projects. The agent also shows its
838+ runs, reviews, plans and checks, and its tokens by model.
815839
816840 Storage, git operations, scans and search embeddings are metered through
817841 the month and charged when it closes; until then they are marked pending.
831855
832856 | Line | What it holds |
833857 | --- | --- |
834− | Agent runs | Runs on g1t's models: the model's cost plus the margin. |
858+ | Agent runs | Runs on g1t's models: the model at the provider's price (cost plus 20% before Oct 8, 2026). |
859+ | Agent rate | The agent rate on runs on g1t's models. |
860+ | Agent rate, your own model key | The agent rate on runs on your own provider. |
835861 | Runs on your own model provider | Older months only: the flat fee runs on your own provider used to carry. |
836862 | Sandbox time | Each sandbox's time. |
837863 | Self-hosted runner time | Each job on your own runners, at $0. |
+59−32
440440
441441 ## Which model runs
442442
443−You do not pick one. You assign the work to `g1t`, the way you would
444−assign an issue to a colleague, and g1t routes it. On g1t's hosted models,
445−each piece of work goes to the least costly of two tiers that can do it:
443+You do not have to pick one. You assign the work to `g1t`, the way you
444+would assign an issue to a colleague, and **Auto** routes each job to the
445+least costly model that can do it, from three tiers:
446446
447−| Tier | Model today |
448−| --- | --- |
449−| Small | Claude Haiku 4.5 |
450−| Large | Claude Sonnet 5.5 |
447+| Tier | Model today | For |
448+| --- | --- | --- |
449+| Fast | Claude Haiku 4.5 | Small, well-bounded work |
450+| Standard | Claude Sonnet 5.5 | Most changes and reviews |
451+| Most capable | Claude Opus 5.5 | Hard work, and work that failed on the standard model |
451452
452−The work decides the tier:
453+The job starts on its tier:
453454
454−| Work | Tier |
455+| Work | Starts on |
455456 | --- | --- |
456−| Making a change for an issue, revising it, and answering a mention | Large |
457−| Reviewing a pull request that changes at most 10 files and 200 lines, touches no sensitive path, and is not for an issue labelled `security` | Small |
458−| Reviewing any other pull request, or one whose changed files g1t does not know yet | Large |
459−| Catching up with `main` and resolving conflicts | Small |
460−| Planning an outcome | Small |
461−| Any of these again, after the last attempt at the same work failed or stopped at a guardrail cap | Large |
457+| Making a change for an issue, revising it, and taking over handed-on work | Standard |
458+| Answering a question asked of `@g1t` | Fast |
459+| Reviewing a pull request that changes at most 10 files and 200 lines and touches no sensitive path | Fast |
460+| Reviewing a pull request that changes more than 60 files or 3,000 lines | Most capable |
461+| Reviewing any other pull request, or one whose changed files g1t does not know yet | Standard |
462+| Catching up with the base branch and resolving conflicts | Fast |
463+| Planning an outcome | Standard |
464+
465+Then, in this order:
462466
467+1. **Labels on the issue.** `architecture` sends the work to the most
468+ capable model. `security` keeps it off the fast one. `documentation`,
469+ `docs` and `typo` let a change or an answer start on the fast one.
470+2. **Failures.** When the last attempt at the same work failed or stopped
471+ at a guardrail cap, the next goes one tier up; after two in a row, to
472+ the most capable. A revision counts each round before it. When the
473+ last attempt finished but left a change g1t had
474+ [low confidence](#how-sure-the-agent-is) in, the next goes one tier up.
475+3. **What worked here.** g1t looks at the repository's last 20 runs of
476+ the same kind. When the tier below finished at least 9 in 10 of at
477+ least 5, the work goes down a tier; when this tier failed half of at
478+ least 5, it goes up. Work that touches a sensitive path or carries
479+ one of the labels above is never moved down.
480+
463481 Sensitive paths are the ones that run, configure or guard things: CI
464482 workflows, `.g1t/` and `.github/`, `CODEOWNERS`, secrets such as `.env`
465483 and `.pem` files, and infrastructure such as Dockerfiles, Terraform and
466484 `wrangler.*` files. They are the same paths that lower a change's
467485 [confidence](#how-sure-the-agent-is).
468486
469−The agent's own small background steps run on the small tier.
487+The agent's own small background steps run on the fast tier.
470488
471−Every session opens with a note naming the model that ran, and an agent's
472−review says which model wrote it, so what you got is always on the record.
473−When a better model for a tier appears, g1t changes the route and nothing
474−you have set up needs to change.
489+Every run says which model it used and why, in one line: as the first
490+step on its run, and at the top of its pull request's session. For
491+example, *Used a fast model (Claude Haiku 4.5): small change, 3 files and
492+80 lines.* An agent's review also says which model wrote it. When a better
493+model for a tier appears, g1t changes the route and nothing you have set
494+up needs to change.
475495
496+To choose instead of Auto, an owner picks **Fast**, **Standard** or **Most
497+capable** for a kind of work under
498+[which model does which work](/guides/models/#choose-which-model-does-which-work).
499+Every run of that kind then uses it, and says the workspace chose it.
500+
476501 A workspace that routes its work to [its own provider](/guides/models/)
477−is not routed by tier: its work runs on the model its route names.
502+runs the model its route names. On an Anthropic key with no model named,
503+Auto chooses the tier's Claude model, as on g1t's models.
478504
479505 A pull request g1t opens has `g1t` as its author and as its `agent` in the
480506 API, and its commits are authored `g1t <g1t@users.noreply.g1t.sh>`.
525551
526552 | Setting | Where | What it does |
527553 | --- | --- | --- |
528−| `AGENT_ROUTING` | Runner | JSON. `tiers`: the model behind `small` and `large`, each `{ "modelName", "model" }`. `tasks`: the tier of `implement`, `review`, `update` and `plan`, or `change` to decide by the change. `smallChange`: the most `files` and `lines` a `change` review runs on the small tier with. `largeLabels`: issue labels that keep a review on the large tier. Anything left out takes the defaults above. |
554+| `AGENT_ROUTING` | Runner | JSON. `tiers`: the model behind `small`, `large` and `frontier`, each `{ "modelName", "model", "price" }` (`price`, dollars per million `input`, `output`, `cacheRead` and `cacheWrite` tokens, is for estimates only). `tasks`: the tier `implement`, `revise`, `answer`, `review`, `update` and `plan` start on, or `change` to decide by the change. `smallChange` and `largeChange`: the most `files` and `lines` of a small change, and the least of a large one. `smallLabels`, `largeLabels` and `frontierLabels`: issue labels that move work. `frontierAfter`: failures in a row before the most capable tier. `learning`: `window`, `minRuns`, `stepDownAt` and `stepUpAt`. Anything left out takes the defaults above. |
529555 | `MODELS_URL` | Runner | Where sandboxes send model requests: the model proxy. |
530556 | `AI_GATEWAY_ID` | Model proxy | The gateway hosted requests go through. Empty sends them to the provider directly. |
531557 | `AI_GATEWAY_TOKEN` | Model proxy | Secret. Authenticates to the gateway. |
533559
534560 ## What it costs
535561
536−A workspace pays for g1t's runs on its repositories, after
537−they run: each run is charged its sandbox by the second, at cost plus 20%,
538−and, on g1t's hosted models, what AI Gateway priced its model requests at,
539−plus 20%. A workspace's [own provider](/guides/models/) bills it for the
540−model directly. See
541−[Usage and billing](/guides/usage-and-billing/) for how prices are set and
542−the limits on usage not yet paid for.
543−The workspace's **Usage** page shows what its agents have cost, by day,
544−kind of work, repository, model and pull request. See
545−[usage and billing](/guides/usage-and-billing/).
562+A workspace pays for g1t's runs on its repositories, after they run: each
563+run is charged its sandbox by the second, at cost plus 20%, and the
564+[agent rate](/guides/usage-and-billing/#the-agent-rate) on the tokens it
565+used. On g1t's hosted models, the model is charged at what AI Gateway
566+priced its requests at, the provider's price with no markup. A
567+workspace's [own provider](/guides/models/) bills it for the model
568+directly; the agent rate is still charged, as **Agent rate, your own model
569+key**. See [Usage and billing](/guides/usage-and-billing/) for how prices
570+are set and the limits on usage not yet paid for. The workspace's
571+**Usage** page shows what its agents have cost, by day, kind of work,
572+repository and pull request, and their tokens by model.
546573
547574 ## What a sandbox has
548575
+11−4
229229 revisions (`max_revisions`, 2 by default) only a required check still
230230 failing holds it for a person. There is no model or
231231 agent count to choose: to put more agents to work, assign more issues.
232− On g1t's hosted models, g1t routes each piece of work to a small or a
233− large tier: planning, catching up and reviews of small changes that
234− touch no sensitive path run small; making changes, other reviews, and
235− any retry after a failed attempt run large.
232+ On g1t's hosted models, Auto routes each job to the cheapest of three
233+ tiers that can do it, fast, standard and most capable: catching up,
234+ answering and reviews of small changes that touch no sensitive path
235+ start fast; making changes, revising, planning and other reviews start
236+ standard; reviews of very large changes and issues labelled
237+ `architecture` start most capable. A failed attempt moves the next one
238+ up a tier (two in a row: most capable), and a repository's own recent
239+ runs move work down or up. Each run states its model and why in one
240+ line, on the run and in the session. An owner can pin a tier per kind of
241+ work instead (`PUT /workspaces/{workspace}/model-routes`, `model`
242+ `small`, `large` or `frontier` with `connection_id` null).
236243 - **Who a g1t pull request is for:** g1t is the `author` (`username`
237244 `g1t`, `kind` `agent`) of every pull request it makes and every issue it
238245 files at work; `requested_by` is the person who asked (null when nobody
+56−8
276276 | Agent run through the model proxy (`services/models`) on g1t's hosted models | g1t | `runs` row; session `ms_…` in `cf-aig-metadata` | On finish, the sandbox's figure (Claude Code's `total_cost_usd`, at its own price table, cache tokens included) | Yes, every 15 minutes |
277277 | Agent run straight to the gateway (no `MODELS_URL`) | g1t | `runs` row; session `rs_…` in `cf-aig-metadata` (`services/runner` `gatewaySession`) | As above | Yes |
278278 | Agent run with no gateway (`AI_GATEWAY_ID` empty, self-hosting) | g1t's key | `runs` row, no session | The sandbox's figure | No: nothing to settle against |
279−| Agent run on a workspace's own provider | The workspace | `runs` row, `billed_to = 'workspace'`, no session | None (no cost to g1t) | No; never on g1t's gateway |
279+| Agent run on a workspace's own provider | The workspace | `runs` row, `billed_to = 'workspace'`, session `ms_…` (the proxy counts its tokens by it) | No model cost (none to g1t); the agent rate on `agent_tokens_own`, line `<run>/agent-own` | No; never on g1t's gateway. Closed by the cron's `settle_own_runs`, which charges tokens counted late |
280280 | A sandbox that died before reporting | g1t | as its route | Charged from the gateway when settled | Yes |
281281 | Embeddings (indexing) | g1t | none (Workers AI) | Month-end `context` meter | No: Cloudflare's bill, `embeddings` bucket |
282282 | Embeddings (queries, search and agent context) | g1t | none | None: not charged, by design | No: in Cloudflare's `embeddings` line, shared out |
494494 `card_fee_percent` (29,000 micros per dollar) and `card_fee_fixed`
495495 (300,000). Changing any is a price-book change, never a deploy. `finish_run`
496496 and `settle` charge models at `agent_models`' markup; `charge_agent_rate`
497−charges the tokens `token_usage` counted for the run's session since it was
498−last charged (`runs.agent_tokens`, claimed with a compare-and-set), on a
499−line `<run>/agent` (later `<run>/agent/<tokens>`), with `quantity` the
500−tokens. Runs on a workspace's own provider have no session here and are not
501−charged the rate. **Card fee switch:** sudo → Costs → Guardrails → *Card fee
497+charges the weighted tokens of the run's session since it was last charged
498+(`runs.agent_tokens`, the weighted tokens charged so far, claimed with a
499+compare-and-set), on a line `<run>/agent` (later `<run>/agent/<tokens>`),
500+with `quantity` the weighted tokens. What it counts is the more of what
501+`token_usage` holds for the session and what the sandbox reported with its
502+cost (`finish_run`'s `tokens`, from Claude Code's closing `usage`), each
503+weighted by kind. **Card fee switch:** sudo → Costs → Guardrails → *Card fee
502504 on AI credit bought by card* (`cost_settings.card_fee`, `on`/`off`).
503505
506+**On a workspace's own model key** (migration `0041_agent_rate_own_key.sql`):
507+the run keeps its model session (`runs.session_id`, `ms_…`) so the proxy's
508+counts reach it, and the agent rate is charged at `agent_tokens_own` ($0 until
509+2026-10-22, then $0.25 a million, a rise from nothing with its notice), on
510+`<run>/agent-own` (later `<run>/agent-own/<tokens>`), `billed_to = 'g1t'`
511+(g1t's own charge: it counts toward limits and spend), named *Agent rate,
512+your own model key* on Usage and the statement. The model is never charged.
513+`settle_runs` skips these runs (nothing on g1t's gateway); the cron's
514+`settle_own_runs` closes them 5 minutes after they finish (3 hours after
515+they start, for a sandbox that never reported) and charges tokens counted
516+late. Runs from before have no session and are never charged the rate.
517+
518+### The agent rate's token weights
519+
520+How much each kind of token counts toward the agent rate, on g1t's models
521+and own keys alike, is four price-book meters (migration
522+`0042_agent_rate_weights.sql`): `agent_token_weight_input`, `_output`,
523+`_cache_read` and `_cache_write`, each a weight in millionths in
524+`cost_micros` (1,000,000 counts a token once). All start at 1, which is what
525+the rate always counted. A cached agent run reads most of its context from
526+cache (about 90% of its tokens on a typical Sonnet implement run), so the
527+cache-read weight is the lever: at 1 the rate adds about 44% to such a run's
528+model cost; at 0.1, far less.
529+
530+To count cache reads at a tenth:
531+
532+1. Add the version, effective at once (a lower weight is a fall):
533+
534+ ```sql
535+ INSERT INTO price_versions (id, meter, version, cost_micros, markup_percent, effective_at, reason, created_by, created_at)
536+ VALUES ('pv_agent_token_weight_cache_read_2', 'agent_token_weight_cache_read', 2, 100000, 0,
537+ '2026-10-08T00:00:00Z', 'Cache reads count a tenth toward the agent rate', 'staff', '2026-10-08T00:00:00Z');
538+ ```
539+
540+ in a migration, or with `npx wrangler d1 execute g1t-billing --remote`
541+ until sudo has a form.
542+2. The daily run applies it once `effective_at` has come and writes the
543+ public `price_changes` record (**Run the analysis now** applies it at
544+ once). Raising a weight later is a rise: give it an `effective_at` 14
545+ days out, and owners on the plan are emailed.
546+3. Check `/pricing`: the agent rate's row lists the weights, and Usage's
547+ agent-rate lines name them.
548+
549+`weighted` (`ai.rs`) rounds down to a whole token; a weight is never below 0.
550+
504551 ### Budgets
505552
506553 The owners' spend limit is the budget. `limits.alert_levels` (comma
532579 `token_usage` reads a window (42 days by default, 366 at most) for the
533580 workspace or one person: totals, every day's tokens and the active days,
534581 with `costMicros` the window's run charges from the ledger, measured as
535−`usage` measures them. These counts are for views only: runs are still
536−priced from AI Gateway's logs, never from `token_usage`.
582+`usage` measures them. Usage's report also lists tokens by model
583+(`UsageReport.models`). The agent rate is charged on these counts (above);
584+a run's model is still priced from AI Gateway's logs, never from them.
537585
538586 ## Tables (migration `0022_costs_and_margin.sql`)
539587
+66−10
10431043 - **All hosted model traffic goes through Cloudflare AI Gateway.** That gives
10441044 one place for spend tracking, budgets, rate limits, fallback and logs,
10451045 whichever provider or endpoint is behind it.
1046−- **Nobody picks a model.** A person assigns work to `g1t`, as they
1047− would assign an issue to Copilot, and g1t routes it. Today the kind of
1048− work decides (implementing, reviewing, catching up), from one setting on
1049− the runner, and each request is tagged at the gateway with that kind, the
1050− repository and the pull request. The session records which model ran.
1051− The gateway's own dynamic routes cannot make the choice yet: they work
1052− only on its OpenAI-compatible endpoint, and the harness speaks
1053− Anthropic's.
1046+- **Nobody has to pick a model.** A person assigns work to `g1t`, as
1047+ they would assign an issue to a colleague, and **Auto** routes each job
1048+ to the cheapest model that can do it (see
1049+ [Routing for cost](#routing-for-cost) below). A workspace can pin a tier
1050+ per kind of work instead. Each request is tagged at the gateway with the
1051+ kind of work, the tier, the repository and the pull request, and each
1052+ run says which model ran and why. The gateway's own dynamic routes
1053+ cannot make the choice yet: they work only on its OpenAI-compatible
1054+ endpoint, and the harness speaks Anthropic's.
10541055 - **Subscriptions stay local.** A Claude subscription cannot be used by a
10551056 hosted sandbox; it needs an API key. People on subscriptions use their own
10561057 Claude Code session, which is a full participant.
1057−- **The workspace pays.** A workspace buys credit by card and each agent
1058− run deducts what the model cost plus a margin. The billing service asks
1058+- **The workspace pays.** A workspace buys AI credit by card and each
1059+ agent run deducts the model at the provider's price plus the agent rate
1060+ per million (weighted) tokens; on its own model key, only the agent rate,
1061+ counted from the model proxy and the sandbox's own report. The billing service asks
10591062 nothing of the others: the runner asks it before starting a sandbox and
10601063 is refused when there is no credit, and the sandbox reports what its run
10611064 cost with a token only it holds. Where no card processor is configured
11001103 an assigned issue to a commit on `main` with nobody in between. Required
11011104 human approval per path, and risk tiers, are still to come.
11021105
1106+### Routing for cost
1107+
1108+*Built (runner `route` in `services/runner/src/model-env.ts`).* The goal
1109+is cost per merged change, not cost per request: a cheap attempt that
1110+fails and is retried on the same model costs more than one that finishes.
1111+
1112+- **Tiers and catalogue.** `small` (Claude Haiku 4.5, $1/$5 per million
1113+ input/output), `large` (Claude Sonnet 5.5, $2/$10) and `frontier`
1114+ (Claude Opus 5.5, $4/$20). Models, names and list prices are
1115+ configuration (`AGENT_ROUTING`), never code; prices there are for
1116+ estimates only, runs are charged what AI Gateway priced them at.
1117+- **Starting tier by job.** Catch-up, answering a question, and reviews of
1118+ at most 10 files and 200 lines touching no sensitive path: small.
1119+ Changes, revisions, plans and other reviews: large. Reviews over 60 files
1120+ or 3,000 lines: frontier. Labels: `architecture` frontier, `security` off
1121+ small, `docs`/`documentation`/`typo` let changes and answers start small.
1122+- **Escalation.** A failed (or guardrail-stopped) attempt at the same work
1123+ goes one tier up; two in a row, frontier; a revision counts its rounds;
1124+ a change left at low confidence sends the next attempt up.
1125+- **Learning, per repository.** From the last 20 runs of the same kind:
1126+ one tier down when the cheaper tier finished at least 90% of at least 5
1127+ (never for sensitive or labelled work, never on a retry); one tier up
1128+ when this tier failed at least half of at least 5. No new tables: it
1129+ reads work's `agent_runs` (model, status, confidence).
1130+- **Explained.** Every run's first step and session note is one line:
1131+ *Used a fast model (Claude Haiku 4.5): small change, 3 files and 80
1132+ lines.*
1133+- **Chosen instead.** `model_routes` rows to g1t's models name `small`,
1134+ `large` or `frontier`, or nothing for Auto (Integrations → Models).
1135+ A workspace's own Anthropic key with no model named is routed by Auto
1136+ too.
1137+- **Measured.** `scripts/ops/routing-savings.mjs` replays tasks through
1138+ the router offline, priced from the catalogue, against routing before
1139+ Auto and against the frontier model for everything, net of failed
1140+ attempts, with cost per merged change; `--live` reads billing's runs and
1141+ counted tokens. On the bundled sample (13 tasks, assumed failures): Auto
1142+ costs 7% less than the frontier model for everything and about 7% more
1143+ per run than routing before it, but half as much per merged change,
1144+ because it finishes the hard tasks the old routing gave up on. At
1145+ current prices Opus 5.5 and Sonnet 5.5 cost the same per cache read, and
1146+ cache reads are most of an agent run's tokens, so moving off the
1147+ frontier model saves less than its list price suggests; the fast tier
1148+ and fewer failed attempts are where the money is. Run `--live` monthly
1149+ and after any routing change.
1150+- **A cheaper route for the simplest jobs (designed, off).** A fourth tier
1151+ on Workers AI through AI Gateway (an open model, billed on Cloudflare's
1152+ invoice) for classification-sized jobs: commit messages, triage,
1153+ summaries. Behind the same router as a tier with its own catalogue entry
1154+ and `tasks` rules, off by default. It needs the proxy to translate the
1155+ harness's Anthropic requests to the gateway's OpenAI-compatible
1156+ endpoint (it already does for workspaces' own OpenAI-shaped providers)
1157+ and a quality bar from the savings harness before any job moves to it.
1158+
11031159 ### Choosing the right agent automatically
11041160
11051161 Because several agents can work on the same issue, every issue with more
+3−2
195195 console.log(`${name.padEnd(9)} ${dollars(total.cost).padStart(10)} merged ${total.merged} per merged change ${dollars(total.perMerged)}`);
196196 }
197197 console.log("");
198− console.log(`Auto against routing before it: ${(savings.vsBefore * 100).toFixed(1)}% ${savings.vsBefore >= 0 ? "less" : "more"}`);
199− console.log(`Auto against the most capable model for everything: ${(savings.vsFrontier * 100).toFixed(1)}% less`);
198+ const say = (share) => `${Math.abs(share * 100).toFixed(1)}% ${share >= 0 ? "less" : "more"}`;
199+ console.log(`Auto against routing before it: ${say(savings.vsBefore)}`);
200+ console.log(`Auto against the most capable model for everything: ${say(savings.vsFrontier)}`);
200201 }
201202
202203 async function main() {