Docs: Auto, choosing a model, and the agent rate on your own key
The models guide explains Auto, choosing a tier per kind of work, and what own-key runs are charged; Working with g1t has the full routing rules and the AGENT_ROUTING settings; Usage and billing gains the agent rate (how it is counted and weighted), its own-key line, and the statement's agent-rate rows. BILLING_OPERATIONS covers own-key metering, the cron that closes those runs, and how to set the token weights (cache reads at a tenth). PLAN records routing for cost, its measured savings, and the Workers AI route designed and left off. llms.txt follows.
7 files+306−850/7 viewed
| 1 | 1 | --- | |
| 2 | 2 | title: Model providers | |
| 3 | − | description: Connect Anthropic, OpenAI, Gemini or any compatible endpoint, choose which model does which work, and pay for it where you choose. | |
| 3 | + | description: Let Auto choose the model for each job, or choose it yourself; connect Anthropic, OpenAI, Gemini or any compatible endpoint, and pay for models where you choose. | |
| 4 | 4 | --- | |
| 5 | 5 | ||
| 6 | 6 | Each workspace decides where its agents' model spend goes: | |
| 7 | 7 | ||
| 8 | − | - **g1t's hosted models.** g1t chooses the model for each kind of work, pays | |
| 9 | − | the provider, and charges your workspace what it cost plus 20%. The | |
| 10 | − | plan's included usage and [the trial](/guides/usage-and-billing/#the-trial) | |
| 8 | + | - **g1t's hosted models.** **Auto** chooses the model for each job, g1t pays | |
| 9 | + | the provider, and your workspace is charged the provider's price, with no | |
| 10 | + | markup, plus the [agent rate](/guides/usage-and-billing/#the-agent-rate). | |
| 11 | + | The plan's included usage and [the trial](/guides/usage-and-billing/#the-trial) | |
| 11 | 12 | pay for it first. While payments are in test mode, they are open only | |
| 12 | 13 | to a few invited workspaces, g1t's own among them; a card check or a | |
| 13 | 14 | trial does not open them. Every other workspace connects its own | |
| 15 | 16 | that says so. Once payments go live, they are open to all. | |
| 16 | 17 | - **Your own providers.** Connect as many as you use, then choose, for each | |
| 17 | 18 | kind of work, which provider and model it runs on. Each provider bills | |
| 18 | − | you directly. Open to every workspace now. | |
| 19 | + | you for the model directly. Open to every workspace now. | |
| 19 | 20 | ||
| 20 | − | g1t's own routing is fixed; yours is not. | |
| 21 | + | ## Auto | |
| 21 | 22 | ||
| 23 | + | On g1t's models you do not have to pick a model. **Auto**, the default, | |
| 24 | + | sends each job to the least costly model that can do it, from three tiers: | |
| 25 | + | **fast** (Claude Haiku 4.5 today), **standard** (Claude Sonnet 5.5) and | |
| 26 | + | **most capable** (Claude Opus 5.5). It decides by the kind of job, the | |
| 27 | + | size of the change it reads, the issue's labels, whether the last attempt | |
| 28 | + | at the same work failed, and what has worked in the repository before: | |
| 29 | + | ||
| 30 | + | - Catching up, answering a question and reviewing a small change that | |
| 31 | + | touches no sensitive path start on the fast model. | |
| 32 | + | - Making and revising changes, planning, and most reviews start on the | |
| 33 | + | standard model. | |
| 34 | + | - A review of a very large change, work on an issue labelled | |
| 35 | + | `architecture`, and work that failed twice in a row go to the most | |
| 36 | + | capable model. One failure moves the next attempt up one tier. | |
| 37 | + | - When the cheaper model finished nearly all of a repository's recent runs | |
| 38 | + | of the same kind, Auto moves that work down a tier there; when a model | |
| 39 | + | keeps failing, up. | |
| 40 | + | ||
| 41 | + | Every run says which model it used and why, in one line on its run and in | |
| 42 | + | its pull request's session, such as *Used a fast model (Claude Haiku 4.5): | |
| 43 | + | small change, 3 files and 80 lines.* The full rules are in | |
| 44 | + | [which model runs](/guides/working-with-g1t/#which-model-runs). | |
| 45 | + | ||
| 22 | 46 | ## Providers | |
| 23 | 47 | ||
| 24 | 48 | Labs and platforms need only a key; g1t knows where they are. | |
| 75 | 99 | | Catching up | Bringing a change up to date with `main`. | | |
| 76 | 100 | ||
| 77 | 101 | Each can go to g1t's models, or to any of your providers on any of its | |
| 78 | − | models. An Anthropic provider also offers **g1t's choice of Claude**, | |
| 79 | − | which runs g1t's large-tier model, Claude Sonnet 5.5 today, on your key. | |
| 80 | − | Routing between tiers by the size of the work is only for g1t's hosted | |
| 81 | − | models; see [which model runs](/guides/working-with-g1t/#which-model-runs). For example: make | |
| 82 | − | changes on Claude through your Anthropic key, review on GPT through your | |
| 83 | − | OpenAI key, and catch up on a small model through OpenRouter. | |
| 102 | + | models. | |
| 103 | + | ||
| 104 | + | - **On g1t's models**, choose **Auto** (the default), or a tier for every | |
| 105 | + | run of that kind: **Fast**, **Standard** or **Most capable**. A run on a | |
| 106 | + | chosen tier says the workspace chose it. | |
| 107 | + | - **On an Anthropic provider**, leave the model empty for g1t's choice of | |
| 108 | + | Claude: Auto picks the tier's model for each job, as on g1t's models, on | |
| 109 | + | your key. Name a model to run that one every time. | |
| 110 | + | - **On any other provider**, name the model. | |
| 111 | + | ||
| 112 | + | For example: make changes on Claude through your Anthropic key, review on | |
| 113 | + | GPT through your OpenAI key, catch up on g1t's models on Fast, and plan on | |
| 114 | + | Most capable. | |
| 84 | 115 | ||
| 85 | 116 | **Save routing**, and the next runs use it. Without any routing, work goes | |
| 86 | 117 | to g1t's models where they are open to the workspace, and otherwise to the | |
| 90 | 121 | ||
| 91 | 122 | ## What it costs | |
| 92 | 123 | ||
| 93 | − | Your providers bill you for the models. g1t charges only each run's | |
| 94 | − | [sandbox time](/guides/usage-and-billing/#sandbox-time), at what it costs | |
| 95 | − | g1t plus 20%, by the second. A change, a review, a revision, a catch-up | |
| 96 | − | and a plan each run in a sandbox, and each sandbox is a line on the | |
| 97 | − | statement. See [Usage and billing](/guides/usage-and-billing/). | |
| 124 | + | On g1t's models, each run is charged the model at the provider's price | |
| 125 | + | (what AI Gateway priced its requests at, with no markup), the agent rate | |
| 126 | + | on its tokens, and its sandbox time. Auto keeps the first of those down: | |
| 127 | + | the fast model costs about half what the standard one does, and the most | |
| 128 | + | capable up to twice as much, so most of what it saves comes from sending small | |
| 129 | + | jobs to the fast model and from finishing hard ones instead of retrying | |
| 130 | + | them on the same model. | |
| 98 | 131 | ||
| 99 | − | That sandbox time counts toward the workspace's usage limit like any | |
| 100 | − | other. | |
| 132 | + | On your own providers, they bill you for the models. g1t charges each | |
| 133 | + | run's [sandbox time](/guides/usage-and-billing/#sandbox-time), at what it | |
| 134 | + | costs g1t plus 20%, by the second, and the | |
| 135 | + | [agent rate](/guides/usage-and-billing/#the-agent-rate) on the tokens the | |
| 136 | + | run used, from Oct 22, 2026: $0.25 per million, as on g1t's models. Tokens | |
| 137 | + | are counted by g1t's model proxy as answers pass, and by the agent in the | |
| 138 | + | sandbox; the more of the two is charged. On **Usage** it is the line | |
| 139 | + | **Agent rate, your own model key**, with its tokens weighted as the | |
| 140 | + | pricing page says. | |
| 141 | + | ||
| 142 | + | Both count toward the workspace's usage limit like any other charge. | |
| 143 | + | **Usage** also shows the agent's tokens by model. | |
| 101 | 144 | ||
| 102 | 145 | ## Your keys never reach a sandbox | |
| 103 | 146 | ||
| 114 | 157 | ||
| 115 | 158 | As each answer passes, the proxy reads how many tokens it used (input, | |
| 116 | 159 | output, and cache reads and writes) and counts them for the run, under the | |
| 117 | − | person it was for. Those counts are for usage views; they never change what | |
| 118 | − | a run is charged. | |
| 160 | + | person it was for. They show on **Usage** by model, and the | |
| 161 | + | [agent rate](/guides/usage-and-billing/#the-agent-rate) is charged on them. | |
| 162 | + | On g1t's models, the model itself is charged at what AI Gateway priced it | |
| 163 | + | at, never from these counts. | |
| 119 | 164 | ||
| 120 | 165 | The token stops working within seconds of the run finishing, however it | |
| 121 | 166 | ends, and within seconds if you disconnect the provider. A run whose end | |
| 142 | 187 | `task` is `default`, `implement`, `review`, `plan` or `update`. | |
| 143 | 188 | `connection_id` is null for g1t's hosted models. `model` is null for the | |
| 144 | 189 | provider's default, or for an Anthropic provider, g1t's choice of Claude. | |
| 190 | + | On g1t's hosted models, `model` is `small` (Fast), `large` (Standard) or | |
| 191 | + | `frontier` (Most capable), or null for Auto: | |
| 192 | + | ||
| 193 | + | ```sh | |
| 194 | + | curl -X PUT https://api.g1t.sh/workspaces/acme/model-routes \ | |
| 195 | + | -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \ | |
| 196 | + | -d '{"routes": [ | |
| 197 | + | {"task": "default", "connection_id": null, "model": null}, | |
| 198 | + | {"task": "plan", "connection_id": null, "model": "frontier"} | |
| 199 | + | ]}' | |
| 200 | + | ``` |
| 81 | 81 | | What | Unit | Costs g1t | You pay | | |
| 82 | 82 | | --- | --- | --- | --- | | |
| 83 | 83 | | Agent models | A run | What the provider charged | The provider's price, from [AI credit](#ai-credit) | | |
| 84 | − | | g1t agent rate | Million tokens a run uses (input, output and cached) | — | $0.25, from Oct 22, 2026 | | |
| 84 | + | | g1t agent rate | Million tokens a run uses (input, output and cached), [weighted by kind](#the-agent-rate) | — | $0.25, from Oct 22, 2026 | | |
| 85 | + | | g1t agent rate, your own model key | The same, on runs that use [your own provider](/guides/models/) | — | $0.25, from Oct 22, 2026 | | |
| 85 | 86 | | AI Gateway | A request | What the provider charged | The provider's price: free of markup during beta | | |
| 86 | 87 | | Sandbox time (agents, workflows, the merge queue) | Second | About $0.001 a minute | About $0.0012 a minute | | |
| 87 | 88 | | [Larger machines](#workflow-jobs-on-larger-machines) for workflow jobs (`g1t-2core`, `g1t-4core`) | Second | About 2.8 and 5.1 times a sandbox second | Cost + 20% | | |
| 232 | 233 | are dated changes on the pricing page. | |
| 233 | 234 | ||
| 234 | 235 | Work a workspace routes to [its own model providers](/guides/models/) is | |
| 235 | − | paid for at those providers instead. Such a run is charged here only for | |
| 236 | − | its [sandbox time](#sandbox-time), like any other sandbox. | |
| 236 | + | paid for at those providers instead. Such a run is charged here for its | |
| 237 | + | [sandbox time](#sandbox-time), like any other sandbox, and the agent rate | |
| 238 | + | on the tokens it used, on a line of its own (*g1t agent rate, your own | |
| 239 | + | model key: 980,000 tokens for work on acme/api#12*). | |
| 240 | + | ||
| 241 | + | ### The agent rate | |
| 242 | + | ||
| 243 | + | The agent rate pays for what g1t adds around the model: context, memory, | |
| 244 | + | routing and orchestration. It is charged per million tokens a run used, | |
| 245 | + | on g1t's models and on your own model key alike: | |
| 246 | + | ||
| 247 | + | 1. g1t's model proxy counts each answer's tokens as it passes: input, | |
| 248 | + | output, and prompt-cache reads and writes. The sandbox reports what its | |
| 249 | + | agent counted too, and the rate is charged on the more of the two. | |
| 250 | + | 2. Each kind of token counts at its weight. Today every token counts | |
| 251 | + | once: input ×1, output ×1, cache reads ×1, cache writes ×1. The weights | |
| 252 | + | are on [g1t.sh/pricing](https://g1t.sh/pricing) under the rate, and a | |
| 253 | + | change to them is a dated price change like any other. | |
| 254 | + | 3. The run is charged when it reports, and again for tokens counted after | |
| 255 | + | that, never twice for the same token. | |
| 237 | 256 | ||
| 257 | + | On the **Usage** page the agent rate's lines count weighted tokens and name | |
| 258 | + | the weights: **Agent rate** for runs on g1t's models, and **Agent rate, your | |
| 259 | + | own model key** for runs on your own provider. | |
| 260 | + | ||
| 238 | 261 | The charge goes to the workspace that owns the repository, whoever | |
| 239 | 262 | assigned the issue. That is why putting g1t to work on a | |
| 240 | 263 | repository needs the Write [role](/guides/access-and-roles/) or higher on it. | |
| 808 | 831 | Hover or focus a column for each product's part; **Show as a table** has | |
| 809 | 832 | every number. | |
| 810 | 833 | - **The breakdown**: each product family with its meters (the agent's | |
| 811 | − | model tokens, agent rate and sandbox time; sandbox time; builds; git | |
| 812 | − | operations and private storage with what is free; and so on), each with a | |
| 813 | − | trend line, how much was used and its charge at price. Open a meter for | |
| 814 | − | its projects. The agent also shows its runs, reviews, plans and checks. | |
| 834 | + | model tokens, agent rate, agent rate on your own model key and sandbox | |
| 835 | + | time; sandbox time; builds; git operations and private storage with what | |
| 836 | + | is free; and so on), each with a trend line, how much was used and its | |
| 837 | + | charge at price. Open a meter for its projects. The agent also shows its | |
| 838 | + | runs, reviews, plans and checks, and its tokens by model. | |
| 815 | 839 | ||
| 816 | 840 | Storage, git operations, scans and search embeddings are metered through | |
| 817 | 841 | the month and charged when it closes; until then they are marked pending. | |
| 831 | 855 | ||
| 832 | 856 | | Line | What it holds | | |
| 833 | 857 | | --- | --- | | |
| 834 | − | | Agent runs | Runs on g1t's models: the model's cost plus the margin. | | |
| 858 | + | | Agent runs | Runs on g1t's models: the model at the provider's price (cost plus 20% before Oct 8, 2026). | | |
| 859 | + | | Agent rate | The agent rate on runs on g1t's models. | | |
| 860 | + | | Agent rate, your own model key | The agent rate on runs on your own provider. | | |
| 835 | 861 | | Runs on your own model provider | Older months only: the flat fee runs on your own provider used to carry. | | |
| 836 | 862 | | Sandbox time | Each sandbox's time. | | |
| 837 | 863 | | Self-hosted runner time | Each job on your own runners, at $0. | |
| 440 | 440 | ||
| 441 | 441 | ## Which model runs | |
| 442 | 442 | ||
| 443 | − | You do not pick one. You assign the work to `g1t`, the way you would | |
| 444 | − | assign an issue to a colleague, and g1t routes it. On g1t's hosted models, | |
| 445 | − | each piece of work goes to the least costly of two tiers that can do it: | |
| 443 | + | You do not have to pick one. You assign the work to `g1t`, the way you | |
| 444 | + | would assign an issue to a colleague, and **Auto** routes each job to the | |
| 445 | + | least costly model that can do it, from three tiers: | |
| 446 | 446 | ||
| 447 | − | | Tier | Model today | | |
| 448 | − | | --- | --- | | |
| 449 | − | | Small | Claude Haiku 4.5 | | |
| 450 | − | | Large | Claude Sonnet 5.5 | | |
| 447 | + | | Tier | Model today | For | | |
| 448 | + | | --- | --- | --- | | |
| 449 | + | | Fast | Claude Haiku 4.5 | Small, well-bounded work | | |
| 450 | + | | Standard | Claude Sonnet 5.5 | Most changes and reviews | | |
| 451 | + | | Most capable | Claude Opus 5.5 | Hard work, and work that failed on the standard model | | |
| 451 | 452 | ||
| 452 | − | The work decides the tier: | |
| 453 | + | The job starts on its tier: | |
| 453 | 454 | ||
| 454 | − | | Work | Tier | | |
| 455 | + | | Work | Starts on | | |
| 455 | 456 | | --- | --- | | |
| 456 | − | | Making a change for an issue, revising it, and answering a mention | Large | | |
| 457 | − | | Reviewing a pull request that changes at most 10 files and 200 lines, touches no sensitive path, and is not for an issue labelled `security` | Small | | |
| 458 | − | | Reviewing any other pull request, or one whose changed files g1t does not know yet | Large | | |
| 459 | − | | Catching up with `main` and resolving conflicts | Small | | |
| 460 | − | | Planning an outcome | Small | | |
| 461 | − | | Any of these again, after the last attempt at the same work failed or stopped at a guardrail cap | Large | | |
| 457 | + | | Making a change for an issue, revising it, and taking over handed-on work | Standard | | |
| 458 | + | | Answering a question asked of `@g1t` | Fast | | |
| 459 | + | | Reviewing a pull request that changes at most 10 files and 200 lines and touches no sensitive path | Fast | | |
| 460 | + | | Reviewing a pull request that changes more than 60 files or 3,000 lines | Most capable | | |
| 461 | + | | Reviewing any other pull request, or one whose changed files g1t does not know yet | Standard | | |
| 462 | + | | Catching up with the base branch and resolving conflicts | Fast | | |
| 463 | + | | Planning an outcome | Standard | | |
| 464 | + | ||
| 465 | + | Then, in this order: | |
| 462 | 466 | ||
| 467 | + | 1. **Labels on the issue.** `architecture` sends the work to the most | |
| 468 | + | capable model. `security` keeps it off the fast one. `documentation`, | |
| 469 | + | `docs` and `typo` let a change or an answer start on the fast one. | |
| 470 | + | 2. **Failures.** When the last attempt at the same work failed or stopped | |
| 471 | + | at a guardrail cap, the next goes one tier up; after two in a row, to | |
| 472 | + | the most capable. A revision counts each round before it. When the | |
| 473 | + | last attempt finished but left a change g1t had | |
| 474 | + | [low confidence](#how-sure-the-agent-is) in, the next goes one tier up. | |
| 475 | + | 3. **What worked here.** g1t looks at the repository's last 20 runs of | |
| 476 | + | the same kind. When the tier below finished at least 9 in 10 of at | |
| 477 | + | least 5, the work goes down a tier; when this tier failed half of at | |
| 478 | + | least 5, it goes up. Work that touches a sensitive path or carries | |
| 479 | + | one of the labels above is never moved down. | |
| 480 | + | ||
| 463 | 481 | Sensitive paths are the ones that run, configure or guard things: CI | |
| 464 | 482 | workflows, `.g1t/` and `.github/`, `CODEOWNERS`, secrets such as `.env` | |
| 465 | 483 | and `.pem` files, and infrastructure such as Dockerfiles, Terraform and | |
| 466 | 484 | `wrangler.*` files. They are the same paths that lower a change's | |
| 467 | 485 | [confidence](#how-sure-the-agent-is). | |
| 468 | 486 | ||
| 469 | − | The agent's own small background steps run on the small tier. | |
| 487 | + | The agent's own small background steps run on the fast tier. | |
| 470 | 488 | ||
| 471 | − | Every session opens with a note naming the model that ran, and an agent's | |
| 472 | − | review says which model wrote it, so what you got is always on the record. | |
| 473 | − | When a better model for a tier appears, g1t changes the route and nothing | |
| 474 | − | you have set up needs to change. | |
| 489 | + | Every run says which model it used and why, in one line: as the first | |
| 490 | + | step on its run, and at the top of its pull request's session. For | |
| 491 | + | example, *Used a fast model (Claude Haiku 4.5): small change, 3 files and | |
| 492 | + | 80 lines.* An agent's review also says which model wrote it. When a better | |
| 493 | + | model for a tier appears, g1t changes the route and nothing you have set | |
| 494 | + | up needs to change. | |
| 475 | 495 | ||
| 496 | + | To choose instead of Auto, an owner picks **Fast**, **Standard** or **Most | |
| 497 | + | capable** for a kind of work under | |
| 498 | + | [which model does which work](/guides/models/#choose-which-model-does-which-work). | |
| 499 | + | Every run of that kind then uses it, and says the workspace chose it. | |
| 500 | + | ||
| 476 | 501 | A workspace that routes its work to [its own provider](/guides/models/) | |
| 477 | − | is not routed by tier: its work runs on the model its route names. | |
| 502 | + | runs the model its route names. On an Anthropic key with no model named, | |
| 503 | + | Auto chooses the tier's Claude model, as on g1t's models. | |
| 478 | 504 | ||
| 479 | 505 | A pull request g1t opens has `g1t` as its author and as its `agent` in the | |
| 480 | 506 | API, and its commits are authored `g1t <g1t@users.noreply.g1t.sh>`. | |
| 525 | 551 | ||
| 526 | 552 | | Setting | Where | What it does | | |
| 527 | 553 | | --- | --- | --- | | |
| 528 | − | | `AGENT_ROUTING` | Runner | JSON. `tiers`: the model behind `small` and `large`, each `{ "modelName", "model" }`. `tasks`: the tier of `implement`, `review`, `update` and `plan`, or `change` to decide by the change. `smallChange`: the most `files` and `lines` a `change` review runs on the small tier with. `largeLabels`: issue labels that keep a review on the large tier. Anything left out takes the defaults above. | | |
| 554 | + | | `AGENT_ROUTING` | Runner | JSON. `tiers`: the model behind `small`, `large` and `frontier`, each `{ "modelName", "model", "price" }` (`price`, dollars per million `input`, `output`, `cacheRead` and `cacheWrite` tokens, is for estimates only). `tasks`: the tier `implement`, `revise`, `answer`, `review`, `update` and `plan` start on, or `change` to decide by the change. `smallChange` and `largeChange`: the most `files` and `lines` of a small change, and the least of a large one. `smallLabels`, `largeLabels` and `frontierLabels`: issue labels that move work. `frontierAfter`: failures in a row before the most capable tier. `learning`: `window`, `minRuns`, `stepDownAt` and `stepUpAt`. Anything left out takes the defaults above. | | |
| 529 | 555 | | `MODELS_URL` | Runner | Where sandboxes send model requests: the model proxy. | | |
| 530 | 556 | | `AI_GATEWAY_ID` | Model proxy | The gateway hosted requests go through. Empty sends them to the provider directly. | | |
| 531 | 557 | | `AI_GATEWAY_TOKEN` | Model proxy | Secret. Authenticates to the gateway. | | |
| 533 | 559 | ||
| 534 | 560 | ## What it costs | |
| 535 | 561 | ||
| 536 | − | A workspace pays for g1t's runs on its repositories, after | |
| 537 | − | they run: each run is charged its sandbox by the second, at cost plus 20%, | |
| 538 | − | and, on g1t's hosted models, what AI Gateway priced its model requests at, | |
| 539 | − | plus 20%. A workspace's [own provider](/guides/models/) bills it for the | |
| 540 | − | model directly. See | |
| 541 | − | [Usage and billing](/guides/usage-and-billing/) for how prices are set and | |
| 542 | − | the limits on usage not yet paid for. | |
| 543 | − | The workspace's **Usage** page shows what its agents have cost, by day, | |
| 544 | − | kind of work, repository, model and pull request. See | |
| 545 | − | [usage and billing](/guides/usage-and-billing/). | |
| 562 | + | A workspace pays for g1t's runs on its repositories, after they run: each | |
| 563 | + | run is charged its sandbox by the second, at cost plus 20%, and the | |
| 564 | + | [agent rate](/guides/usage-and-billing/#the-agent-rate) on the tokens it | |
| 565 | + | used. On g1t's hosted models, the model is charged at what AI Gateway | |
| 566 | + | priced its requests at, the provider's price with no markup. A | |
| 567 | + | workspace's [own provider](/guides/models/) bills it for the model | |
| 568 | + | directly; the agent rate is still charged, as **Agent rate, your own model | |
| 569 | + | key**. See [Usage and billing](/guides/usage-and-billing/) for how prices | |
| 570 | + | are set and the limits on usage not yet paid for. The workspace's | |
| 571 | + | **Usage** page shows what its agents have cost, by day, kind of work, | |
| 572 | + | repository and pull request, and their tokens by model. | |
| 546 | 573 | ||
| 547 | 574 | ## What a sandbox has | |
| 548 | 575 |
| 229 | 229 | revisions (`max_revisions`, 2 by default) only a required check still | |
| 230 | 230 | failing holds it for a person. There is no model or | |
| 231 | 231 | agent count to choose: to put more agents to work, assign more issues. | |
| 232 | − | On g1t's hosted models, g1t routes each piece of work to a small or a | |
| 233 | − | large tier: planning, catching up and reviews of small changes that | |
| 234 | − | touch no sensitive path run small; making changes, other reviews, and | |
| 235 | − | any retry after a failed attempt run large. | |
| 232 | + | On g1t's hosted models, Auto routes each job to the cheapest of three | |
| 233 | + | tiers that can do it, fast, standard and most capable: catching up, | |
| 234 | + | answering and reviews of small changes that touch no sensitive path | |
| 235 | + | start fast; making changes, revising, planning and other reviews start | |
| 236 | + | standard; reviews of very large changes and issues labelled | |
| 237 | + | `architecture` start most capable. A failed attempt moves the next one | |
| 238 | + | up a tier (two in a row: most capable), and a repository's own recent | |
| 239 | + | runs move work down or up. Each run states its model and why in one | |
| 240 | + | line, on the run and in the session. An owner can pin a tier per kind of | |
| 241 | + | work instead (`PUT /workspaces/{workspace}/model-routes`, `model` | |
| 242 | + | `small`, `large` or `frontier` with `connection_id` null). | |
| 236 | 243 | - **Who a g1t pull request is for:** g1t is the `author` (`username` | |
| 237 | 244 | `g1t`, `kind` `agent`) of every pull request it makes and every issue it | |
| 238 | 245 | files at work; `requested_by` is the person who asked (null when nobody |
| 276 | 276 | | Agent run through the model proxy (`services/models`) on g1t's hosted models | g1t | `runs` row; session `ms_…` in `cf-aig-metadata` | On finish, the sandbox's figure (Claude Code's `total_cost_usd`, at its own price table, cache tokens included) | Yes, every 15 minutes | | |
| 277 | 277 | | Agent run straight to the gateway (no `MODELS_URL`) | g1t | `runs` row; session `rs_…` in `cf-aig-metadata` (`services/runner` `gatewaySession`) | As above | Yes | | |
| 278 | 278 | | Agent run with no gateway (`AI_GATEWAY_ID` empty, self-hosting) | g1t's key | `runs` row, no session | The sandbox's figure | No: nothing to settle against | | |
| 279 | − | | Agent run on a workspace's own provider | The workspace | `runs` row, `billed_to = 'workspace'`, no session | None (no cost to g1t) | No; never on g1t's gateway | | |
| 279 | + | | Agent run on a workspace's own provider | The workspace | `runs` row, `billed_to = 'workspace'`, session `ms_…` (the proxy counts its tokens by it) | No model cost (none to g1t); the agent rate on `agent_tokens_own`, line `<run>/agent-own` | No; never on g1t's gateway. Closed by the cron's `settle_own_runs`, which charges tokens counted late | | |
| 280 | 280 | | A sandbox that died before reporting | g1t | as its route | Charged from the gateway when settled | Yes | | |
| 281 | 281 | | Embeddings (indexing) | g1t | none (Workers AI) | Month-end `context` meter | No: Cloudflare's bill, `embeddings` bucket | | |
| 282 | 282 | | Embeddings (queries, search and agent context) | g1t | none | None: not charged, by design | No: in Cloudflare's `embeddings` line, shared out | | |
| 494 | 494 | `card_fee_percent` (29,000 micros per dollar) and `card_fee_fixed` | |
| 495 | 495 | (300,000). Changing any is a price-book change, never a deploy. `finish_run` | |
| 496 | 496 | and `settle` charge models at `agent_models`' markup; `charge_agent_rate` | |
| 497 | − | charges the tokens `token_usage` counted for the run's session since it was | |
| 498 | − | last charged (`runs.agent_tokens`, claimed with a compare-and-set), on a | |
| 499 | − | line `<run>/agent` (later `<run>/agent/<tokens>`), with `quantity` the | |
| 500 | − | tokens. Runs on a workspace's own provider have no session here and are not | |
| 501 | − | charged the rate. **Card fee switch:** sudo → Costs → Guardrails → *Card fee | |
| 497 | + | charges the weighted tokens of the run's session since it was last charged | |
| 498 | + | (`runs.agent_tokens`, the weighted tokens charged so far, claimed with a | |
| 499 | + | compare-and-set), on a line `<run>/agent` (later `<run>/agent/<tokens>`), | |
| 500 | + | with `quantity` the weighted tokens. What it counts is the more of what | |
| 501 | + | `token_usage` holds for the session and what the sandbox reported with its | |
| 502 | + | cost (`finish_run`'s `tokens`, from Claude Code's closing `usage`), each | |
| 503 | + | weighted by kind. **Card fee switch:** sudo → Costs → Guardrails → *Card fee | |
| 502 | 504 | on AI credit bought by card* (`cost_settings.card_fee`, `on`/`off`). | |
| 503 | 505 | ||
| 506 | + | **On a workspace's own model key** (migration `0041_agent_rate_own_key.sql`): | |
| 507 | + | the run keeps its model session (`runs.session_id`, `ms_…`) so the proxy's | |
| 508 | + | counts reach it, and the agent rate is charged at `agent_tokens_own` ($0 until | |
| 509 | + | 2026-10-22, then $0.25 a million, a rise from nothing with its notice), on | |
| 510 | + | `<run>/agent-own` (later `<run>/agent-own/<tokens>`), `billed_to = 'g1t'` | |
| 511 | + | (g1t's own charge: it counts toward limits and spend), named *Agent rate, | |
| 512 | + | your own model key* on Usage and the statement. The model is never charged. | |
| 513 | + | `settle_runs` skips these runs (nothing on g1t's gateway); the cron's | |
| 514 | + | `settle_own_runs` closes them 5 minutes after they finish (3 hours after | |
| 515 | + | they start, for a sandbox that never reported) and charges tokens counted | |
| 516 | + | late. Runs from before have no session and are never charged the rate. | |
| 517 | + | ||
| 518 | + | ### The agent rate's token weights | |
| 519 | + | ||
| 520 | + | How much each kind of token counts toward the agent rate, on g1t's models | |
| 521 | + | and own keys alike, is four price-book meters (migration | |
| 522 | + | `0042_agent_rate_weights.sql`): `agent_token_weight_input`, `_output`, | |
| 523 | + | `_cache_read` and `_cache_write`, each a weight in millionths in | |
| 524 | + | `cost_micros` (1,000,000 counts a token once). All start at 1, which is what | |
| 525 | + | the rate always counted. A cached agent run reads most of its context from | |
| 526 | + | cache (about 90% of its tokens on a typical Sonnet implement run), so the | |
| 527 | + | cache-read weight is the lever: at 1 the rate adds about 44% to such a run's | |
| 528 | + | model cost; at 0.1, far less. | |
| 529 | + | ||
| 530 | + | To count cache reads at a tenth: | |
| 531 | + | ||
| 532 | + | 1. Add the version, effective at once (a lower weight is a fall): | |
| 533 | + | ||
| 534 | + | ```sql | |
| 535 | + | INSERT INTO price_versions (id, meter, version, cost_micros, markup_percent, effective_at, reason, created_by, created_at) | |
| 536 | + | VALUES ('pv_agent_token_weight_cache_read_2', 'agent_token_weight_cache_read', 2, 100000, 0, | |
| 537 | + | '2026-10-08T00:00:00Z', 'Cache reads count a tenth toward the agent rate', 'staff', '2026-10-08T00:00:00Z'); | |
| 538 | + | ``` | |
| 539 | + | ||
| 540 | + | in a migration, or with `npx wrangler d1 execute g1t-billing --remote` | |
| 541 | + | until sudo has a form. | |
| 542 | + | 2. The daily run applies it once `effective_at` has come and writes the | |
| 543 | + | public `price_changes` record (**Run the analysis now** applies it at | |
| 544 | + | once). Raising a weight later is a rise: give it an `effective_at` 14 | |
| 545 | + | days out, and owners on the plan are emailed. | |
| 546 | + | 3. Check `/pricing`: the agent rate's row lists the weights, and Usage's | |
| 547 | + | agent-rate lines name them. | |
| 548 | + | ||
| 549 | + | `weighted` (`ai.rs`) rounds down to a whole token; a weight is never below 0. | |
| 550 | + | ||
| 504 | 551 | ### Budgets | |
| 505 | 552 | ||
| 506 | 553 | The owners' spend limit is the budget. `limits.alert_levels` (comma | |
| 532 | 579 | `token_usage` reads a window (42 days by default, 366 at most) for the | |
| 533 | 580 | workspace or one person: totals, every day's tokens and the active days, | |
| 534 | 581 | with `costMicros` the window's run charges from the ledger, measured as | |
| 535 | − | `usage` measures them. These counts are for views only: runs are still | |
| 536 | − | priced from AI Gateway's logs, never from `token_usage`. | |
| 582 | + | `usage` measures them. Usage's report also lists tokens by model | |
| 583 | + | (`UsageReport.models`). The agent rate is charged on these counts (above); | |
| 584 | + | a run's model is still priced from AI Gateway's logs, never from them. | |
| 537 | 585 | ||
| 538 | 586 | ## Tables (migration `0022_costs_and_margin.sql`) | |
| 539 | 587 |
| 1043 | 1043 | - **All hosted model traffic goes through Cloudflare AI Gateway.** That gives | |
| 1044 | 1044 | one place for spend tracking, budgets, rate limits, fallback and logs, | |
| 1045 | 1045 | whichever provider or endpoint is behind it. | |
| 1046 | − | - **Nobody picks a model.** A person assigns work to `g1t`, as they | |
| 1047 | − | would assign an issue to Copilot, and g1t routes it. Today the kind of | |
| 1048 | − | work decides (implementing, reviewing, catching up), from one setting on | |
| 1049 | − | the runner, and each request is tagged at the gateway with that kind, the | |
| 1050 | − | repository and the pull request. The session records which model ran. | |
| 1051 | − | The gateway's own dynamic routes cannot make the choice yet: they work | |
| 1052 | − | only on its OpenAI-compatible endpoint, and the harness speaks | |
| 1053 | − | Anthropic's. | |
| 1046 | + | - **Nobody has to pick a model.** A person assigns work to `g1t`, as | |
| 1047 | + | they would assign an issue to a colleague, and **Auto** routes each job | |
| 1048 | + | to the cheapest model that can do it (see | |
| 1049 | + | [Routing for cost](#routing-for-cost) below). A workspace can pin a tier | |
| 1050 | + | per kind of work instead. Each request is tagged at the gateway with the | |
| 1051 | + | kind of work, the tier, the repository and the pull request, and each | |
| 1052 | + | run says which model ran and why. The gateway's own dynamic routes | |
| 1053 | + | cannot make the choice yet: they work only on its OpenAI-compatible | |
| 1054 | + | endpoint, and the harness speaks Anthropic's. | |
| 1054 | 1055 | - **Subscriptions stay local.** A Claude subscription cannot be used by a | |
| 1055 | 1056 | hosted sandbox; it needs an API key. People on subscriptions use their own | |
| 1056 | 1057 | Claude Code session, which is a full participant. | |
| 1057 | − | - **The workspace pays.** A workspace buys credit by card and each agent | |
| 1058 | − | run deducts what the model cost plus a margin. The billing service asks | |
| 1058 | + | - **The workspace pays.** A workspace buys AI credit by card and each | |
| 1059 | + | agent run deducts the model at the provider's price plus the agent rate | |
| 1060 | + | per million (weighted) tokens; on its own model key, only the agent rate, | |
| 1061 | + | counted from the model proxy and the sandbox's own report. The billing service asks | |
| 1059 | 1062 | nothing of the others: the runner asks it before starting a sandbox and | |
| 1060 | 1063 | is refused when there is no credit, and the sandbox reports what its run | |
| 1061 | 1064 | cost with a token only it holds. Where no card processor is configured | |
| 1100 | 1103 | an assigned issue to a commit on `main` with nobody in between. Required | |
| 1101 | 1104 | human approval per path, and risk tiers, are still to come. | |
| 1102 | 1105 | ||
| 1106 | + | ### Routing for cost | |
| 1107 | + | ||
| 1108 | + | *Built (runner `route` in `services/runner/src/model-env.ts`).* The goal | |
| 1109 | + | is cost per merged change, not cost per request: a cheap attempt that | |
| 1110 | + | fails and is retried on the same model costs more than one that finishes. | |
| 1111 | + | ||
| 1112 | + | - **Tiers and catalogue.** `small` (Claude Haiku 4.5, $1/$5 per million | |
| 1113 | + | input/output), `large` (Claude Sonnet 5.5, $2/$10) and `frontier` | |
| 1114 | + | (Claude Opus 5.5, $4/$20). Models, names and list prices are | |
| 1115 | + | configuration (`AGENT_ROUTING`), never code; prices there are for | |
| 1116 | + | estimates only, runs are charged what AI Gateway priced them at. | |
| 1117 | + | - **Starting tier by job.** Catch-up, answering a question, and reviews of | |
| 1118 | + | at most 10 files and 200 lines touching no sensitive path: small. | |
| 1119 | + | Changes, revisions, plans and other reviews: large. Reviews over 60 files | |
| 1120 | + | or 3,000 lines: frontier. Labels: `architecture` frontier, `security` off | |
| 1121 | + | small, `docs`/`documentation`/`typo` let changes and answers start small. | |
| 1122 | + | - **Escalation.** A failed (or guardrail-stopped) attempt at the same work | |
| 1123 | + | goes one tier up; two in a row, frontier; a revision counts its rounds; | |
| 1124 | + | a change left at low confidence sends the next attempt up. | |
| 1125 | + | - **Learning, per repository.** From the last 20 runs of the same kind: | |
| 1126 | + | one tier down when the cheaper tier finished at least 90% of at least 5 | |
| 1127 | + | (never for sensitive or labelled work, never on a retry); one tier up | |
| 1128 | + | when this tier failed at least half of at least 5. No new tables: it | |
| 1129 | + | reads work's `agent_runs` (model, status, confidence). | |
| 1130 | + | - **Explained.** Every run's first step and session note is one line: | |
| 1131 | + | *Used a fast model (Claude Haiku 4.5): small change, 3 files and 80 | |
| 1132 | + | lines.* | |
| 1133 | + | - **Chosen instead.** `model_routes` rows to g1t's models name `small`, | |
| 1134 | + | `large` or `frontier`, or nothing for Auto (Integrations → Models). | |
| 1135 | + | A workspace's own Anthropic key with no model named is routed by Auto | |
| 1136 | + | too. | |
| 1137 | + | - **Measured.** `scripts/ops/routing-savings.mjs` replays tasks through | |
| 1138 | + | the router offline, priced from the catalogue, against routing before | |
| 1139 | + | Auto and against the frontier model for everything, net of failed | |
| 1140 | + | attempts, with cost per merged change; `--live` reads billing's runs and | |
| 1141 | + | counted tokens. On the bundled sample (13 tasks, assumed failures): Auto | |
| 1142 | + | costs 7% less than the frontier model for everything and about 7% more | |
| 1143 | + | per run than routing before it, but half as much per merged change, | |
| 1144 | + | because it finishes the hard tasks the old routing gave up on. At | |
| 1145 | + | current prices Opus 5.5 and Sonnet 5.5 cost the same per cache read, and | |
| 1146 | + | cache reads are most of an agent run's tokens, so moving off the | |
| 1147 | + | frontier model saves less than its list price suggests; the fast tier | |
| 1148 | + | and fewer failed attempts are where the money is. Run `--live` monthly | |
| 1149 | + | and after any routing change. | |
| 1150 | + | - **A cheaper route for the simplest jobs (designed, off).** A fourth tier | |
| 1151 | + | on Workers AI through AI Gateway (an open model, billed on Cloudflare's | |
| 1152 | + | invoice) for classification-sized jobs: commit messages, triage, | |
| 1153 | + | summaries. Behind the same router as a tier with its own catalogue entry | |
| 1154 | + | and `tasks` rules, off by default. It needs the proxy to translate the | |
| 1155 | + | harness's Anthropic requests to the gateway's OpenAI-compatible | |
| 1156 | + | endpoint (it already does for workspaces' own OpenAI-shaped providers) | |
| 1157 | + | and a quality bar from the savings harness before any job moves to it. | |
| 1158 | + | ||
| 1103 | 1159 | ### Choosing the right agent automatically | |
| 1104 | 1160 | ||
| 1105 | 1161 | Because several agents can work on the same issue, every issue with more |
| 195 | 195 | console.log(`${name.padEnd(9)} ${dollars(total.cost).padStart(10)} merged ${total.merged} per merged change ${dollars(total.perMerged)}`); | |
| 196 | 196 | } | |
| 197 | 197 | console.log(""); | |
| 198 | − | console.log(`Auto against routing before it: ${(savings.vsBefore * 100).toFixed(1)}% ${savings.vsBefore >= 0 ? "less" : "more"}`); | |
| 199 | − | console.log(`Auto against the most capable model for everything: ${(savings.vsFrontier * 100).toFixed(1)}% less`); | |
| 198 | + | const say = (share) => `${Math.abs(share * 100).toFixed(1)}% ${share >= 0 ? "less" : "more"}`; | |
| 199 | + | console.log(`Auto against routing before it: ${say(savings.vsBefore)}`); | |
| 200 | + | console.log(`Auto against the most capable model for everything: ${say(savings.vsFrontier)}`); | |
| 200 | 201 | } | |
| 201 | 202 | ||
| 202 | 203 | async function main() { |