Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier
22 files+470−700/22 viewed
| 434 | 434 | ## Which model runs | |
| 435 | 435 | ||
| 436 | 436 | You do not pick one. You assign the work to `g1t-agent`, the way you would | |
| 437 | − | assign an issue to a colleague, and g1t routes it. The kind of work decides: | |
| 437 | + | assign an issue to a colleague, and g1t routes it. On g1t's hosted models, | |
| 438 | + | each piece of work goes to the least costly of two tiers that can do it: | |
| 438 | 439 | ||
| 439 | − | | Work | Model today | | |
| 440 | + | | Tier | Model today | | |
| 441 | + | | --- | --- | | |
| 442 | + | | Small | Claude Haiku 4.5 | | |
| 443 | + | | Large | Claude Sonnet 5.5 | | |
| 444 | + | ||
| 445 | + | The work decides the tier: | |
| 446 | + | ||
| 447 | + | | Work | Tier | | |
| 440 | 448 | | --- | --- | | |
| 441 | − | | Making a change for an issue, and revising it | Claude Sonnet 5.5 | | |
| 442 | − | | Reviewing a pull request | Claude Sonnet 5.5 | | |
| 443 | − | | Catching up with `main` and resolving conflicts | Claude Sonnet 5.5 | | |
| 444 | − | | Planning an outcome | Claude Sonnet 5.5 | | |
| 449 | + | | Making a change for an issue, revising it, and answering a mention | Large | | |
| 450 | + | | Reviewing a pull request that changes at most 10 files and 200 lines, touches no sensitive path, and is not for an issue labelled `security` | Small | | |
| 451 | + | | Reviewing any other pull request, or one whose changed files g1t does not know yet | Large | | |
| 452 | + | | Catching up with `main` and resolving conflicts | Small | | |
| 453 | + | | Planning an outcome | Small | | |
| 454 | + | | Any of these again, after the last attempt at the same work failed or stopped at a guardrail cap | Large | | |
| 455 | + | ||
| 456 | + | Sensitive paths are the ones that run, configure or guard things: CI | |
| 457 | + | workflows, `.g1t/` and `.github/`, `CODEOWNERS`, secrets such as `.env` | |
| 458 | + | and `.pem` files, and infrastructure such as Dockerfiles, Terraform and | |
| 459 | + | `wrangler.*` files. They are the same paths that lower a change's | |
| 460 | + | [confidence](#how-sure-the-agent-is). | |
| 461 | + | ||
| 462 | + | The agent's own small background steps run on the small tier. | |
| 445 | 463 | ||
| 446 | 464 | Every session opens with a note naming the model that ran, and an agent's | |
| 447 | 465 | review says which model wrote it, so what you got is always on the record. | |
| 448 | − | When a better model for a kind of work appears, g1t changes the route and | |
| 449 | − | nothing you have set up needs to change. | |
| 466 | + | When a better model for a tier appears, g1t changes the route and nothing | |
| 467 | + | you have set up needs to change. | |
| 450 | 468 | ||
| 469 | + | A workspace that routes its work to [its own provider](/guides/models/) | |
| 470 | + | is not routed by tier: its work runs on the model its route names. | |
| 471 | + | ||
| 451 | 472 | A pull request made by a g1t agent carries the label `g1t-agent`, and its | |
| 452 | 473 | commits are authored by `g1t agent`. | |
| 453 | 474 | ||
| 459 | 480 | Requests for g1t's hosted models go on through | |
| 460 | 481 | [Cloudflare AI Gateway](https://developers.cloudflare.com/ai-gateway/), | |
| 461 | 482 | which holds g1t's key. Each of those requests is tagged with the kind of | |
| 462 | − | work, the repository and the pull request, so spend can be read per pull | |
| 463 | − | request. Requests for a workspace's own provider go to that provider. | |
| 483 | + | work, the tier, the repository and the pull request, so spend can be read | |
| 484 | + | per tier and per pull request. Requests for a workspace's own provider go to that provider. | |
| 464 | 485 | ||
| 465 | 486 | If you run your own copy of g1t, these settings control it: | |
| 466 | 487 | ||
| 467 | 488 | | Setting | Where | What it does | | |
| 468 | 489 | | --- | --- | --- | | |
| 469 | − | | `AGENT_ROUTES` | Runner | The model for each kind of work: `implement`, `review`, `update` and `plan`. | | |
| 490 | + | | `AGENT_ROUTING` | Runner | JSON. `tiers`: the model behind `small` and `large`, each `{ "modelName", "model" }`. `tasks`: the tier of `implement`, `review`, `update` and `plan`, or `change` to decide by the change. `smallChange`: the most `files` and `lines` a `change` review runs on the small tier with. `largeLabels`: issue labels that keep a review on the large tier. Anything left out takes the defaults above. | | |
| 470 | 491 | | `MODELS_URL` | Runner | Where sandboxes send model requests: the model proxy. | | |
| 471 | 492 | | `AI_GATEWAY_ID` | Model proxy | The gateway hosted requests go through. Empty sends them to the provider directly. | | |
| 472 | 493 | | `AI_GATEWAY_TOKEN` | Model proxy | Secret. Authenticates to the gateway. | |
| 76 | 76 | ||
| 77 | 77 | Each can go to g1t's models, or to any of your providers on any of its | |
| 78 | 78 | models. An Anthropic provider also offers **g1t's choice of Claude**, | |
| 79 | − | which runs g1t's pick for that kind of work on your key. For example: make | |
| 79 | + | which runs g1t's large-tier model, Claude Sonnet 5.5 today, on your key. | |
| 80 | + | Routing between tiers by the size of the work is only for g1t's hosted | |
| 81 | + | models; see [which model runs](/guides/g1t-agents/#which-model-runs). For example: make | |
| 80 | 82 | changes on Claude through your Anthropic key, review on GPT through your | |
| 81 | 83 | OpenAI key, and catch up on a small model through OpenRouter. | |
| 82 | 84 |
| 225 | 225 | revisions (`max_revisions`, 2 by default) only a required check still | |
| 226 | 226 | failing holds it for a person. There is no model or | |
| 227 | 227 | agent count to choose: to put more agents to work, assign more issues. | |
| 228 | + | On g1t's hosted models, g1t routes each piece of work to a small or a | |
| 229 | + | large tier: planning, catching up and reviews of small changes that | |
| 230 | + | touch no sensitive path run small; making changes, other reviews, and | |
| 231 | + | any retry after a failed attempt run large. | |
| 228 | 232 | - **Put an agent on something in one step:** `POST {repo}/issues/delegate` | |
| 229 | 233 | with `title` and `body` (what to do, in plain words, and what done | |
| 230 | 234 | means if you know it). It opens the issue and assigns g1t-agent at once; it needs the |
| 270 | 270 | /// Gateway: settling charges the run what the gateway priced them at. | |
| 271 | 271 | #[serde(default)] | |
| 272 | 272 | pub session: Option<String>, | |
| 273 | + | /// `small` or `large`: the tier g1t routed the run to, when g1t pays | |
| 274 | + | /// for its model. None on the workspace's own provider. | |
| 275 | + | #[serde(default)] | |
| 276 | + | pub tier: Option<String>, | |
| 273 | 277 | } | |
| 274 | 278 | ||
| 275 | 279 | #[derive(Clone, Debug, Serialize, Deserialize)] |
| 384 | 384 | /// The session's id; see `ModelSession::id`. | |
| 385 | 385 | #[serde(default)] | |
| 386 | 386 | pub session: String, | |
| 387 | + | /// For `g1t`: the tier the run was routed to, `small` or `large`. | |
| 388 | + | #[serde(default)] | |
| 389 | + | pub tier: Option<String>, | |
| 387 | 390 | /// For `endpoint`: where to send requests. | |
| 388 | 391 | pub base_url: Option<String>, | |
| 389 | 392 | /// For `anthropic` and `endpoint`: the workspace's key. | |
| 558 | 561 | /// decides that; this service only follows the routes. | |
| 559 | 562 | #[serde(default = "yes")] | |
| 560 | 563 | pub hosted_open: bool, | |
| 564 | + | /// `small` or `large`: the tier the runner routed the run to on g1t's | |
| 565 | + | /// hosted models, tagged on its requests at the gateway. Kept only | |
| 566 | + | /// when the run goes to g1t's models. | |
| 567 | + | #[serde(default)] | |
| 568 | + | pub tier: Option<String>, | |
| 561 | 569 | } | |
| 562 | 570 | ||
| 563 | 571 | /// `routes`: a workspace's model routes, one per kind of work that has its |
| 1621 | 1621 | pub issue: Option<Issue>, | |
| 1622 | 1622 | /// Who opened the pull request, and so can read its source. | |
| 1623 | 1623 | pub author: User, | |
| 1624 | + | /// The files it changes, as of its latest push: how large the change | |
| 1625 | + | /// is, which decides the model that reviews it. | |
| 1626 | + | #[serde(default)] | |
| 1627 | + | pub files: Vec<ChangedFile>, | |
| 1628 | + | /// What among them runs, configures or guards things (CI workflows, | |
| 1629 | + | /// secrets, infrastructure), once each. Any sends the review to the | |
| 1630 | + | /// larger model. | |
| 1631 | + | #[serde(default)] | |
| 1632 | + | pub sensitive: Vec<String>, | |
| 1624 | 1633 | } | |
| 1625 | 1634 | ||
| 1626 | 1635 | /// A comment on one line, as a reviewing agent reports it. |
| 73 | 73 | | **Cache API** | `services/repos/src/store.rs` (trees and blobs up to 1 MiB, by hash), `apps/web/workers/app.ts` (avatars), `services/og` | thin, optional | Miniflare's cache, or none. Every use tolerates a miss. | | |
| 74 | 74 | | **Vectorize** | `services/context/src/index.ts` (`VECTORS.upsert/deleteByIds/query`); optional, guarded by `if (!AI \|\| !VECTORS)` | thin | **sqlite-vec** (default: one file, next to D1), pgvector or Qdrant behind a `VectorIndex` port; or off, which already degrades to keyword search. | | |
| 75 | 75 | | **Workers AI** | `services/context` only: `@cf/baai/bge-base-en-v1.5` embeddings (768 dims) | thin | An OpenAI-compatible `/v1/embeddings` endpoint (Ollama, vLLM, LM Studio, or a hosted API) behind an `Embedder` port. Changing models means re-embedding (a backfill job already exists). | | |
| 76 | − | | **AI Gateway** | `services/runner/src/model-env.ts:56`, `services/models/src/route.ts:70` (gateway URL, `cf-aig-*` headers), `services/billing/src/keeper.rs` (reads gateway logs to settle) | thin | Optional already: an empty `AI_GATEWAY_ID` goes straight to the provider. Any Anthropic- or OpenAI-compatible base URL works for a workspace's own provider. | | |
| 76 | + | | **AI Gateway** | `services/runner/src/model-env.ts` (`modelEnv`), `services/models/src/route.ts:70` (gateway URL, `cf-aig-*` headers), `services/billing/src/keeper.rs` (reads gateway logs to settle) | thin | Optional already: an empty `AI_GATEWAY_ID` goes straight to the provider. Any Anthropic- or OpenAI-compatible base URL works for a workspace's own provider. | | |
| 77 | 77 | | **Workers for Platforms** | `services/pages` (dispatcher: `env.APPS.get(script).fetch`), `services/deployments/src/cloudflare.ts` (script and asset upload through the REST API) | woven | Phase 1: off. Later: a self-hosted app host in workerd, using the Worker Loader binding to load uploaded scripts, or a workerd per app (see [Deployments](#deployments)). | | |
| 78 | 78 | | **Cloudflare for SaaS** (custom hostnames) | `services/deployments/src/custom-hostnames.ts` (`/zones/{id}/custom_hostnames`) | thin | Caddy with on-demand TLS, asking g1t whether a hostname is allowed. | | |
| 79 | 79 | | **Cloudflare REST API** | deployments: script upload, list, delete, assets, GraphQL usage. Billing keeper: AI Gateway logs, `billable-usage`, GraphQL container usage. Ops scripts in `scripts/`. | thin (deployments), woven (keeper pricing) | Deployments: the app-host adapter. Keeper: off when self-hosted, because there is no bill to reconcile. | |
| 125 | 125 | computes them (`overlaps`). | |
| 126 | 126 | - Previous reviews of the same PR, beyond comments via `peopleSaid`. | |
| 127 | 127 | ||
| 128 | − | **Model.** The production route for `review` is **Claude Sonnet 5.5** | |
| 129 | − | (`claude-sonnet-5-5`, `AGENT_ROUTES` in `services/runner/wrangler.jsonc`). | |
| 128 | + | **Model.** The production route for `review` was **Claude Sonnet 5.5** | |
| 129 | + | (`claude-sonnet-5-5`) when this was measured. Reviews are now routed by | |
| 130 | + | tier (`AGENT_ROUTING` in `services/runner/wrangler.jsonc`): small changes | |
| 131 | + | that touch no sensitive path go to the small tier (Claude Haiku 4.5), the | |
| 132 | + | rest stay on Sonnet 5.5. | |
| 130 | 133 | A workspace can route reviews to its own provider through the model proxy | |
| 131 | 134 | (`openModelSession`). No effort level is set, so it uses the Claude Code | |
| 132 | − | default. The model-env test uses Opus 5.5 for review as a fixture only. | |
| 135 | + | default. | |
| 133 | 136 | ||
| 134 | 137 | **Outputs and post-processing** (`reviews.rs`, `report_review`): | |
| 135 | 138 | ||
| 318 | 321 | high-severity comment counts for more than three nits. | |
| 319 | 322 | 6. **Model and effort.** Run the same sample on Sonnet 5.5 at | |
| 320 | 323 | `medium`/`high` effort and on Opus 5.5. The review route can change in | |
| 321 | − | `AGENT_ROUTES` without a code change. Opus 5.5 is about 1.4x the cost | |
| 324 | + | `AGENT_ROUTING` without a code change. Opus 5.5 is about 1.4x the cost | |
| 322 | 325 | per review at our token profile. | |
| 323 | 326 | 7. **Product-side follow-ups** (not benchmark-visible): | |
| 324 | 327 | - allow file-level comments on unchanged files when the change breaks |
| 938 | 938 | billedTo?: "g1t" | "workspace"; | |
| 939 | 939 | /** The model session's id, so the run can be settled at AI Gateway's price. */ | |
| 940 | 940 | session?: string | null; | |
| 941 | + | /** `small` or `large`: the tier g1t routed the run to, on its hosted models. */ | |
| 942 | + | tier?: "small" | "large" | null; | |
| 941 | 943 | }): Promise<Result<RunTicket | null>>; | |
| 942 | 944 | } | |
| 943 | 945 |
| 142 | 142 | task: string; | |
| 143 | 143 | /** The session's id; see `ModelSession.id`. */ | |
| 144 | 144 | session: string; | |
| 145 | + | /** For `g1t`: the tier the run was routed to. */ | |
| 146 | + | tier?: "small" | "large" | null; | |
| 145 | 147 | baseUrl: string | null; | |
| 146 | 148 | apiKey: string | null; | |
| 147 | 149 | authHeader: string | null; | |
| 183 | 185 | number: number; | |
| 184 | 186 | task: string; | |
| 185 | 187 | hostedOpen: boolean; | |
| 188 | + | /** The tier the run is routed to on g1t's hosted models, for the gateway's logs. */ | |
| 189 | + | tier?: "small" | "large" | null; | |
| 186 | 190 | }): Promise<Result<ModelSession>>; | |
| 187 | 191 | routes(workspace: string, viewer: Viewer): Promise<Result<ModelRoute[]>>; | |
| 188 | 192 | setRoutes(actor: User, workspace: string, routes: ModelRoute[]): Promise<Result<ModelRoute[]>>; |
| 660 | 660 | issue: Issue | null; | |
| 661 | 661 | /** Who opened the pull request, and so can read its source. */ | |
| 662 | 662 | author: User; | |
| 663 | + | /** | |
| 664 | + | * The files it changes, as of its latest push: how large the change is, | |
| 665 | + | * which decides the model that reviews it. | |
| 666 | + | */ | |
| 667 | + | files: ChangedFile[]; | |
| 668 | + | /** | |
| 669 | + | * What among them runs, configures or guards things (CI workflows, | |
| 670 | + | * secrets, infrastructure), once each. Any sends the review to the | |
| 671 | + | * larger model. | |
| 672 | + | */ | |
| 673 | + | sensitive: string[]; | |
| 663 | 674 | }; | |
| 664 | 675 | ||
| 665 | 676 | export type OpenIssueInput = { |
| 1 | + | -- The tier g1t routed a run to on its hosted models, small or large, so | |
| 2 | + | -- what the model cost can be read per tier. Null on a workspace's own | |
| 3 | + | -- provider, and for runs from before routing by tier. | |
| 4 | + | ALTER TABLE runs ADD COLUMN tier TEXT; |
| 639 | 639 | let token = hex::encode(bytes); | |
| 640 | 640 | self.db | |
| 641 | 641 | .prepare( | |
| 642 | − | "INSERT INTO runs (id, workspace, repo, number, task, model, token_hash, created_at, billed_to, session_id) | |
| 643 | − | VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)", | |
| 642 | + | "INSERT INTO runs (id, workspace, repo, number, task, model, token_hash, created_at, billed_to, session_id, tier) | |
| 643 | + | VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)", | |
| 644 | 644 | ) | |
| 645 | 645 | .bind(&[ | |
| 646 | 646 | run_id.as_str().into(), | |
| 653 | 653 | rfc3339(now).into(), | |
| 654 | 654 | if a.billed_to == "workspace" { "workspace" } else { "g1t" }.into(), | |
| 655 | 655 | optional(a.session.as_deref().filter(|_| a.billed_to != "workspace")), | |
| 656 | + | optional( | |
| 657 | + | a.tier | |
| 658 | + | .as_deref() | |
| 659 | + | .filter(|tier| a.billed_to != "workspace" && matches!(*tier, "small" | "large")), | |
| 660 | + | ), | |
| 656 | 661 | ])? | |
| 657 | 662 | .run() | |
| 658 | 663 | .await?; |
| 1 | + | -- The tier a run on g1t's hosted models was routed to, small or large. | |
| 2 | + | -- The model proxy tags each of the run's requests with it at the gateway. | |
| 3 | + | -- Null for a workspace's own provider. | |
| 4 | + | ALTER TABLE model_sessions ADD COLUMN tier TEXT; |
| 137 | 137 | number: u32, | |
| 138 | 138 | task: String, | |
| 139 | 139 | model: Option<String>, | |
| 140 | + | #[serde(default)] | |
| 141 | + | tier: Option<String>, | |
| 140 | 142 | } | |
| 141 | 143 | ||
| 142 | 144 | #[derive(Deserialize)] | |
| 1227 | 1229 | .bind(&[rfc3339(now).into()])?, | |
| 1228 | 1230 | self.db | |
| 1229 | 1231 | .prepare( | |
| 1230 | − | "INSERT INTO model_sessions (token_hash, workspace, connection_id, repo, number, task, expires_at, model) | |
| 1231 | − | VALUES (?, ?, ?, ?, ?, ?, ?, ?)", | |
| 1232 | + | "INSERT INTO model_sessions (token_hash, workspace, connection_id, repo, number, task, expires_at, model, tier) | |
| 1233 | + | VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)", | |
| 1232 | 1234 | ) | |
| 1233 | 1235 | .bind(&[ | |
| 1234 | 1236 | crypto::sha256_hex(&token).into(), | |
| 1239 | 1241 | a.task.as_str().into(), | |
| 1240 | 1242 | rfc3339(now + MODEL_SESSION_SECONDS * 1000).into(), | |
| 1241 | 1243 | optional(model.as_deref()), | |
| 1244 | + | // The tier is g1t's routing; it means nothing on the workspace's own provider. | |
| 1245 | + | optional( | |
| 1246 | + | a.tier | |
| 1247 | + | .as_deref() | |
| 1248 | + | .filter(|tier| connection.is_none() && matches!(*tier, "small" | "large")), | |
| 1249 | + | ), | |
| 1242 | 1250 | ])?, | |
| 1243 | 1251 | ]) | |
| 1244 | 1252 | .await?; | |
| 1273 | 1281 | number: session.number, | |
| 1274 | 1282 | task: session.task, | |
| 1275 | 1283 | session: session_id(&session.token_hash), | |
| 1284 | + | tier: session.tier, | |
| 1276 | 1285 | base_url: None, | |
| 1277 | 1286 | api_key: None, | |
| 1278 | 1287 | auth_header: None, |
| 39 | 39 | }); | |
| 40 | 40 | }); | |
| 41 | 41 | ||
| 42 | + | test("a run's tier is tagged at the gateway, so spend can be read per tier", () => { | |
| 43 | + | const upstream: ModelUpstream = { ...run, route: "g1t", tier: "small" }; | |
| 44 | + | const { headers } = upstreamRequest(upstream, hosted, "/v1/messages", incoming()); | |
| 45 | + | const metadata = JSON.parse(headers.get("cf-aig-metadata") ?? "{}"); | |
| 46 | + | assert.deepEqual(metadata, { | |
| 47 | + | task: "implement", | |
| 48 | + | tier: "small", | |
| 49 | + | repo: "acme/web", | |
| 50 | + | pull: 7, | |
| 51 | + | session: "ms_abc", | |
| 52 | + | }); | |
| 53 | + | // The gateway keeps five entries; the session, which billing settles | |
| 54 | + | // by, must be one of them. | |
| 55 | + | assert.ok(Object.keys(metadata).length <= 5); | |
| 56 | + | }); | |
| 57 | + | ||
| 42 | 58 | test("a workspace's own Anthropic key goes to Anthropic, and only there", () => { | |
| 43 | 59 | const upstream: ModelUpstream = { ...run, route: "anthropic", baseUrl: "https://api.anthropic.com", apiKey: "sk-ant-theirs", authHeader: "x-api-key" }; | |
| 44 | 60 | const { url, headers } = upstreamRequest(upstream, hosted, "/v1/messages", incoming()); |
| 52 | 52 | return { url: `https://api.anthropic.com${path}`, headers }; | |
| 53 | 53 | } | |
| 54 | 54 | // The gateway logs these with every request, so spend and failures can | |
| 55 | − | // be read per kind of work, workspace, repository and pull request. | |
| 55 | + | // be read per kind of work, tier, repository and pull request. It keeps | |
| 56 | + | // five entries and drops the rest, so the tier takes the workspace's | |
| 57 | + | // place: the repository names the workspace too. A run from before | |
| 58 | + | // routing by tier has no tier, and keeps the workspace. | |
| 56 | 59 | headers.set( | |
| 57 | 60 | "cf-aig-metadata", | |
| 58 | 61 | JSON.stringify({ | |
| 59 | 62 | task: upstream.task, | |
| 60 | − | workspace: upstream.workspace, | |
| 63 | + | ...(upstream.tier ? { tier: upstream.tier } : { workspace: upstream.workspace }), | |
| 61 | 64 | repo: upstream.repo, | |
| 62 | 65 | pull: upstream.number, | |
| 63 | 66 | // What billing finds the run's requests by, to charge what they cost. |
| 60 | 60 | instanceNamed, | |
| 61 | 61 | } from "@g1t/contracts"; | |
| 62 | 62 | ||
| 63 | − | import { type AgentRoutes, type AgentTask, canReachModel, modelEnv } from "./model-env"; | |
| 63 | + | import { | |
| 64 | + | type AgentTask, | |
| 65 | + | type RouteSignals, | |
| 66 | + | type Tier, | |
| 67 | + | canReachModel, | |
| 68 | + | changeSize, | |
| 69 | + | chooseTier, | |
| 70 | + | lastAttemptFailed, | |
| 71 | + | modelEnv, | |
| 72 | + | parseRouting, | |
| 73 | + | tierVars, | |
| 74 | + | } from "./model-env"; | |
| 64 | 75 | import { hubContext } from "./hub"; | |
| 65 | 76 | import { hostedOpen } from "./hosted"; | |
| 66 | 77 | import { delegateInput, noModelMessage, notStarted, queued, started } from "./delegate"; | |
| 137 | 148 | */ | |
| 138 | 149 | HOSTED_AGENT_WORKSPACES: string; | |
| 139 | 150 | /** | |
| 140 | − | * Which model each kind of work runs on, as JSON: | |
| 141 | − | * `{ implement, review, update }`, each `{ modelName, model }`. | |
| 142 | − | * `modelName` is what people see; `model` is sent to the provider. | |
| 151 | + | * How g1t routes the work it pays the model for, as JSON (`AgentRouting` | |
| 152 | + | * in model-env.ts): `tiers`, the model behind `small` and `large`, each | |
| 153 | + | * `{ modelName, model }`; `tasks`, the tier of each kind of work, or | |
| 154 | + | * `change` to decide a review by its change; `smallChange`, the largest | |
| 155 | + | * change reviewed on the small tier; `largeLabels`, issue labels that | |
| 156 | + | * keep a review large. Anything left out takes the default. | |
| 143 | 157 | */ | |
| 144 | − | AGENT_ROUTES: string; | |
| 158 | + | AGENT_ROUTING?: string; | |
| 145 | 159 | /** | |
| 146 | 160 | * A Cloudflare AI Gateway id. When set, model traffic goes through that | |
| 147 | 161 | * gateway, which is where logging, spend limits, caching and fallback | |
| 167 | 181 | ABUSE_WATCH?: string; | |
| 168 | 182 | } | |
| 169 | 183 | ||
| 184 | + | /** | |
| 185 | + | * What routing knows about one piece of work, and how to ask whether it is | |
| 186 | + | * a retry (asked only when the answer matters). | |
| 187 | + | */ | |
| 188 | + | type RouteInput = RouteSignals & { retried?: () => Promise<boolean> }; | |
| 189 | + | ||
| 170 | 190 | /** A run that takes longer than this has its token expire under it. */ | |
| 171 | 191 | const TOKEN_TTL_SECONDS = 2 * 60 * 60; | |
| 172 | 192 | /** How g1t's own agent is labelled. What runs behind it is g1t's choice. */ | |
| 1308 | 1328 | * What a sandbox needs to reach the model routed for `task`, having | |
| 1309 | 1329 | * opened the run the repository's workspace will be charged for. Refused | |
| 1310 | 1330 | * when that workspace has no credit. | |
| 1331 | + | * | |
| 1332 | + | * On g1t's hosted models the work goes to the cheapest tier that can do | |
| 1333 | + | * it (`chooseTier`), by what `route` says about it. Whether this is a | |
| 1334 | + | * retry is asked only when it would change the answer: when the work | |
| 1335 | + | * would otherwise go to the small tier. | |
| 1311 | 1336 | */ | |
| 1312 | 1337 | private async modelEnv( | |
| 1313 | 1338 | task: AgentTask, | |
| 1314 | 1339 | repo: RepoPath, | |
| 1315 | 1340 | pull: number, | |
| 1341 | + | route: RouteInput = {}, | |
| 1316 | 1342 | ): Promise<Result<Record<string, string>>> { | |
| 1317 | − | const routes: AgentRoutes = JSON.parse(this.env.AGENT_ROUTES); | |
| 1343 | + | const routing = parseRouting(this.env.AGENT_ROUTING); | |
| 1344 | + | let tier: Tier = chooseTier(task, route, routing); | |
| 1345 | + | if (tier === "small" && route.retried && (await route.retried().catch(() => false))) { | |
| 1346 | + | tier = chooseTier(task, { ...route, retry: true }, routing); | |
| 1347 | + | } | |
| 1318 | 1348 | const tags = { repo: `${repo.namespace}/${repo.name}`, pull }; | |
| 1319 | 1349 | // Where the run's model requests go, by the workspace's routes: g1t's | |
| 1320 | 1350 | // hosted models, or one of its own providers. | |
| 1326 | 1356 | number: pull, | |
| 1327 | 1357 | task, | |
| 1328 | 1358 | hostedOpen: (await this.modelAccess(repo.namespace)).hosted, | |
| 1359 | + | tier, | |
| 1329 | 1360 | }); | |
| 1330 | 1361 | if (!opened.ok) return opened; | |
| 1331 | 1362 | session = opened.value; | |
| 1332 | 1363 | } | |
| 1333 | 1364 | const own = session?.billedTo === "workspace"; | |
| 1334 | − | const model = session?.model ?? routes[task].model; | |
| 1335 | − | const modelName = session?.model ?? routes[task].modelName; | |
| 1365 | + | // A workspace's own provider is not routed by tier: it runs the model | |
| 1366 | + | // its route names, or for an Anthropic provider, the large tier's. | |
| 1367 | + | const routed = routing.tiers[own ? "large" : tier]; | |
| 1368 | + | const model = session?.model ?? routed.model; | |
| 1369 | + | const modelName = session?.model ?? routed.modelName; | |
| 1336 | 1370 | const ticket = await billingClient(this.env.BILLING).startRun({ | |
| 1337 | 1371 | workspace: repo.namespace, | |
| 1338 | 1372 | repo, | |
| 1341 | 1375 | model: own ? `${modelName} (${session?.providerName ?? "own provider"})` : modelName, | |
| 1342 | 1376 | billedTo: own ? "workspace" : "g1t", | |
| 1343 | 1377 | session: own ? null : (session?.id ?? null), | |
| 1378 | + | tier: own ? null : tier, | |
| 1344 | 1379 | }); | |
| 1345 | 1380 | if (!ticket.ok) return ticket; | |
| 1346 | 1381 | const vars: Record<string, string> = session | |
| 1347 | 1382 | ? { | |
| 1383 | + | // On g1t's models, the tier's model, and the small tier's for the | |
| 1384 | + | // harness's own small tasks. | |
| 1385 | + | ...(own ? {} : tierVars(routing, tier)), | |
| 1348 | 1386 | ANTHROPIC_MODEL: model, | |
| 1349 | 1387 | AGENT_MODEL_NAME: own ? `${modelName}, through ${session.providerName}` : modelName, | |
| 1350 | 1388 | ANTHROPIC_BASE_URL: `${this.env.MODELS_URL!.replace(/\/+$/, "")}/anthropic`, | |
| 1354 | 1392 | // the harness's small tasks too. | |
| 1355 | 1393 | ...(session.model ? { ANTHROPIC_SMALL_FAST_MODEL: session.model } : {}), | |
| 1356 | 1394 | } | |
| 1357 | − | : modelEnv(this.env, routes, task, tags); | |
| 1395 | + | : modelEnv(this.env, routing, task, tier, tags); | |
| 1358 | 1396 | if (ticket.value) { | |
| 1359 | 1397 | // How the sandbox says what the run cost. Kept from the agent. | |
| 1360 | 1398 | vars.BILLING_RUN = ticket.value.runId; | |
| 1490 | 1528 | task: AgentTask, | |
| 1491 | 1529 | repo: RepoPath, | |
| 1492 | 1530 | pull: number, | |
| 1531 | + | route: RouteInput = {}, | |
| 1493 | 1532 | ): Promise<Record<string, string>> { | |
| 1494 | − | const vars = await this.modelEnv(task, repo, pull); | |
| 1533 | + | const vars = await this.modelEnv(task, repo, pull, route); | |
| 1495 | 1534 | if (!vars.ok) throw new Error(vars.error.message); | |
| 1496 | 1535 | return vars.value; | |
| 1497 | 1536 | } | |
| 1498 | 1537 | ||
| 1538 | + | /** | |
| 1539 | + | * Whether the latest run of the same work failed, so that this one is a | |
| 1540 | + | * retry: the same kind of run on the same pull request, or for a plan, | |
| 1541 | + | * the latest plan with the same brief. Read as the one the run is for; | |
| 1542 | + | * unknown counts as not. | |
| 1543 | + | */ | |
| 1544 | + | private async failedBefore( | |
| 1545 | + | viewer: User, | |
| 1546 | + | repo: RepoPath, | |
| 1547 | + | kind: "review" | "update" | "plan", | |
| 1548 | + | number: number | null, | |
| 1549 | + | title?: string, | |
| 1550 | + | ): Promise<boolean> { | |
| 1551 | + | const runs = await agentsClient(this.env.WORK) | |
| 1552 | + | .listRuns(viewer, { repo, kind, ...(number != null ? { number } : {}), limit: 1 }) | |
| 1553 | + | .catch(() => null); | |
| 1554 | + | return runs?.ok ? lastAttemptFailed(runs.value, title) : false; | |
| 1555 | + | } | |
| 1556 | + | ||
| 1499 | 1557 | /** Whether sandboxes have a way to reach a model at all. */ | |
| 1500 | 1558 | private modelsReachable(): boolean { | |
| 1501 | 1559 | return Boolean(this.env.MODELS_URL) || canReachModel(this.env); | |
| 2432 | 2490 | repo, | |
| 2433 | 2491 | actor, | |
| 2434 | 2492 | ), | |
| 2435 | − | ...(await this.modelEnvOrThrow("update", repo, number)), | |
| 2493 | + | ...(await this.modelEnvOrThrow("update", repo, number, { | |
| 2494 | + | retried: () => this.failedBefore(actor, repo, "update", number), | |
| 2495 | + | })), | |
| 2436 | 2496 | }, | |
| 2437 | 2497 | }); | |
| 2438 | 2498 | } | |
| 2486 | 2546 | `It is for issue #${job.issue.number}: ${job.issue.title}\n\n${job.issue.body}`, | |
| 2487 | 2547 | await this.peopleSaid(job.author, repo, number), | |
| 2488 | 2548 | ]; | |
| 2489 | − | const model = await this.modelEnv("review", repo, number); | |
| 2549 | + | const model = await this.modelEnv("review", repo, number, { | |
| 2550 | + | change: job.files?.length ? changeSize(job.files, job.sensitive ?? []) : null, | |
| 2551 | + | labels: job.issue?.labels ?? [], | |
| 2552 | + | retried: () => this.failedBefore(job.author, repo, "review", number), | |
| 2553 | + | }); | |
| 2490 | 2554 | if (!model.ok) { | |
| 2491 | 2555 | await workClient(this.env.WORK).failReview(job.runId, job.token, model.error.message); | |
| 2492 | 2556 | return model; | |
| 2546 | 2610 | const started = await work.startPlan(actor, repo, brief); | |
| 2547 | 2611 | if (!started.ok) return started; | |
| 2548 | 2612 | const job = started.value; | |
| 2549 | − | const model = await this.modelEnv("plan", repo, 0); | |
| 2613 | + | const model = await this.modelEnv("plan", repo, 0, { | |
| 2614 | + | retried: () => this.failedBefore(actor, repo, "plan", null, job.brief), | |
| 2615 | + | }); | |
| 2550 | 2616 | if (!model.ok) { | |
| 2551 | 2617 | await work.failPlan(job.planId, job.token, model.error.message); | |
| 2552 | 2618 | return model; |
| 2 | 2 | import { createServer } from "node:http"; | |
| 3 | 3 | import { test } from "node:test"; | |
| 4 | 4 | ||
| 5 | − | import { type AgentRoutes, canReachModel, modelEnv } from "./model-env.ts"; | |
| 5 | + | import { | |
| 6 | + | type AgentRouting, | |
| 7 | + | type ChangeSize, | |
| 8 | + | DEFAULT_ROUTING, | |
| 9 | + | canReachModel, | |
| 10 | + | changeSize, | |
| 11 | + | chooseTier, | |
| 12 | + | lastAttemptFailed, | |
| 13 | + | modelEnv, | |
| 14 | + | parseRouting, | |
| 15 | + | } from "./model-env.ts"; | |
| 6 | 16 | ||
| 7 | − | const routes: AgentRoutes = { | |
| 8 | − | implement: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" }, | |
| 9 | − | review: { modelName: "Claude Opus 5.5", model: "claude-opus-5-5" }, | |
| 10 | − | update: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" }, | |
| 11 | − | plan: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" }, | |
| 17 | + | const routes: AgentRouting = { | |
| 18 | + | ...DEFAULT_ROUTING, | |
| 19 | + | tiers: { | |
| 20 | + | small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" }, | |
| 21 | + | large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" }, | |
| 22 | + | }, | |
| 12 | 23 | }; | |
| 13 | 24 | const tags = { repo: "acme/site", pull: 12 }; | |
| 14 | 25 | const direct = { ANTHROPIC_API_KEY: "sk-test", AI_GATEWAY_ID: "", CLOUDFLARE_ACCOUNT_ID: "acct" }; | |
| 23 | 34 | ); | |
| 24 | 35 | } | |
| 25 | 36 | ||
| 26 | − | test("the kind of work decides the model", () => { | |
| 27 | − | assert.equal(modelEnv(direct, routes, "implement", tags).ANTHROPIC_MODEL, "claude-sonnet-5-5"); | |
| 28 | − | const review = modelEnv(direct, routes, "review", tags); | |
| 29 | − | assert.equal(review.ANTHROPIC_MODEL, "claude-opus-5-5"); | |
| 30 | − | assert.equal(review.AGENT_MODEL_NAME, "Claude Opus 5.5"); | |
| 37 | + | const small: ChangeSize = { files: 3, lines: 80, sensitive: [] }; | |
| 38 | + | ||
| 39 | + | test("planning and catching up run on the small tier, making a change on the large", () => { | |
| 40 | + | assert.equal(chooseTier("plan", {}, routes), "small"); | |
| 41 | + | assert.equal(chooseTier("update", {}, routes), "small"); | |
| 42 | + | assert.equal(chooseTier("implement", {}, routes), "large"); | |
| 43 | + | assert.equal(chooseTier("implement", { change: small }, routes), "large"); | |
| 44 | + | }); | |
| 45 | + | ||
| 46 | + | test("a review is small only for a small change that touches nothing sensitive", () => { | |
| 47 | + | assert.equal(chooseTier("review", { change: small }, routes), "small"); | |
| 48 | + | assert.equal(chooseTier("review", { change: { ...small, lines: 200, files: 10 } }, routes), "small"); | |
| 49 | + | assert.equal(chooseTier("review", { change: { ...small, lines: 201 } }, routes), "large"); | |
| 50 | + | assert.equal(chooseTier("review", { change: { ...small, files: 11 } }, routes), "large"); | |
| 51 | + | assert.equal(chooseTier("review", { change: { ...small, sensitive: ["CI workflows"] } }, routes), "large"); | |
| 52 | + | assert.equal(chooseTier("review", { change: small, labels: ["Security"] }, routes), "large"); | |
| 53 | + | assert.equal(chooseTier("review", { change: small, labels: ["docs"] }, routes), "small"); | |
| 54 | + | }); | |
| 55 | + | ||
| 56 | + | test("a review of a change g1t cannot size runs on the large tier", () => { | |
| 57 | + | assert.equal(chooseTier("review", {}, routes), "large"); | |
| 58 | + | assert.equal(chooseTier("review", { change: null }, routes), "large"); | |
| 59 | + | assert.equal(chooseTier("review", { change: { files: 0, lines: 0, sensitive: [] } }, routes), "large"); | |
| 60 | + | }); | |
| 61 | + | ||
| 62 | + | test("a retry after a failed attempt goes up to the large tier", () => { | |
| 63 | + | assert.equal(chooseTier("plan", { retry: true }, routes), "large"); | |
| 64 | + | assert.equal(chooseTier("update", { retry: true }, routes), "large"); | |
| 65 | + | assert.equal(chooseTier("review", { change: small, retry: true }, routes), "large"); | |
| 66 | + | }); | |
| 67 | + | ||
| 68 | + | test("a retry is the same work again after its latest attempt failed", () => { | |
| 69 | + | assert.equal(lastAttemptFailed([]), false); | |
| 70 | + | assert.equal(lastAttemptFailed([{ status: "failed" }, { status: "succeeded" }]), true); | |
| 71 | + | assert.equal(lastAttemptFailed([{ status: "succeeded" }, { status: "failed" }]), false); | |
| 72 | + | assert.equal(lastAttemptFailed([{ status: "stopped", halted: "budget" }]), true); | |
| 73 | + | assert.equal(lastAttemptFailed([{ status: "stopped", halted: null }]), false); | |
| 74 | + | assert.equal(lastAttemptFailed([{ status: "running" }]), false); | |
| 75 | + | assert.equal(lastAttemptFailed([{ status: "failed", title: "Add search" }], "Add search "), true); | |
| 76 | + | assert.equal(lastAttemptFailed([{ status: "failed", title: "Add search" }], "Add billing"), false); | |
| 77 | + | }); | |
| 78 | + | ||
| 79 | + | test("the configuration decides the tiers, the rules and the limits", () => { | |
| 80 | + | const parsed = parseRouting( | |
| 81 | + | JSON.stringify({ | |
| 82 | + | tiers: { small: { modelName: "Small", model: "small-1" } }, | |
| 83 | + | tasks: { update: "large" }, | |
| 84 | + | smallChange: { lines: 50 }, | |
| 85 | + | }), | |
| 86 | + | ); | |
| 87 | + | assert.deepEqual(parsed.tiers.small, { modelName: "Small", model: "small-1" }); | |
| 88 | + | assert.deepEqual(parsed.tiers.large, DEFAULT_ROUTING.tiers.large); | |
| 89 | + | assert.equal(chooseTier("update", {}, parsed), "large"); | |
| 90 | + | assert.equal(chooseTier("plan", {}, parsed), "small"); | |
| 91 | + | assert.equal(chooseTier("review", { change: { ...small, lines: 51 } }, parsed), "large"); | |
| 92 | + | assert.equal(chooseTier("review", { change: { ...small, files: 10, lines: 50 } }, parsed), "small"); | |
| 93 | + | assert.deepEqual(parseRouting(undefined), DEFAULT_ROUTING); | |
| 94 | + | assert.deepEqual(parseRouting("not json"), DEFAULT_ROUTING); | |
| 31 | 95 | }); | |
| 32 | 96 | ||
| 97 | + | test("a change's size is its files and the lines added and removed", () => { | |
| 98 | + | assert.deepEqual( | |
| 99 | + | changeSize([{ additions: 10, deletions: 2 }, { additions: 0, deletions: 5 }], ["secrets"]), | |
| 100 | + | { files: 2, lines: 17, sensitive: ["secrets"] }, | |
| 101 | + | ); | |
| 102 | + | }); | |
| 103 | + | ||
| 104 | + | test("the tier decides the model, and the harness's small tasks use the small tier", () => { | |
| 105 | + | const large = modelEnv(direct, routes, "implement", "large", tags); | |
| 106 | + | assert.equal(large.ANTHROPIC_MODEL, "claude-sonnet-5-5"); | |
| 107 | + | assert.equal(large.AGENT_MODEL_NAME, "Claude Sonnet 5.5"); | |
| 108 | + | assert.equal(large.ANTHROPIC_SMALL_FAST_MODEL, "claude-haiku-4-5-20251001"); | |
| 109 | + | assert.equal(large.ANTHROPIC_DEFAULT_HAIKU_MODEL, "claude-haiku-4-5-20251001"); | |
| 110 | + | const review = modelEnv(direct, routes, "review", "small", tags); | |
| 111 | + | assert.equal(review.ANTHROPIC_MODEL, "claude-haiku-4-5-20251001"); | |
| 112 | + | assert.equal(review.AGENT_MODEL_NAME, "Claude Haiku 4.5"); | |
| 113 | + | }); | |
| 114 | + | ||
| 33 | 115 | test("without a gateway, requests go to the provider directly", () => { | |
| 34 | − | const vars = modelEnv(direct, routes, "implement", tags); | |
| 116 | + | const vars = modelEnv(direct, routes, "implement", "large", tags); | |
| 35 | 117 | assert.equal(vars.ANTHROPIC_BASE_URL, undefined); | |
| 36 | 118 | assert.equal(vars.ANTHROPIC_CUSTOM_HEADERS, undefined); | |
| 37 | 119 | assert.equal(vars.ANTHROPIC_API_KEY, "sk-test"); | |
| 38 | 120 | }); | |
| 39 | 121 | ||
| 40 | 122 | test("with a gateway, requests go through it and say what they are for", () => { | |
| 41 | − | const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t" }, routes, "review", tags); | |
| 123 | + | const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t" }, routes, "review", "large", tags); | |
| 42 | 124 | assert.equal(vars.ANTHROPIC_BASE_URL, "https://gateway.ai.cloudflare.com/v1/acct/g1t/anthropic"); | |
| 43 | 125 | assert.deepEqual(customHeaders(vars), { | |
| 44 | − | "cf-aig-metadata": '{"task":"review","repo":"acme/site","pull":12}', | |
| 126 | + | "cf-aig-metadata": '{"task":"review","tier":"large","repo":"acme/site","pull":12}', | |
| 45 | 127 | }); | |
| 46 | 128 | }); | |
| 47 | 129 | ||
| 48 | 130 | test("an authenticated gateway is sent its token", () => { | |
| 49 | − | const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", tags); | |
| 131 | + | const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", "large", tags); | |
| 50 | 132 | assert.equal(customHeaders(vars)["cf-aig-authorization"], "Bearer tok"); | |
| 51 | 133 | assert.equal(vars.ANTHROPIC_API_KEY, "sk-test"); | |
| 52 | 134 | }); | |
| 54 | 136 | test("when the gateway holds the provider's key, the sandbox never gets it", () => { | |
| 55 | 137 | const gatewayOnly = { AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok", CLOUDFLARE_ACCOUNT_ID: "acct" }; | |
| 56 | 138 | assert.equal(canReachModel(gatewayOnly), true); | |
| 57 | − | assert.equal(modelEnv(gatewayOnly, routes, "implement", tags).ANTHROPIC_API_KEY, "tok"); | |
| 139 | + | assert.equal(modelEnv(gatewayOnly, routes, "implement", "large", tags).ANTHROPIC_API_KEY, "tok"); | |
| 58 | 140 | assert.equal(canReachModel({ AI_GATEWAY_ID: "g1t", CLOUDFLARE_ACCOUNT_ID: "acct" }), false); | |
| 59 | 141 | assert.equal(canReachModel(direct), true); | |
| 60 | 142 | }); | |
| 79 | 161 | await new Promise<void>((resolve) => server.listen(0, resolve)); | |
| 80 | 162 | const { port } = server.address() as { port: number }; | |
| 81 | 163 | ||
| 82 | − | const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", tags); | |
| 164 | + | const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", "large", tags); | |
| 83 | 165 | // Same path as the real gateway, on the stand-in's address. | |
| 84 | 166 | const base = vars.ANTHROPIC_BASE_URL.replace("https://gateway.ai.cloudflare.com", `http://localhost:${port}`); | |
| 85 | 167 | await fetch(`${base}/v1/messages`, { | |
| 93 | 175 | url: "/v1/acct/g1t/anthropic/v1/messages", | |
| 94 | 176 | key: "sk-test", | |
| 95 | 177 | gateway: "Bearer tok", | |
| 96 | − | metadata: '{"task":"implement","repo":"acme/site","pull":12}', | |
| 178 | + | metadata: '{"task":"implement","tier":"large","repo":"acme/site","pull":12}', | |
| 97 | 179 | model: "claude-sonnet-5-5", | |
| 98 | 180 | }); | |
| 99 | 181 | }); |
| 1 | 1 | /** The kinds of work a g1t agent does. Each is routed on its own. */ | |
| 2 | 2 | export type AgentTask = "implement" | "review" | "update" | "plan"; | |
| 3 | 3 | ||
| 4 | + | /** | |
| 5 | + | * How capable, and how costly, a model is. g1t's hosted models come in | |
| 6 | + | * two: `small` for work a smaller model does as well, `large` for the rest. | |
| 7 | + | */ | |
| 8 | + | export type Tier = "small" | "large"; | |
| 9 | + | ||
| 4 | 10 | /** Where one kind of work goes: what people see, and what is sent. */ | |
| 5 | 11 | export type ModelRoute = { | |
| 6 | 12 | /** The model's public name, e.g. `Claude Sonnet 5.5`. */ | |
| 10 | 16 | }; | |
| 11 | 17 | ||
| 12 | 18 | /** | |
| 13 | − | * g1t's routing policy. Nobody assigning an agent picks a model; the kind | |
| 14 | − | * of work decides, here, and the operator changes it in one place. | |
| 19 | + | * g1t's routing policy for the runs it pays the model for. Nobody | |
| 20 | + | * assigning an agent picks a model; the work decides, here, and the | |
| 21 | + | * operator changes it in one place (`AGENT_ROUTING` in wrangler.jsonc). | |
| 22 | + | */ | |
| 23 | + | export type AgentRouting = { | |
| 24 | + | /** The model behind each tier. */ | |
| 25 | + | tiers: Record<Tier, ModelRoute>; | |
| 26 | + | /** | |
| 27 | + | * The tier each kind of work runs on. `change` decides by the change the | |
| 28 | + | * work reads: small when it is small and touches nothing sensitive. | |
| 29 | + | */ | |
| 30 | + | tasks: Record<AgentTask, Tier | "change">; | |
| 31 | + | /** The largest change `change` sends to the small tier. */ | |
| 32 | + | smallChange: { files: number; lines: number }; | |
| 33 | + | /** Labels on the issue behind the work that send `change` to the large tier. */ | |
| 34 | + | largeLabels: string[]; | |
| 35 | + | }; | |
| 36 | + | ||
| 37 | + | /** What a change is, as far as routing cares. */ | |
| 38 | + | export type ChangeSize = { | |
| 39 | + | files: number; | |
| 40 | + | /** Lines added and removed. */ | |
| 41 | + | lines: number; | |
| 42 | + | /** | |
| 43 | + | * What it touches that runs, configures or guards things: CI, secrets, | |
| 44 | + | * infrastructure, ownership (work's confidence.rs `sensitive`). | |
| 45 | + | */ | |
| 46 | + | sensitive: string[]; | |
| 47 | + | }; | |
| 48 | + | ||
| 49 | + | /** What g1t knows about one piece of work when it routes it. */ | |
| 50 | + | export type RouteSignals = { | |
| 51 | + | /** The change the work reads; null or absent when g1t does not know it. */ | |
| 52 | + | change?: ChangeSize | null; | |
| 53 | + | /** Labels on the issue the work is for. */ | |
| 54 | + | labels?: string[]; | |
| 55 | + | /** The last attempt at the same work failed. */ | |
| 56 | + | retry?: boolean; | |
| 57 | + | }; | |
| 58 | + | ||
| 59 | + | /** The routing g1t ships with, for whatever the configuration leaves out. */ | |
| 60 | + | export const DEFAULT_ROUTING: AgentRouting = { | |
| 61 | + | tiers: { | |
| 62 | + | small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" }, | |
| 63 | + | large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" }, | |
| 64 | + | }, | |
| 65 | + | tasks: { implement: "large", review: "change", update: "small", plan: "small" }, | |
| 66 | + | smallChange: { files: 10, lines: 200 }, | |
| 67 | + | largeLabels: ["security"], | |
| 68 | + | }; | |
| 69 | + | ||
| 70 | + | /** | |
| 71 | + | * The routing in `AGENT_ROUTING`, with anything it leaves out taken from | |
| 72 | + | * `DEFAULT_ROUTING`. An unset or unreadable value is the default. | |
| 73 | + | */ | |
| 74 | + | export function parseRouting(json: string | undefined): AgentRouting { | |
| 75 | + | let given: Partial<AgentRouting> = {}; | |
| 76 | + | try { | |
| 77 | + | given = json ? (JSON.parse(json) as Partial<AgentRouting>) : {}; | |
| 78 | + | } catch { | |
| 79 | + | console.log("AGENT_ROUTING is not JSON; using the default routing"); | |
| 80 | + | } | |
| 81 | + | return { | |
| 82 | + | tiers: { ...DEFAULT_ROUTING.tiers, ...given.tiers }, | |
| 83 | + | tasks: { ...DEFAULT_ROUTING.tasks, ...given.tasks }, | |
| 84 | + | smallChange: { ...DEFAULT_ROUTING.smallChange, ...given.smallChange }, | |
| 85 | + | largeLabels: given.largeLabels ?? DEFAULT_ROUTING.largeLabels, | |
| 86 | + | }; | |
| 87 | + | } | |
| 88 | + | ||
| 89 | + | /** | |
| 90 | + | * The tier one piece of work runs on: the cheapest that can do it. | |
| 91 | + | * Planning and catching up are small; making a change is large; a review | |
| 92 | + | * is small for a small change that touches nothing sensitive, and large | |
| 93 | + | * for anything else, including a change g1t does not know the size of. | |
| 94 | + | * A retry after a failed attempt is always large, so that work the small | |
| 95 | + | * tier could not finish goes up rather than failing again the same way. | |
| 15 | 96 | */ | |
| 16 | − | export type AgentRoutes = Record<AgentTask, ModelRoute>; | |
| 97 | + | export function chooseTier(task: AgentTask, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Tier { | |
| 98 | + | if (signals.retry) return "large"; | |
| 99 | + | const rule = routing.tasks[task] ?? "large"; | |
| 100 | + | if (rule !== "change") return rule; | |
| 101 | + | const change = signals.change; | |
| 102 | + | if (!change || change.files === 0) return "large"; | |
| 103 | + | const large = new Set(routing.largeLabels.map((label) => label.toLowerCase())); | |
| 104 | + | if ((signals.labels ?? []).some((label) => large.has(label.toLowerCase()))) return "large"; | |
| 105 | + | const small = | |
| 106 | + | change.sensitive.length === 0 && | |
| 107 | + | change.files <= routing.smallChange.files && | |
| 108 | + | change.lines <= routing.smallChange.lines; | |
| 109 | + | return small ? "small" : "large"; | |
| 110 | + | } | |
| 17 | 111 | ||
| 18 | 112 | /** The settings that decide where model requests go. */ | |
| 19 | 113 | export type ModelRouting = { | |
| 37 | 131 | return Boolean(env.ANTHROPIC_API_KEY || (env.AI_GATEWAY_ID && env.AI_GATEWAY_TOKEN)); | |
| 38 | 132 | } | |
| 39 | 133 | ||
| 134 | + | /** | |
| 135 | + | * The model variables of a run on g1t's hosted models: the tier's model | |
| 136 | + | * for the work, and the small tier's for the harness's own small tasks. | |
| 137 | + | */ | |
| 138 | + | export function tierVars(routing: AgentRouting, tier: Tier): Record<string, string> { | |
| 139 | + | const route = routing.tiers[tier]; | |
| 140 | + | return { | |
| 141 | + | ANTHROPIC_MODEL: route.model, | |
| 142 | + | // Recorded at the top of the session, so anyone can see what ran. | |
| 143 | + | AGENT_MODEL_NAME: route.modelName, | |
| 144 | + | ANTHROPIC_DEFAULT_HAIKU_MODEL: routing.tiers.small.model, | |
| 145 | + | ANTHROPIC_SMALL_FAST_MODEL: routing.tiers.small.model, | |
| 146 | + | }; | |
| 147 | + | } | |
| 148 | + | ||
| 40 | 149 | /** Where the sandbox sends model requests, and what it sends with them. */ | |
| 41 | 150 | export function modelEnv( | |
| 42 | 151 | env: ModelRouting, | |
| 43 | − | routes: AgentRoutes, | |
| 152 | + | routing: AgentRouting, | |
| 44 | 153 | task: AgentTask, | |
| 154 | + | tier: Tier, | |
| 45 | 155 | tags: RunTags, | |
| 46 | 156 | ): Record<string, string> { | |
| 47 | − | const route = routes[task]; | |
| 48 | − | const vars: Record<string, string> = { | |
| 49 | − | ANTHROPIC_MODEL: route.model, | |
| 50 | − | // Recorded at the top of the session, so anyone can see what ran. | |
| 51 | − | AGENT_MODEL_NAME: route.modelName, | |
| 52 | − | }; | |
| 157 | + | const vars = tierVars(routing, tier); | |
| 53 | 158 | if (env.ANTHROPIC_API_KEY) vars.ANTHROPIC_API_KEY = env.ANTHROPIC_API_KEY; | |
| 54 | 159 | if (!env.AI_GATEWAY_ID) return vars; | |
| 55 | 160 | ||
| 56 | 161 | vars.ANTHROPIC_BASE_URL = `https://gateway.ai.cloudflare.com/v1/${env.CLOUDFLARE_ACCOUNT_ID}/${env.AI_GATEWAY_ID}/anthropic`; | |
| 57 | 162 | // The gateway logs these with every request, so spend and failures can | |
| 58 | − | // be read per kind of work, repository and pull request. | |
| 59 | − | const headers = [`cf-aig-metadata: ${JSON.stringify({ task, ...tags })}`]; | |
| 163 | + | // be read per kind of work, tier, repository and pull request. | |
| 164 | + | const headers = [`cf-aig-metadata: ${JSON.stringify({ task, tier, ...tags })}`]; | |
| 60 | 165 | if (env.AI_GATEWAY_TOKEN) { | |
| 61 | 166 | vars.AI_GATEWAY_TOKEN = env.AI_GATEWAY_TOKEN; | |
| 62 | 167 | headers.push(`cf-aig-authorization: Bearer ${env.AI_GATEWAY_TOKEN}`); | |
| 67 | 172 | vars.ANTHROPIC_CUSTOM_HEADERS = headers.join("\n"); | |
| 68 | 173 | return vars; | |
| 69 | 174 | } | |
| 175 | + | ||
| 176 | + | /** Lines added and removed across a change's files. */ | |
| 177 | + | export function changeSize(files: { additions: number; deletions: number }[], sensitive: string[]): ChangeSize { | |
| 178 | + | return { | |
| 179 | + | files: files.length, | |
| 180 | + | lines: files.reduce((sum, file) => sum + file.additions + file.deletions, 0), | |
| 181 | + | sensitive, | |
| 182 | + | }; | |
| 183 | + | } | |
| 184 | + | ||
| 185 | + | /** A past run of the same work, as the work service lists it. */ | |
| 186 | + | export type PastAttempt = { status: string; halted?: string | null; title?: string | null }; | |
| 187 | + | ||
| 188 | + | /** | |
| 189 | + | * Whether the latest attempt at the same work failed: it failed, or g1t | |
| 190 | + | * stopped it at a cap of its guardrails. A person stopping it is not a | |
| 191 | + | * failure. `title` narrows it to the same plan, whose runs have no pull | |
| 192 | + | * request to tell them apart. | |
| 193 | + | */ | |
| 194 | + | export function lastAttemptFailed(newestFirst: PastAttempt[], title?: string): boolean { | |
| 195 | + | const last = newestFirst[0]; | |
| 196 | + | if (!last) return false; | |
| 197 | + | if (title !== undefined && (last.title ?? "").trim() !== title.trim()) return false; | |
| 198 | + | return last.status === "failed" || (last.status === "stopped" && Boolean(last.halted)); | |
| 199 | + | } |
| 75 | 75 | // credit. While billing takes no real money, hosted models are open | |
| 76 | 76 | // only to these workspaces; once it does, to every workspace. | |
| 77 | 77 | "HOSTED_AGENT_WORKSPACES": "flagon-io", | |
| 78 | − | // Which model each kind of work runs on. Nobody assigning an agent | |
| 79 | − | // chooses; this is g1t's policy. "modelName" is shown to people in the | |
| 80 | − | // session; "model" is sent to the provider. | |
| 81 | − | "AGENT_ROUTES": "{\"implement\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"},\"review\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"},\"update\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"},\"plan\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"}}", | |
| 78 | + | // How g1t routes the work it pays the model for. Nobody assigning an | |
| 79 | + | // agent chooses; this is g1t's policy (AgentRouting in src/model-env.ts). | |
| 80 | + | // "tiers": the model behind each tier; "modelName" is shown to people in | |
| 81 | + | // the session, "model" is sent to the provider. "tasks": the tier of each | |
| 82 | + | // kind of work; "change" reviews a change of at most "smallChange" that | |
| 83 | + | // touches nothing sensitive, for an issue without one of "largeLabels", | |
| 84 | + | // on the small tier, and anything else on the large. A retry after a | |
| 85 | + | // failed attempt always runs on the large tier. | |
| 86 | + | "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\"},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"}},\"tasks\":{\"implement\":\"large\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"small\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeLabels\":[\"security\"]}", | |
| 82 | 87 | // Where sandboxes send model requests, with a token for their run. | |
| 83 | 88 | // The proxy holds the keys: g1t's gateway's, or the workspace's own. | |
| 84 | 89 | "MODELS_URL": "https://models.g1t.sh", |
| 248 | 248 | namespace: repo.namespace, | |
| 249 | 249 | name: repo.name, | |
| 250 | 250 | }; | |
| 251 | + | let mut sensitive: Vec<String> = Vec::new(); | |
| 252 | + | for kind in pull.files.iter().filter_map(|file| crate::confidence::sensitive(&file.path)) { | |
| 253 | + | if !sensitive.iter().any(|seen| seen == kind) { | |
| 254 | + | sensitive.push(kind.to_owned()); | |
| 255 | + | } | |
| 256 | + | } | |
| 251 | 257 | Ok(Outcome::Ok(ReviewJob { | |
| 252 | 258 | run_id, | |
| 253 | 259 | token, | |
| 260 | 266 | description: pull.body.unwrap_or_default(), | |
| 261 | 267 | issue, | |
| 262 | 268 | author: pull.author, | |
| 269 | + | files: pull.files, | |
| 270 | + | sensitive, | |
| 263 | 271 | })) | |
| 264 | 272 | } | |
| 265 | 273 |