Commit

Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier

syntaqxcommitted Parentff663e5Browse files
22 files+470−700/22 viewed
+32−11
434434 ## Which model runs
435435
436436 You do not pick one. You assign the work to `g1t-agent`, the way you would
437−assign an issue to a colleague, and g1t routes it. The kind of work decides:
437+assign an issue to a colleague, and g1t routes it. On g1t's hosted models,
438+each piece of work goes to the least costly of two tiers that can do it:
438439
439−| Work | Model today |
440+| Tier | Model today |
441+| --- | --- |
442+| Small | Claude Haiku 4.5 |
443+| Large | Claude Sonnet 5.5 |
444+
445+The work decides the tier:
446+
447+| Work | Tier |
440448 | --- | --- |
441−| Making a change for an issue, and revising it | Claude Sonnet 5.5 |
442−| Reviewing a pull request | Claude Sonnet 5.5 |
443−| Catching up with `main` and resolving conflicts | Claude Sonnet 5.5 |
444−| Planning an outcome | Claude Sonnet 5.5 |
449+| Making a change for an issue, revising it, and answering a mention | Large |
450+| Reviewing a pull request that changes at most 10 files and 200 lines, touches no sensitive path, and is not for an issue labelled `security` | Small |
451+| Reviewing any other pull request, or one whose changed files g1t does not know yet | Large |
452+| Catching up with `main` and resolving conflicts | Small |
453+| Planning an outcome | Small |
454+| Any of these again, after the last attempt at the same work failed or stopped at a guardrail cap | Large |
455+
456+Sensitive paths are the ones that run, configure or guard things: CI
457+workflows, `.g1t/` and `.github/`, `CODEOWNERS`, secrets such as `.env`
458+and `.pem` files, and infrastructure such as Dockerfiles, Terraform and
459+`wrangler.*` files. They are the same paths that lower a change's
460+[confidence](#how-sure-the-agent-is).
461+
462+The agent's own small background steps run on the small tier.
445463
446464 Every session opens with a note naming the model that ran, and an agent's
447465 review says which model wrote it, so what you got is always on the record.
448−When a better model for a kind of work appears, g1t changes the route and
449−nothing you have set up needs to change.
466+When a better model for a tier appears, g1t changes the route and nothing
467+you have set up needs to change.
450468
469+A workspace that routes its work to [its own provider](/guides/models/)
470+is not routed by tier: its work runs on the model its route names.
471+
451472 A pull request made by a g1t agent carries the label `g1t-agent`, and its
452473 commits are authored by `g1t agent`.
453474
459480 Requests for g1t's hosted models go on through
460481 [Cloudflare AI Gateway](https://developers.cloudflare.com/ai-gateway/),
461482 which holds g1t's key. Each of those requests is tagged with the kind of
462−work, the repository and the pull request, so spend can be read per pull
463−request. Requests for a workspace's own provider go to that provider.
483+work, the tier, the repository and the pull request, so spend can be read
484+per tier and per pull request. Requests for a workspace's own provider go to that provider.
464485
465486 If you run your own copy of g1t, these settings control it:
466487
467488 | Setting | Where | What it does |
468489 | --- | --- | --- |
469−| `AGENT_ROUTES` | Runner | The model for each kind of work: `implement`, `review`, `update` and `plan`. |
490+| `AGENT_ROUTING` | Runner | JSON. `tiers`: the model behind `small` and `large`, each `{ "modelName", "model" }`. `tasks`: the tier of `implement`, `review`, `update` and `plan`, or `change` to decide by the change. `smallChange`: the most `files` and `lines` a `change` review runs on the small tier with. `largeLabels`: issue labels that keep a review on the large tier. Anything left out takes the defaults above. |
470491 | `MODELS_URL` | Runner | Where sandboxes send model requests: the model proxy. |
471492 | `AI_GATEWAY_ID` | Model proxy | The gateway hosted requests go through. Empty sends them to the provider directly. |
472493 | `AI_GATEWAY_TOKEN` | Model proxy | Secret. Authenticates to the gateway. |
+3−1
7676
7777 Each can go to g1t's models, or to any of your providers on any of its
7878 models. An Anthropic provider also offers **g1t's choice of Claude**,
79−which runs g1t's pick for that kind of work on your key. For example: make
79+which runs g1t's large-tier model, Claude Sonnet 5.5 today, on your key.
80+Routing between tiers by the size of the work is only for g1t's hosted
81+models; see [which model runs](/guides/g1t-agents/#which-model-runs). For example: make
8082 changes on Claude through your Anthropic key, review on GPT through your
8183 OpenAI key, and catch up on a small model through OpenRouter.
8284
+4−0
225225 revisions (`max_revisions`, 2 by default) only a required check still
226226 failing holds it for a person. There is no model or
227227 agent count to choose: to put more agents to work, assign more issues.
228+ On g1t's hosted models, g1t routes each piece of work to a small or a
229+ large tier: planning, catching up and reviews of small changes that
230+ touch no sensitive path run small; making changes, other reviews, and
231+ any retry after a failed attempt run large.
228232 - **Put an agent on something in one step:** `POST {repo}/issues/delegate`
229233 with `title` and `body` (what to do, in plain words, and what done
230234 means if you know it). It opens the issue and assigns g1t-agent at once; it needs the
+4−0
270270 /// Gateway: settling charges the run what the gateway priced them at.
271271 #[serde(default)]
272272 pub session: Option<String>,
273+ /// `small` or `large`: the tier g1t routed the run to, when g1t pays
274+ /// for its model. None on the workspace's own provider.
275+ #[serde(default)]
276+ pub tier: Option<String>,
273277 }
274278
275279 #[derive(Clone, Debug, Serialize, Deserialize)]
+8−0
384384 /// The session's id; see `ModelSession::id`.
385385 #[serde(default)]
386386 pub session: String,
387+ /// For `g1t`: the tier the run was routed to, `small` or `large`.
388+ #[serde(default)]
389+ pub tier: Option<String>,
387390 /// For `endpoint`: where to send requests.
388391 pub base_url: Option<String>,
389392 /// For `anthropic` and `endpoint`: the workspace's key.
558561 /// decides that; this service only follows the routes.
559562 #[serde(default = "yes")]
560563 pub hosted_open: bool,
564+ /// `small` or `large`: the tier the runner routed the run to on g1t's
565+ /// hosted models, tagged on its requests at the gateway. Kept only
566+ /// when the run goes to g1t's models.
567+ #[serde(default)]
568+ pub tier: Option<String>,
561569 }
562570
563571 /// `routes`: a workspace's model routes, one per kind of work that has its
+9−0
16211621 pub issue: Option<Issue>,
16221622 /// Who opened the pull request, and so can read its source.
16231623 pub author: User,
1624+ /// The files it changes, as of its latest push: how large the change
1625+ /// is, which decides the model that reviews it.
1626+ #[serde(default)]
1627+ pub files: Vec<ChangedFile>,
1628+ /// What among them runs, configures or guards things (CI workflows,
1629+ /// secrets, infrastructure), once each. Any sends the review to the
1630+ /// larger model.
1631+ #[serde(default)]
1632+ pub sensitive: Vec<String>,
16241633 }
16251634
16261635 /// A comment on one line, as a reviewing agent reports it.
+1−1
7373 | **Cache API** | `services/repos/src/store.rs` (trees and blobs up to 1 MiB, by hash), `apps/web/workers/app.ts` (avatars), `services/og` | thin, optional | Miniflare's cache, or none. Every use tolerates a miss. |
7474 | **Vectorize** | `services/context/src/index.ts` (`VECTORS.upsert/deleteByIds/query`); optional, guarded by `if (!AI \|\| !VECTORS)` | thin | **sqlite-vec** (default: one file, next to D1), pgvector or Qdrant behind a `VectorIndex` port; or off, which already degrades to keyword search. |
7575 | **Workers AI** | `services/context` only: `@cf/baai/bge-base-en-v1.5` embeddings (768 dims) | thin | An OpenAI-compatible `/v1/embeddings` endpoint (Ollama, vLLM, LM Studio, or a hosted API) behind an `Embedder` port. Changing models means re-embedding (a backfill job already exists). |
76−| **AI Gateway** | `services/runner/src/model-env.ts:56`, `services/models/src/route.ts:70` (gateway URL, `cf-aig-*` headers), `services/billing/src/keeper.rs` (reads gateway logs to settle) | thin | Optional already: an empty `AI_GATEWAY_ID` goes straight to the provider. Any Anthropic- or OpenAI-compatible base URL works for a workspace's own provider. |
76+| **AI Gateway** | `services/runner/src/model-env.ts` (`modelEnv`), `services/models/src/route.ts:70` (gateway URL, `cf-aig-*` headers), `services/billing/src/keeper.rs` (reads gateway logs to settle) | thin | Optional already: an empty `AI_GATEWAY_ID` goes straight to the provider. Any Anthropic- or OpenAI-compatible base URL works for a workspace's own provider. |
7777 | **Workers for Platforms** | `services/pages` (dispatcher: `env.APPS.get(script).fetch`), `services/deployments/src/cloudflare.ts` (script and asset upload through the REST API) | woven | Phase 1: off. Later: a self-hosted app host in workerd, using the Worker Loader binding to load uploaded scripts, or a workerd per app (see [Deployments](#deployments)). |
7878 | **Cloudflare for SaaS** (custom hostnames) | `services/deployments/src/custom-hostnames.ts` (`/zones/{id}/custom_hostnames`) | thin | Caddy with on-demand TLS, asking g1t whether a hostname is allowed. |
7979 | **Cloudflare REST API** | deployments: script upload, list, delete, assets, GraphQL usage. Billing keeper: AI Gateway logs, `billable-usage`, GraphQL container usage. Ops scripts in `scripts/`. | thin (deployments), woven (keeper pricing) | Deployments: the app-host adapter. Keeper: off when self-hosted, because there is no bill to reconcile. |
+7−4
125125 computes them (`overlaps`).
126126 - Previous reviews of the same PR, beyond comments via `peopleSaid`.
127127
128−**Model.** The production route for `review` is **Claude Sonnet 5.5**
129−(`claude-sonnet-5-5`, `AGENT_ROUTES` in `services/runner/wrangler.jsonc`).
128+**Model.** The production route for `review` was **Claude Sonnet 5.5**
129+(`claude-sonnet-5-5`) when this was measured. Reviews are now routed by
130+tier (`AGENT_ROUTING` in `services/runner/wrangler.jsonc`): small changes
131+that touch no sensitive path go to the small tier (Claude Haiku 4.5), the
132+rest stay on Sonnet 5.5.
130133 A workspace can route reviews to its own provider through the model proxy
131134 (`openModelSession`). No effort level is set, so it uses the Claude Code
132−default. The model-env test uses Opus 5.5 for review as a fixture only.
135+default.
133136
134137 **Outputs and post-processing** (`reviews.rs`, `report_review`):
135138
318321 high-severity comment counts for more than three nits.
319322 6. **Model and effort.** Run the same sample on Sonnet 5.5 at
320323 `medium`/`high` effort and on Opus 5.5. The review route can change in
321− `AGENT_ROUTES` without a code change. Opus 5.5 is about 1.4x the cost
324+ `AGENT_ROUTING` without a code change. Opus 5.5 is about 1.4x the cost
322325 per review at our token profile.
323326 7. **Product-side follow-ups** (not benchmark-visible):
324327 - allow file-level comments on unchanged files when the change breaks
+2−0
938938 billedTo?: "g1t" | "workspace";
939939 /** The model session's id, so the run can be settled at AI Gateway's price. */
940940 session?: string | null;
941+ /** `small` or `large`: the tier g1t routed the run to, on its hosted models. */
942+ tier?: "small" | "large" | null;
941943 }): Promise<Result<RunTicket | null>>;
942944 }
943945
+4−0
142142 task: string;
143143 /** The session's id; see `ModelSession.id`. */
144144 session: string;
145+ /** For `g1t`: the tier the run was routed to. */
146+ tier?: "small" | "large" | null;
145147 baseUrl: string | null;
146148 apiKey: string | null;
147149 authHeader: string | null;
183185 number: number;
184186 task: string;
185187 hostedOpen: boolean;
188+ /** The tier the run is routed to on g1t's hosted models, for the gateway's logs. */
189+ tier?: "small" | "large" | null;
186190 }): Promise<Result<ModelSession>>;
187191 routes(workspace: string, viewer: Viewer): Promise<Result<ModelRoute[]>>;
188192 setRoutes(actor: User, workspace: string, routes: ModelRoute[]): Promise<Result<ModelRoute[]>>;
+11−0
660660 issue: Issue | null;
661661 /** Who opened the pull request, and so can read its source. */
662662 author: User;
663+ /**
664+ * The files it changes, as of its latest push: how large the change is,
665+ * which decides the model that reviews it.
666+ */
667+ files: ChangedFile[];
668+ /**
669+ * What among them runs, configures or guards things (CI workflows,
670+ * secrets, infrastructure), once each. Any sends the review to the
671+ * larger model.
672+ */
673+ sensitive: string[];
663674 };
664675
665676 export type OpenIssueInput = {
+4−0
1+-- The tier g1t routed a run to on its hosted models, small or large, so
2+-- what the model cost can be read per tier. Null on a workspace's own
3+-- provider, and for runs from before routing by tier.
4+ALTER TABLE runs ADD COLUMN tier TEXT;
+7−2
639639 let token = hex::encode(bytes);
640640 self.db
641641 .prepare(
642− "INSERT INTO runs (id, workspace, repo, number, task, model, token_hash, created_at, billed_to, session_id)
643− VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
642+ "INSERT INTO runs (id, workspace, repo, number, task, model, token_hash, created_at, billed_to, session_id, tier)
643+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
644644 )
645645 .bind(&[
646646 run_id.as_str().into(),
653653 rfc3339(now).into(),
654654 if a.billed_to == "workspace" { "workspace" } else { "g1t" }.into(),
655655 optional(a.session.as_deref().filter(|_| a.billed_to != "workspace")),
656+ optional(
657+ a.tier
658+ .as_deref()
659+ .filter(|tier| a.billed_to != "workspace" && matches!(*tier, "small" | "large")),
660+ ),
656661 ])?
657662 .run()
658663 .await?;
+4−0
1+-- The tier a run on g1t's hosted models was routed to, small or large.
2+-- The model proxy tags each of the run's requests with it at the gateway.
3+-- Null for a workspace's own provider.
4+ALTER TABLE model_sessions ADD COLUMN tier TEXT;
+11−2
137137 number: u32,
138138 task: String,
139139 model: Option<String>,
140+ #[serde(default)]
141+ tier: Option<String>,
140142 }
141143
142144 #[derive(Deserialize)]
12271229 .bind(&[rfc3339(now).into()])?,
12281230 self.db
12291231 .prepare(
1230− "INSERT INTO model_sessions (token_hash, workspace, connection_id, repo, number, task, expires_at, model)
1231− VALUES (?, ?, ?, ?, ?, ?, ?, ?)",
1232+ "INSERT INTO model_sessions (token_hash, workspace, connection_id, repo, number, task, expires_at, model, tier)
1233+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
12321234 )
12331235 .bind(&[
12341236 crypto::sha256_hex(&token).into(),
12391241 a.task.as_str().into(),
12401242 rfc3339(now + MODEL_SESSION_SECONDS * 1000).into(),
12411243 optional(model.as_deref()),
1244+ // The tier is g1t's routing; it means nothing on the workspace's own provider.
1245+ optional(
1246+ a.tier
1247+ .as_deref()
1248+ .filter(|tier| connection.is_none() && matches!(*tier, "small" | "large")),
1249+ ),
12421250 ])?,
12431251 ])
12441252 .await?;
12731281 number: session.number,
12741282 task: session.task,
12751283 session: session_id(&session.token_hash),
1284+ tier: session.tier,
12761285 base_url: None,
12771286 api_key: None,
12781287 auth_header: None,
+16−0
3939 });
4040 });
4141
42+test("a run's tier is tagged at the gateway, so spend can be read per tier", () => {
43+ const upstream: ModelUpstream = { ...run, route: "g1t", tier: "small" };
44+ const { headers } = upstreamRequest(upstream, hosted, "/v1/messages", incoming());
45+ const metadata = JSON.parse(headers.get("cf-aig-metadata") ?? "{}");
46+ assert.deepEqual(metadata, {
47+ task: "implement",
48+ tier: "small",
49+ repo: "acme/web",
50+ pull: 7,
51+ session: "ms_abc",
52+ });
53+ // The gateway keeps five entries; the session, which billing settles
54+ // by, must be one of them.
55+ assert.ok(Object.keys(metadata).length <= 5);
56+});
57+
4258 test("a workspace's own Anthropic key goes to Anthropic, and only there", () => {
4359 const upstream: ModelUpstream = { ...run, route: "anthropic", baseUrl: "https://api.anthropic.com", apiKey: "sk-ant-theirs", authHeader: "x-api-key" };
4460 const { url, headers } = upstreamRequest(upstream, hosted, "/v1/messages", incoming());
+5−2
5252 return { url: `https://api.anthropic.com${path}`, headers };
5353 }
5454 // The gateway logs these with every request, so spend and failures can
55− // be read per kind of work, workspace, repository and pull request.
55+ // be read per kind of work, tier, repository and pull request. It keeps
56+ // five entries and drops the rest, so the tier takes the workspace's
57+ // place: the repository names the workspace too. A run from before
58+ // routing by tier has no tier, and keeps the workspace.
5659 headers.set(
5760 "cf-aig-metadata",
5861 JSON.stringify({
5962 task: upstream.task,
60− workspace: upstream.workspace,
63+ ...(upstream.tier ? { tier: upstream.tier } : { workspace: upstream.workspace }),
6164 repo: upstream.repo,
6265 pull: upstream.number,
6366 // What billing finds the run's requests by, to charge what they cost.
+79−13
6060 instanceNamed,
6161 } from "@g1t/contracts";
6262
63−import { type AgentRoutes, type AgentTask, canReachModel, modelEnv } from "./model-env";
63+import {
64+ type AgentTask,
65+ type RouteSignals,
66+ type Tier,
67+ canReachModel,
68+ changeSize,
69+ chooseTier,
70+ lastAttemptFailed,
71+ modelEnv,
72+ parseRouting,
73+ tierVars,
74+} from "./model-env";
6475 import { hubContext } from "./hub";
6576 import { hostedOpen } from "./hosted";
6677 import { delegateInput, noModelMessage, notStarted, queued, started } from "./delegate";
137148 */
138149 HOSTED_AGENT_WORKSPACES: string;
139150 /**
140− * Which model each kind of work runs on, as JSON:
141− * `{ implement, review, update }`, each `{ modelName, model }`.
142− * `modelName` is what people see; `model` is sent to the provider.
151+ * How g1t routes the work it pays the model for, as JSON (`AgentRouting`
152+ * in model-env.ts): `tiers`, the model behind `small` and `large`, each
153+ * `{ modelName, model }`; `tasks`, the tier of each kind of work, or
154+ * `change` to decide a review by its change; `smallChange`, the largest
155+ * change reviewed on the small tier; `largeLabels`, issue labels that
156+ * keep a review large. Anything left out takes the default.
143157 */
144− AGENT_ROUTES: string;
158+ AGENT_ROUTING?: string;
145159 /**
146160 * A Cloudflare AI Gateway id. When set, model traffic goes through that
147161 * gateway, which is where logging, spend limits, caching and fallback
167181 ABUSE_WATCH?: string;
168182 }
169183
184+/**
185+ * What routing knows about one piece of work, and how to ask whether it is
186+ * a retry (asked only when the answer matters).
187+ */
188+type RouteInput = RouteSignals & { retried?: () => Promise<boolean> };
189+
170190 /** A run that takes longer than this has its token expire under it. */
171191 const TOKEN_TTL_SECONDS = 2 * 60 * 60;
172192 /** How g1t's own agent is labelled. What runs behind it is g1t's choice. */
13081328 * What a sandbox needs to reach the model routed for `task`, having
13091329 * opened the run the repository's workspace will be charged for. Refused
13101330 * when that workspace has no credit.
1331+ *
1332+ * On g1t's hosted models the work goes to the cheapest tier that can do
1333+ * it (`chooseTier`), by what `route` says about it. Whether this is a
1334+ * retry is asked only when it would change the answer: when the work
1335+ * would otherwise go to the small tier.
13111336 */
13121337 private async modelEnv(
13131338 task: AgentTask,
13141339 repo: RepoPath,
13151340 pull: number,
1341+ route: RouteInput = {},
13161342 ): Promise<Result<Record<string, string>>> {
1317− const routes: AgentRoutes = JSON.parse(this.env.AGENT_ROUTES);
1343+ const routing = parseRouting(this.env.AGENT_ROUTING);
1344+ let tier: Tier = chooseTier(task, route, routing);
1345+ if (tier === "small" && route.retried && (await route.retried().catch(() => false))) {
1346+ tier = chooseTier(task, { ...route, retry: true }, routing);
1347+ }
13181348 const tags = { repo: `${repo.namespace}/${repo.name}`, pull };
13191349 // Where the run's model requests go, by the workspace's routes: g1t's
13201350 // hosted models, or one of its own providers.
13261356 number: pull,
13271357 task,
13281358 hostedOpen: (await this.modelAccess(repo.namespace)).hosted,
1359+ tier,
13291360 });
13301361 if (!opened.ok) return opened;
13311362 session = opened.value;
13321363 }
13331364 const own = session?.billedTo === "workspace";
1334− const model = session?.model ?? routes[task].model;
1335− const modelName = session?.model ?? routes[task].modelName;
1365+ // A workspace's own provider is not routed by tier: it runs the model
1366+ // its route names, or for an Anthropic provider, the large tier's.
1367+ const routed = routing.tiers[own ? "large" : tier];
1368+ const model = session?.model ?? routed.model;
1369+ const modelName = session?.model ?? routed.modelName;
13361370 const ticket = await billingClient(this.env.BILLING).startRun({
13371371 workspace: repo.namespace,
13381372 repo,
13411375 model: own ? `${modelName} (${session?.providerName ?? "own provider"})` : modelName,
13421376 billedTo: own ? "workspace" : "g1t",
13431377 session: own ? null : (session?.id ?? null),
1378+ tier: own ? null : tier,
13441379 });
13451380 if (!ticket.ok) return ticket;
13461381 const vars: Record<string, string> = session
13471382 ? {
1383+ // On g1t's models, the tier's model, and the small tier's for the
1384+ // harness's own small tasks.
1385+ ...(own ? {} : tierVars(routing, tier)),
13481386 ANTHROPIC_MODEL: model,
13491387 AGENT_MODEL_NAME: own ? `${modelName}, through ${session.providerName}` : modelName,
13501388 ANTHROPIC_BASE_URL: `${this.env.MODELS_URL!.replace(/\/+$/, "")}/anthropic`,
13541392 // the harness's small tasks too.
13551393 ...(session.model ? { ANTHROPIC_SMALL_FAST_MODEL: session.model } : {}),
13561394 }
1357− : modelEnv(this.env, routes, task, tags);
1395+ : modelEnv(this.env, routing, task, tier, tags);
13581396 if (ticket.value) {
13591397 // How the sandbox says what the run cost. Kept from the agent.
13601398 vars.BILLING_RUN = ticket.value.runId;
14901528 task: AgentTask,
14911529 repo: RepoPath,
14921530 pull: number,
1531+ route: RouteInput = {},
14931532 ): Promise<Record<string, string>> {
1494− const vars = await this.modelEnv(task, repo, pull);
1533+ const vars = await this.modelEnv(task, repo, pull, route);
14951534 if (!vars.ok) throw new Error(vars.error.message);
14961535 return vars.value;
14971536 }
14981537
1538+ /**
1539+ * Whether the latest run of the same work failed, so that this one is a
1540+ * retry: the same kind of run on the same pull request, or for a plan,
1541+ * the latest plan with the same brief. Read as the one the run is for;
1542+ * unknown counts as not.
1543+ */
1544+ private async failedBefore(
1545+ viewer: User,
1546+ repo: RepoPath,
1547+ kind: "review" | "update" | "plan",
1548+ number: number | null,
1549+ title?: string,
1550+ ): Promise<boolean> {
1551+ const runs = await agentsClient(this.env.WORK)
1552+ .listRuns(viewer, { repo, kind, ...(number != null ? { number } : {}), limit: 1 })
1553+ .catch(() => null);
1554+ return runs?.ok ? lastAttemptFailed(runs.value, title) : false;
1555+ }
1556+
14991557 /** Whether sandboxes have a way to reach a model at all. */
15001558 private modelsReachable(): boolean {
15011559 return Boolean(this.env.MODELS_URL) || canReachModel(this.env);
24322490 repo,
24332491 actor,
24342492 ),
2435− ...(await this.modelEnvOrThrow("update", repo, number)),
2493+ ...(await this.modelEnvOrThrow("update", repo, number, {
2494+ retried: () => this.failedBefore(actor, repo, "update", number),
2495+ })),
24362496 },
24372497 });
24382498 }
24862546 `It is for issue #${job.issue.number}: ${job.issue.title}\n\n${job.issue.body}`,
24872547 await this.peopleSaid(job.author, repo, number),
24882548 ];
2489− const model = await this.modelEnv("review", repo, number);
2549+ const model = await this.modelEnv("review", repo, number, {
2550+ change: job.files?.length ? changeSize(job.files, job.sensitive ?? []) : null,
2551+ labels: job.issue?.labels ?? [],
2552+ retried: () => this.failedBefore(job.author, repo, "review", number),
2553+ });
24902554 if (!model.ok) {
24912555 await workClient(this.env.WORK).failReview(job.runId, job.token, model.error.message);
24922556 return model;
25462610 const started = await work.startPlan(actor, repo, brief);
25472611 if (!started.ok) return started;
25482612 const job = started.value;
2549− const model = await this.modelEnv("plan", repo, 0);
2613+ const model = await this.modelEnv("plan", repo, 0, {
2614+ retried: () => this.failedBefore(actor, repo, "plan", null, job.brief),
2615+ });
25502616 if (!model.ok) {
25512617 await work.failPlan(job.planId, job.token, model.error.message);
25522618 return model;
+100−18
22 import { createServer } from "node:http";
33 import { test } from "node:test";
44
5−import { type AgentRoutes, canReachModel, modelEnv } from "./model-env.ts";
5+import {
6+ type AgentRouting,
7+ type ChangeSize,
8+ DEFAULT_ROUTING,
9+ canReachModel,
10+ changeSize,
11+ chooseTier,
12+ lastAttemptFailed,
13+ modelEnv,
14+ parseRouting,
15+} from "./model-env.ts";
616
7−const routes: AgentRoutes = {
8− implement: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" },
9− review: { modelName: "Claude Opus 5.5", model: "claude-opus-5-5" },
10− update: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" },
11− plan: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" },
17+const routes: AgentRouting = {
18+ ...DEFAULT_ROUTING,
19+ tiers: {
20+ small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" },
21+ large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" },
22+ },
1223 };
1324 const tags = { repo: "acme/site", pull: 12 };
1425 const direct = { ANTHROPIC_API_KEY: "sk-test", AI_GATEWAY_ID: "", CLOUDFLARE_ACCOUNT_ID: "acct" };
2334 );
2435 }
2536
26−test("the kind of work decides the model", () => {
27− assert.equal(modelEnv(direct, routes, "implement", tags).ANTHROPIC_MODEL, "claude-sonnet-5-5");
28− const review = modelEnv(direct, routes, "review", tags);
29− assert.equal(review.ANTHROPIC_MODEL, "claude-opus-5-5");
30− assert.equal(review.AGENT_MODEL_NAME, "Claude Opus 5.5");
37+const small: ChangeSize = { files: 3, lines: 80, sensitive: [] };
38+
39+test("planning and catching up run on the small tier, making a change on the large", () => {
40+ assert.equal(chooseTier("plan", {}, routes), "small");
41+ assert.equal(chooseTier("update", {}, routes), "small");
42+ assert.equal(chooseTier("implement", {}, routes), "large");
43+ assert.equal(chooseTier("implement", { change: small }, routes), "large");
44+});
45+
46+test("a review is small only for a small change that touches nothing sensitive", () => {
47+ assert.equal(chooseTier("review", { change: small }, routes), "small");
48+ assert.equal(chooseTier("review", { change: { ...small, lines: 200, files: 10 } }, routes), "small");
49+ assert.equal(chooseTier("review", { change: { ...small, lines: 201 } }, routes), "large");
50+ assert.equal(chooseTier("review", { change: { ...small, files: 11 } }, routes), "large");
51+ assert.equal(chooseTier("review", { change: { ...small, sensitive: ["CI workflows"] } }, routes), "large");
52+ assert.equal(chooseTier("review", { change: small, labels: ["Security"] }, routes), "large");
53+ assert.equal(chooseTier("review", { change: small, labels: ["docs"] }, routes), "small");
54+});
55+
56+test("a review of a change g1t cannot size runs on the large tier", () => {
57+ assert.equal(chooseTier("review", {}, routes), "large");
58+ assert.equal(chooseTier("review", { change: null }, routes), "large");
59+ assert.equal(chooseTier("review", { change: { files: 0, lines: 0, sensitive: [] } }, routes), "large");
60+});
61+
62+test("a retry after a failed attempt goes up to the large tier", () => {
63+ assert.equal(chooseTier("plan", { retry: true }, routes), "large");
64+ assert.equal(chooseTier("update", { retry: true }, routes), "large");
65+ assert.equal(chooseTier("review", { change: small, retry: true }, routes), "large");
66+});
67+
68+test("a retry is the same work again after its latest attempt failed", () => {
69+ assert.equal(lastAttemptFailed([]), false);
70+ assert.equal(lastAttemptFailed([{ status: "failed" }, { status: "succeeded" }]), true);
71+ assert.equal(lastAttemptFailed([{ status: "succeeded" }, { status: "failed" }]), false);
72+ assert.equal(lastAttemptFailed([{ status: "stopped", halted: "budget" }]), true);
73+ assert.equal(lastAttemptFailed([{ status: "stopped", halted: null }]), false);
74+ assert.equal(lastAttemptFailed([{ status: "running" }]), false);
75+ assert.equal(lastAttemptFailed([{ status: "failed", title: "Add search" }], "Add search "), true);
76+ assert.equal(lastAttemptFailed([{ status: "failed", title: "Add search" }], "Add billing"), false);
77+});
78+
79+test("the configuration decides the tiers, the rules and the limits", () => {
80+ const parsed = parseRouting(
81+ JSON.stringify({
82+ tiers: { small: { modelName: "Small", model: "small-1" } },
83+ tasks: { update: "large" },
84+ smallChange: { lines: 50 },
85+ }),
86+ );
87+ assert.deepEqual(parsed.tiers.small, { modelName: "Small", model: "small-1" });
88+ assert.deepEqual(parsed.tiers.large, DEFAULT_ROUTING.tiers.large);
89+ assert.equal(chooseTier("update", {}, parsed), "large");
90+ assert.equal(chooseTier("plan", {}, parsed), "small");
91+ assert.equal(chooseTier("review", { change: { ...small, lines: 51 } }, parsed), "large");
92+ assert.equal(chooseTier("review", { change: { ...small, files: 10, lines: 50 } }, parsed), "small");
93+ assert.deepEqual(parseRouting(undefined), DEFAULT_ROUTING);
94+ assert.deepEqual(parseRouting("not json"), DEFAULT_ROUTING);
3195 });
3296
97+test("a change's size is its files and the lines added and removed", () => {
98+ assert.deepEqual(
99+ changeSize([{ additions: 10, deletions: 2 }, { additions: 0, deletions: 5 }], ["secrets"]),
100+ { files: 2, lines: 17, sensitive: ["secrets"] },
101+ );
102+});
103+
104+test("the tier decides the model, and the harness's small tasks use the small tier", () => {
105+ const large = modelEnv(direct, routes, "implement", "large", tags);
106+ assert.equal(large.ANTHROPIC_MODEL, "claude-sonnet-5-5");
107+ assert.equal(large.AGENT_MODEL_NAME, "Claude Sonnet 5.5");
108+ assert.equal(large.ANTHROPIC_SMALL_FAST_MODEL, "claude-haiku-4-5-20251001");
109+ assert.equal(large.ANTHROPIC_DEFAULT_HAIKU_MODEL, "claude-haiku-4-5-20251001");
110+ const review = modelEnv(direct, routes, "review", "small", tags);
111+ assert.equal(review.ANTHROPIC_MODEL, "claude-haiku-4-5-20251001");
112+ assert.equal(review.AGENT_MODEL_NAME, "Claude Haiku 4.5");
113+});
114+
33115 test("without a gateway, requests go to the provider directly", () => {
34− const vars = modelEnv(direct, routes, "implement", tags);
116+ const vars = modelEnv(direct, routes, "implement", "large", tags);
35117 assert.equal(vars.ANTHROPIC_BASE_URL, undefined);
36118 assert.equal(vars.ANTHROPIC_CUSTOM_HEADERS, undefined);
37119 assert.equal(vars.ANTHROPIC_API_KEY, "sk-test");
38120 });
39121
40122 test("with a gateway, requests go through it and say what they are for", () => {
41− const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t" }, routes, "review", tags);
123+ const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t" }, routes, "review", "large", tags);
42124 assert.equal(vars.ANTHROPIC_BASE_URL, "https://gateway.ai.cloudflare.com/v1/acct/g1t/anthropic");
43125 assert.deepEqual(customHeaders(vars), {
44− "cf-aig-metadata": '{"task":"review","repo":"acme/site","pull":12}',
126+ "cf-aig-metadata": '{"task":"review","tier":"large","repo":"acme/site","pull":12}',
45127 });
46128 });
47129
48130 test("an authenticated gateway is sent its token", () => {
49− const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", tags);
131+ const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", "large", tags);
50132 assert.equal(customHeaders(vars)["cf-aig-authorization"], "Bearer tok");
51133 assert.equal(vars.ANTHROPIC_API_KEY, "sk-test");
52134 });
54136 test("when the gateway holds the provider's key, the sandbox never gets it", () => {
55137 const gatewayOnly = { AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok", CLOUDFLARE_ACCOUNT_ID: "acct" };
56138 assert.equal(canReachModel(gatewayOnly), true);
57− assert.equal(modelEnv(gatewayOnly, routes, "implement", tags).ANTHROPIC_API_KEY, "tok");
139+ assert.equal(modelEnv(gatewayOnly, routes, "implement", "large", tags).ANTHROPIC_API_KEY, "tok");
58140 assert.equal(canReachModel({ AI_GATEWAY_ID: "g1t", CLOUDFLARE_ACCOUNT_ID: "acct" }), false);
59141 assert.equal(canReachModel(direct), true);
60142 });
79161 await new Promise<void>((resolve) => server.listen(0, resolve));
80162 const { port } = server.address() as { port: number };
81163
82− const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", tags);
164+ const vars = modelEnv({ ...direct, AI_GATEWAY_ID: "g1t", AI_GATEWAY_TOKEN: "tok" }, routes, "implement", "large", tags);
83165 // Same path as the real gateway, on the stand-in's address.
84166 const base = vars.ANTHROPIC_BASE_URL.replace("https://gateway.ai.cloudflare.com", `http://localhost:${port}`);
85167 await fetch(`${base}/v1/messages`, {
93175 url: "/v1/acct/g1t/anthropic/v1/messages",
94176 key: "sk-test",
95177 gateway: "Bearer tok",
96− metadata: '{"task":"implement","repo":"acme/site","pull":12}',
178+ metadata: '{"task":"implement","tier":"large","repo":"acme/site","pull":12}',
97179 model: "claude-sonnet-5-5",
98180 });
99181 });
+142−12
11 /** The kinds of work a g1t agent does. Each is routed on its own. */
22 export type AgentTask = "implement" | "review" | "update" | "plan";
33
4+/**
5+ * How capable, and how costly, a model is. g1t's hosted models come in
6+ * two: `small` for work a smaller model does as well, `large` for the rest.
7+ */
8+export type Tier = "small" | "large";
9+
410 /** Where one kind of work goes: what people see, and what is sent. */
511 export type ModelRoute = {
612 /** The model's public name, e.g. `Claude Sonnet 5.5`. */
1016 };
1117
1218 /**
13− * g1t's routing policy. Nobody assigning an agent picks a model; the kind
14− * of work decides, here, and the operator changes it in one place.
19+ * g1t's routing policy for the runs it pays the model for. Nobody
20+ * assigning an agent picks a model; the work decides, here, and the
21+ * operator changes it in one place (`AGENT_ROUTING` in wrangler.jsonc).
22+ */
23+export type AgentRouting = {
24+ /** The model behind each tier. */
25+ tiers: Record<Tier, ModelRoute>;
26+ /**
27+ * The tier each kind of work runs on. `change` decides by the change the
28+ * work reads: small when it is small and touches nothing sensitive.
29+ */
30+ tasks: Record<AgentTask, Tier | "change">;
31+ /** The largest change `change` sends to the small tier. */
32+ smallChange: { files: number; lines: number };
33+ /** Labels on the issue behind the work that send `change` to the large tier. */
34+ largeLabels: string[];
35+};
36+
37+/** What a change is, as far as routing cares. */
38+export type ChangeSize = {
39+ files: number;
40+ /** Lines added and removed. */
41+ lines: number;
42+ /**
43+ * What it touches that runs, configures or guards things: CI, secrets,
44+ * infrastructure, ownership (work's confidence.rs `sensitive`).
45+ */
46+ sensitive: string[];
47+};
48+
49+/** What g1t knows about one piece of work when it routes it. */
50+export type RouteSignals = {
51+ /** The change the work reads; null or absent when g1t does not know it. */
52+ change?: ChangeSize | null;
53+ /** Labels on the issue the work is for. */
54+ labels?: string[];
55+ /** The last attempt at the same work failed. */
56+ retry?: boolean;
57+};
58+
59+/** The routing g1t ships with, for whatever the configuration leaves out. */
60+export const DEFAULT_ROUTING: AgentRouting = {
61+ tiers: {
62+ small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" },
63+ large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" },
64+ },
65+ tasks: { implement: "large", review: "change", update: "small", plan: "small" },
66+ smallChange: { files: 10, lines: 200 },
67+ largeLabels: ["security"],
68+};
69+
70+/**
71+ * The routing in `AGENT_ROUTING`, with anything it leaves out taken from
72+ * `DEFAULT_ROUTING`. An unset or unreadable value is the default.
73+ */
74+export function parseRouting(json: string | undefined): AgentRouting {
75+ let given: Partial<AgentRouting> = {};
76+ try {
77+ given = json ? (JSON.parse(json) as Partial<AgentRouting>) : {};
78+ } catch {
79+ console.log("AGENT_ROUTING is not JSON; using the default routing");
80+ }
81+ return {
82+ tiers: { ...DEFAULT_ROUTING.tiers, ...given.tiers },
83+ tasks: { ...DEFAULT_ROUTING.tasks, ...given.tasks },
84+ smallChange: { ...DEFAULT_ROUTING.smallChange, ...given.smallChange },
85+ largeLabels: given.largeLabels ?? DEFAULT_ROUTING.largeLabels,
86+ };
87+}
88+
89+/**
90+ * The tier one piece of work runs on: the cheapest that can do it.
91+ * Planning and catching up are small; making a change is large; a review
92+ * is small for a small change that touches nothing sensitive, and large
93+ * for anything else, including a change g1t does not know the size of.
94+ * A retry after a failed attempt is always large, so that work the small
95+ * tier could not finish goes up rather than failing again the same way.
1596 */
16−export type AgentRoutes = Record<AgentTask, ModelRoute>;
97+export function chooseTier(task: AgentTask, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Tier {
98+ if (signals.retry) return "large";
99+ const rule = routing.tasks[task] ?? "large";
100+ if (rule !== "change") return rule;
101+ const change = signals.change;
102+ if (!change || change.files === 0) return "large";
103+ const large = new Set(routing.largeLabels.map((label) => label.toLowerCase()));
104+ if ((signals.labels ?? []).some((label) => large.has(label.toLowerCase()))) return "large";
105+ const small =
106+ change.sensitive.length === 0 &&
107+ change.files <= routing.smallChange.files &&
108+ change.lines <= routing.smallChange.lines;
109+ return small ? "small" : "large";
110+}
17111
18112 /** The settings that decide where model requests go. */
19113 export type ModelRouting = {
37131 return Boolean(env.ANTHROPIC_API_KEY || (env.AI_GATEWAY_ID && env.AI_GATEWAY_TOKEN));
38132 }
39133
134+/**
135+ * The model variables of a run on g1t's hosted models: the tier's model
136+ * for the work, and the small tier's for the harness's own small tasks.
137+ */
138+export function tierVars(routing: AgentRouting, tier: Tier): Record<string, string> {
139+ const route = routing.tiers[tier];
140+ return {
141+ ANTHROPIC_MODEL: route.model,
142+ // Recorded at the top of the session, so anyone can see what ran.
143+ AGENT_MODEL_NAME: route.modelName,
144+ ANTHROPIC_DEFAULT_HAIKU_MODEL: routing.tiers.small.model,
145+ ANTHROPIC_SMALL_FAST_MODEL: routing.tiers.small.model,
146+ };
147+}
148+
40149 /** Where the sandbox sends model requests, and what it sends with them. */
41150 export function modelEnv(
42151 env: ModelRouting,
43− routes: AgentRoutes,
152+ routing: AgentRouting,
44153 task: AgentTask,
154+ tier: Tier,
45155 tags: RunTags,
46156 ): Record<string, string> {
47− const route = routes[task];
48− const vars: Record<string, string> = {
49− ANTHROPIC_MODEL: route.model,
50− // Recorded at the top of the session, so anyone can see what ran.
51− AGENT_MODEL_NAME: route.modelName,
52− };
157+ const vars = tierVars(routing, tier);
53158 if (env.ANTHROPIC_API_KEY) vars.ANTHROPIC_API_KEY = env.ANTHROPIC_API_KEY;
54159 if (!env.AI_GATEWAY_ID) return vars;
55160
56161 vars.ANTHROPIC_BASE_URL = `https://gateway.ai.cloudflare.com/v1/${env.CLOUDFLARE_ACCOUNT_ID}/${env.AI_GATEWAY_ID}/anthropic`;
57162 // The gateway logs these with every request, so spend and failures can
58− // be read per kind of work, repository and pull request.
59− const headers = [`cf-aig-metadata: ${JSON.stringify({ task, ...tags })}`];
163+ // be read per kind of work, tier, repository and pull request.
164+ const headers = [`cf-aig-metadata: ${JSON.stringify({ task, tier, ...tags })}`];
60165 if (env.AI_GATEWAY_TOKEN) {
61166 vars.AI_GATEWAY_TOKEN = env.AI_GATEWAY_TOKEN;
62167 headers.push(`cf-aig-authorization: Bearer ${env.AI_GATEWAY_TOKEN}`);
67172 vars.ANTHROPIC_CUSTOM_HEADERS = headers.join("\n");
68173 return vars;
69174 }
175+
176+/** Lines added and removed across a change's files. */
177+export function changeSize(files: { additions: number; deletions: number }[], sensitive: string[]): ChangeSize {
178+ return {
179+ files: files.length,
180+ lines: files.reduce((sum, file) => sum + file.additions + file.deletions, 0),
181+ sensitive,
182+ };
183+}
184+
185+/** A past run of the same work, as the work service lists it. */
186+export type PastAttempt = { status: string; halted?: string | null; title?: string | null };
187+
188+/**
189+ * Whether the latest attempt at the same work failed: it failed, or g1t
190+ * stopped it at a cap of its guardrails. A person stopping it is not a
191+ * failure. `title` narrows it to the same plan, whose runs have no pull
192+ * request to tell them apart.
193+ */
194+export function lastAttemptFailed(newestFirst: PastAttempt[], title?: string): boolean {
195+ const last = newestFirst[0];
196+ if (!last) return false;
197+ if (title !== undefined && (last.title ?? "").trim() !== title.trim()) return false;
198+ return last.status === "failed" || (last.status === "stopped" && Boolean(last.halted));
199+}
+9−4
7575 // credit. While billing takes no real money, hosted models are open
7676 // only to these workspaces; once it does, to every workspace.
7777 "HOSTED_AGENT_WORKSPACES": "flagon-io",
78− // Which model each kind of work runs on. Nobody assigning an agent
79− // chooses; this is g1t's policy. "modelName" is shown to people in the
80− // session; "model" is sent to the provider.
81− "AGENT_ROUTES": "{\"implement\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"},\"review\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"},\"update\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"},\"plan\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"}}",
78+ // How g1t routes the work it pays the model for. Nobody assigning an
79+ // agent chooses; this is g1t's policy (AgentRouting in src/model-env.ts).
80+ // "tiers": the model behind each tier; "modelName" is shown to people in
81+ // the session, "model" is sent to the provider. "tasks": the tier of each
82+ // kind of work; "change" reviews a change of at most "smallChange" that
83+ // touches nothing sensitive, for an issue without one of "largeLabels",
84+ // on the small tier, and anything else on the large. A retry after a
85+ // failed attempt always runs on the large tier.
86+ "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\"},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"}},\"tasks\":{\"implement\":\"large\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"small\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeLabels\":[\"security\"]}",
8287 // Where sandboxes send model requests, with a token for their run.
8388 // The proxy holds the keys: g1t's gateway's, or the workspace's own.
8489 "MODELS_URL": "https://models.g1t.sh",
+8−0
248248 namespace: repo.namespace,
249249 name: repo.name,
250250 };
251+ let mut sensitive: Vec<String> = Vec::new();
252+ for kind in pull.files.iter().filter_map(|file| crate::confidence::sensitive(&file.path)) {
253+ if !sensitive.iter().any(|seen| seen == kind) {
254+ sensitive.push(kind.to_owned());
255+ }
256+ }
251257 Ok(Outcome::Ok(ReviewJob {
252258 run_id,
253259 token,
260266 description: pull.body.unwrap_or_default(),
261267 issue,
262268 author: pull.author,
269+ files: pull.files,
270+ sensitive,
263271 }))
264272 }
265273