Skip to content

Commit

Runner: Auto routes each agent job to the cheapest model that can do it

The router (route in model-env.ts) now has three tiers, small (Claude Haiku 4.5), large (Claude Sonnet 5.5) and frontier (Claude Opus 5.5), and routes every kind of job on its own: catching up and answering questions start fast, changes and revisions standard, plans most capable, and a review is sized by its change (small and safe: fast; over 60 files or 3,000 lines: most capable). Issue labels move work (architecture up, security off the fast tier, docs and typos down). A failed attempt goes one tier up, two in a row to the most capable, and a change left at low confidence sends the next attempt up. A repository's own recent runs of the same kind step work down when the cheaper tier finished 9 in 10, and up when this tier failed half. The catalogue (models, list prices for estimates) and every rule are configuration (AGENT_ROUTING), tested to equal the defaults. A workspace can choose a tier per kind of work on g1t's models instead of Auto (model_routes, integrations), and a workspace's own Anthropic key with no model named is routed the same way. Each run says which model and why in one line, as its first step and at the top of its session (AGENT_MODEL_REASON). Own-provider runs keep their model session for the agent rate. The sandbox image needs rebuilding for the session line and the token report.

syntaqxcommitted Parent7c44c2cBrowse files
9 files+651−1460/9 viewed
+19−6
170170 /// The kinds of work a model is chosen for, and `default` for the rest.
171171 pub const MODEL_TASKS: [&str; 5] = ["default", "implement", "review", "plan", "update"];
172172
173+/// The tiers g1t routes its hosted models' work to, cheapest first: `small`
174+/// (fast), `large` (standard) and `frontier` (most capable). A route to
175+/// g1t's models may name one instead of leaving the choice to Auto.
176+pub const MODEL_TIERS: [&str; 3] = ["small", "large", "frontier"];
177+
173178 #[derive(Clone, Copy, Debug, PartialEq, Eq, Serialize, Deserialize)]
174179 #[serde(rename_all = "snake_case")]
175180 pub enum ProviderKind {
286291 /// The workspace's own model connection, or `None` for g1t's hosted
287292 /// models.
288293 pub connection_id: Option<String>,
289− /// The model at that connection; its default model when `None`.
294+ /// The model at that connection; its default model when `None`. On
295+ /// g1t's hosted models, one of [`MODEL_TIERS`], or `None` for Auto.
290296 pub model: Option<String>,
291297 }
292298
354360 /// The model to use instead of g1t's choice, if the connection names one.
355361 pub model: Option<String>,
356362 /// Names the run in AI Gateway's logs (`metadata.session`), so billing
357− /// can charge each run what the gateway priced its requests at. Not a
358− /// secret: it cannot be turned back into the token.
363+ /// can charge each run what the gateway priced its requests at, and its
364+ /// tokens in billing's count (the agent rate). Not a secret: it cannot
365+ /// be turned back into the token.
359366 #[serde(default)]
360367 pub id: String,
368+ /// The tier the workspace chose for this work on g1t's models instead
369+ /// of Auto: the run goes there. `None` for Auto and on its own
370+ /// providers.
371+ #[serde(default)]
372+ pub tier_choice: Option<String>,
361373 }
362374
363375 /// What the model proxy needs to forward one run's requests.
565577 /// decides that; this service only follows the routes.
566578 #[serde(default = "yes")]
567579 pub hosted_open: bool,
568− /// `small` or `large`: the tier the runner routed the run to on g1t's
569− /// hosted models, tagged on its requests at the gateway. Kept only
570− /// when the run goes to g1t's models.
580+ /// `small`, `large` or `frontier`: the tier the runner routed the run
581+ /// to on g1t's hosted models, tagged on its requests at the gateway.
582+ /// Kept only when the run goes to g1t's models, where a tier the
583+ /// workspace chose for the work takes its place.
571584 #[serde(default)]
572585 pub tier: Option<String>,
573586 /// The person the run is for, by username, so usage can be shown per
+5−1
107107 let auth = auth_option(&env("G1T_USER")?, &env("G1T_TOKEN")?);
108108 let workdir = Path::new(WORKDIR);
109109
110− if let Ok(model) = std::env::var("AGENT_MODEL_NAME") {
110+ // Which model, and why: the router's one line names both.
111+ let reason = std::env::var("AGENT_MODEL_REASON").unwrap_or_default();
112+ if !reason.trim().is_empty() {
113+ reporter.record(Entry::new("note", reason.trim()));
114+ } else if let Ok(model) = std::env::var("AGENT_MODEL_NAME") {
111115 reporter.record(Entry::new("note", &format!("Running on {model}.")));
112116 }
113117 reporter.record(Entry::new("prompt", &prompt));
+20−3
117117 lastSeen: string;
118118 };
119119
120+/**
121+ * How capable, and how costly, a model g1t routes to is: `small` (fast),
122+ * `large` (standard) or `frontier` (most capable).
123+ */
124+export type ModelTier = "small" | "large" | "frontier";
125+
126+/** The tiers, cheapest first. */
127+export const MODEL_TIERS: ModelTier[] = ["small", "large", "frontier"];
128+
120129 export type ModelSession = {
121130 token: string;
122131 billedTo: "g1t" | "workspace";
123132 providerName: string | null;
124133 model: string | null;
125− /** Names the run in AI Gateway's logs (`metadata.session`), for billing. */
134+ /**
135+ * Names the run in AI Gateway's logs (`metadata.session`) and its tokens
136+ * in billing's count, on g1t's models and the workspace's own alike.
137+ */
126138 id: string;
139+ /**
140+ * The tier the workspace chose for this work on g1t's models, instead of
141+ * Auto; the run goes there. Null for Auto, and on its own providers.
142+ */
143+ tierChoice?: ModelTier | null;
127144 };
128145
129146 export type ModelUpstream = {
143160 /** The session's id; see `ModelSession.id`. */
144161 session: string;
145162 /** For `g1t`: the tier the run was routed to. */
146− tier?: "small" | "large" | null;
163+ tier?: ModelTier | null;
147164 /** The person the run is for, by username. Null when nobody asked; never `g1t`. */
148165 requestedBy: string | null;
149166 baseUrl: string | null;
188205 task: string;
189206 hostedOpen: boolean;
190207 /** The tier the run is routed to on g1t's hosted models, for the gateway's logs. */
191− tier?: "small" | "large" | null;
208+ tier?: ModelTier | null;
192209 /** The person the run is for, by username, so its tokens show under them. */
193210 requestedBy?: string | null;
194211 }): Promise<Result<ModelSession>>;
+53−5
130130 (!name.is_empty() && name != g1t_contracts::identity::AGENT_NAME).then_some(name)
131131 }
132132
133+/// The tier a workspace chose on g1t's models for `task`: its own route's,
134+/// else the `default` route's, when that route is to g1t's models and names
135+/// one. `None` is Auto.
136+fn hosted_choice(routes: &[RouteRow], task: &str) -> Option<String> {
137+ let chosen = routes.iter().find(|route| route.task == task).or_else(|| routes.iter().find(|route| route.task == "default"))?;
138+ if chosen.connection_id.is_some() {
139+ return None;
140+ }
141+ chosen.model.as_deref().map(str::trim).filter(|model| MODEL_TIERS.contains(model)).map(str::to_owned)
142+}
143+
133144 /// A model session's public id: the start of its token's hash.
134145 fn session_id(token_hash: &str) -> String {
135146 format!("ms_{}", &token_hash[..token_hash.len().min(24)])
11591170 }
11601171 seen.push(route.task.clone());
11611172 let model = route.model.as_deref().map(str::trim).filter(|model| !model.is_empty());
1173+ // On g1t's models a route names a tier, or nothing for Auto.
1174+ if route.connection_id.is_none()
1175+ && let Some(model) = model
1176+ && !MODEL_TIERS.contains(&model)
1177+ {
1178+ return Ok(fail(FailureCode::Invalid, "On g1t's models, choose Auto, fast, standard or most capable."));
1179+ }
11621180 if let Some(id) = &route.connection_id {
11631181 let Some(row) = rows.iter().find(|row| &row.id == id && row.provider().kind() == ProviderKind::Models) else {
11641182 return Ok(fail(FailureCode::NotFound, "A route names a model provider this workspace does not have."));
12331251 Some((row, model)) => (Some(row), model.clone()),
12341252 None => (None, None),
12351253 };
1254+ // On g1t's models, a tier the workspace chose for this work replaces
1255+ // the runner's (Auto's).
1256+ let choice = if connection.is_none() { hosted_choice(&self.route_rows(&workspace).await?, &a.task) } else { None };
1257+ let tier = choice.clone().or_else(|| a.tier.clone());
12361258 self.db
12371259 .batch(vec![
12381260 self.db
12531275 rfc3339(now + MODEL_SESSION_SECONDS * 1000).into(),
12541276 optional(model.as_deref()),
12551277 // The tier is g1t's routing; it means nothing on the workspace's own provider.
1256− optional(
1257− a.tier
1258− .as_deref()
1259− .filter(|tier| connection.is_none() && matches!(*tier, "small" | "large")),
1260− ),
1278+ optional(tier.as_deref().filter(|tier| connection.is_none() && MODEL_TIERS.contains(tier))),
12611279 optional(requester(a.requested_by.as_deref()).as_deref()),
12621280 ])?,
12631281 ])
12691287 provider_name: connection.map(|row| row.name.clone()),
12701288 model,
12711289 id,
1290+ tier_choice: choice,
12721291 }))
12731292 }
12741293
15231542 assert_eq!(closable_hashes(&many).len(), 20);
15241543 }
15251544 }
1545+
1546+#[cfg(test)]
1547+mod choice_tests {
1548+ use super::{RouteRow, hosted_choice};
1549+
1550+ fn route(task: &str, connection: Option<&str>, model: Option<&str>) -> RouteRow {
1551+ RouteRow { task: task.into(), connection_id: connection.map(Into::into), model: model.map(Into::into) }
1552+ }
1553+
1554+ #[test]
1555+ fn a_tier_chosen_on_g1ts_models_is_the_works_own_or_the_defaults() {
1556+ let routes = vec![route("default", None, Some("large")), route("review", None, Some("frontier")), route("plan", None, None)];
1557+ assert_eq!(hosted_choice(&routes, "review").as_deref(), Some("frontier"));
1558+ // No route of its own: the default's.
1559+ assert_eq!(hosted_choice(&routes, "implement").as_deref(), Some("large"));
1560+ // Its own route to g1t's models on Auto: Auto, not the default's.
1561+ assert_eq!(hosted_choice(&routes, "plan"), None);
1562+ // No routes at all: Auto.
1563+ assert_eq!(hosted_choice(&[], "implement"), None);
1564+ }
1565+
1566+ #[test]
1567+ fn a_route_to_the_workspaces_own_provider_chooses_no_tier() {
1568+ let routes = vec![route("default", Some("con_1"), Some("claude-sonnet-5-5")), route("update", None, Some("huge"))];
1569+ assert_eq!(hosted_choice(&routes, "implement"), None);
1570+ // Not a tier: Auto.
1571+ assert_eq!(hosted_choice(&routes, "update"), None);
1572+ }
1573+}
+1−1
1212 session: string;
1313 person: string | null;
1414 model: string;
15− tier: "small" | "large" | null;
15+ tier: "small" | "large" | "frontier" | null;
1616 input: number;
1717 output: number;
1818 cacheRead: number;
+90−68
6262 } from "@g1t/contracts";
6363
6464 import {
65− type AgentTask,
65+ type JobKind,
66+ type PastAttempt,
6667 type RouteSignals,
67− type Tier,
6868 canReachModel,
6969 changeSize,
70− chooseTier,
70+ failuresInARow,
7171 gatewaySession,
72− lastAttemptFailed,
72+ leftLowConfidence,
7373 modelEnv,
74+ outcomesOf,
7475 parseRouting,
76+ route,
77+ taskOf,
7578 tierVars,
7679 } from "./model-env";
7780 import { hubContext } from "./hub";
151154 */
152155 HOSTED_AGENT_WORKSPACES: string;
153156 /**
154− * How g1t routes the work it pays the model for, as JSON (`AgentRouting`
155− * in model-env.ts): `tiers`, the model behind `small` and `large`, each
156− * `{ modelName, model }`; `tasks`, the tier of each kind of work, or
157− * `change` to decide a review by its change; `smallChange`, the largest
158− * change reviewed on the small tier; `largeLabels`, issue labels that
159− * keep a review large. Anything left out takes the default.
157+ * How g1t routes agent work ("Auto"), as JSON (`AgentRouting` in
158+ * model-env.ts): `tiers`, the catalogue (the model behind `small`,
159+ * `large` and `frontier`, each `{ modelName, model, price }`); `tasks`,
160+ * the tier each kind of job starts on, or `change` to size the change;
161+ * `smallChange` and `largeChange`, the bounds of a small and a large
162+ * change; `largeLabels`, `frontierLabels` and `smallLabels`, issue
163+ * labels that move work; `frontierAfter`, failures in a row before the
164+ * frontier tier; `learning`, how a repository's own runs move it.
165+ * Anything left out takes the default.
160166 */
161167 AGENT_ROUTING?: string;
162168 /**
191197 }
192198
193199 /**
194− * What routing knows about one piece of work, and how to ask whether it is
195− * a retry (asked only when the answer matters).
200+ * What routing knows about one piece of work. With `viewer`, the
201+ * repository's recent runs of the same kind are read as them, for
202+ * retries, confidence and learning; `title` narrows the same work to one
203+ * plan's brief, since plans have no pull request.
196204 */
197−type RouteInput = RouteSignals & { retried?: () => Promise<boolean> };
205+type RouteInput = RouteSignals & { viewer?: User; title?: string };
198206
199207 /** A run that takes longer than this has its token expire under it. */
200208 const TOKEN_TTL_SECONDS = 2 * 60 * 60;
561569 return null;
562570 }
563571 await this.ctx.storage.put("agentRun", opened.value);
572+ // Which model it runs on, and why, as the run's first step.
573+ if (envVars.AGENT_MODEL_REASON) {
574+ await reportRun(this.env.WORK, opened.value, { steps: [envVars.AGENT_MODEL_REASON] }).catch(() => undefined);
575+ }
564576 return opened.value;
565577 }
566578
13541366 }
13551367
13561368 /**
1357− * What a sandbox needs to reach the model routed for `task`, having
1358− * opened the run the repository's workspace will be charged for. Refused
1359− * when that workspace has no credit.
1369+ * What a sandbox needs to reach the model routed for `kind`, having
1370+ * opened the run the repository's workspace will be charged for.
1371+ * Refused when that workspace has no credit.
13601372 *
1361− * On g1t's hosted models the work goes to the cheapest tier that can do
1362− * it (`chooseTier`), by what `route` says about it. Whether this is a
1363− * retry is asked only when it would change the answer: when the work
1364− * would otherwise go to the small tier.
1373+ * "Auto" (`route` in model-env.ts) picks the cheapest tier that can do
1374+ * the work, from what `input` says about it and the repository's own
1375+ * recent runs of the same kind (read once, only for a person who can see
1376+ * them), unless the workspace chose a tier for this work. The choice and
1377+ * why go to the sandbox (`AGENT_MODEL_REASON`), which records them on
1378+ * the run and in its session. A workspace's own Anthropic key with no
1379+ * model of its own named is routed the same way.
13651380 *
13661381 * `requestedBy` is the person the run is for, by username, so the run's
13671382 * tokens are counted under them.
13681383 */
13691384 private async modelEnv(
1370− task: AgentTask,
1385+ kind: JobKind,
13711386 repo: RepoPath,
13721387 pull: number,
13731388 requestedBy: string | null,
1374− route: RouteInput = {},
1389+ input: RouteInput = {},
13751390 ): Promise<Result<Record<string, string>>> {
13761391 const routing = parseRouting(this.env.AGENT_ROUTING);
1377− let tier: Tier = chooseTier(task, route, routing);
1378− if (tier === "small" && route.retried && (await route.retried().catch(() => false))) {
1379− tier = chooseTier(task, { ...route, retry: true }, routing);
1392+ const task = taskOf(kind);
1393+ const signals: RouteSignals = { ...input };
1394+ if (input.viewer) {
1395+ // One read: the repository's recent runs of this kind, newest first.
1396+ // The same work's attempts are among them.
1397+ const recent = await agentsClient(this.env.WORK)
1398+ .listRuns(input.viewer, { repo, kind, limit: routing.learning.window })
1399+ .catch(() => null);
1400+ const runs: PastAttempt[] = recent?.ok ? recent.value : [];
1401+ const same = input.title !== undefined ? runs : runs.filter((run) => pull > 0 && run.number === pull);
1402+ signals.failures = Math.max(signals.failures ?? 0, failuresInARow(same, input.title));
1403+ signals.lowConfidence = signals.lowConfidence ?? leftLowConfidence(same);
1404+ signals.history = outcomesOf(runs, routing);
13801405 }
1406+ let routed = route(kind, signals, routing);
13811407 const tags = { repo: `${repo.namespace}/${repo.name}`, pull };
13821408 // Where the run's model requests go, by the workspace's routes: g1t's
13831409 // hosted models, or one of its own providers.
13891415 number: pull,
13901416 task,
13911417 hostedOpen: (await this.modelAccess(repo.namespace)).hosted,
1392− tier,
1418+ tier: routed.tier,
13931419 requestedBy,
13941420 });
13951421 if (!opened.ok) return opened;
13961422 session = opened.value;
1423+ // The workspace chose a tier for this work instead of Auto.
1424+ if (session.tierChoice) routed = route(kind, { chosen: session.tierChoice }, routing);
13971425 }
13981426 const own = session?.billedTo === "workspace";
1427+ const tier = routed.tier;
13991428 // Straight to the gateway, without the proxy: the run still gets a
14001429 // session there, so billing settles it to what the gateway priced it
14011430 // at instead of leaving the sandbox's own figure.
14021431 const direct = !session && this.env.AI_GATEWAY_ID ? gatewaySession() : undefined;
1403− // A workspace's own provider is not routed by tier: it runs the model
1404− // its route names, or for an Anthropic provider, the large tier's.
1405− const routed = routing.tiers[own ? "large" : tier];
1406− const model = session?.model ?? routed.model;
1407− const modelName = session?.model ?? routed.modelName;
1432+ // A workspace's own provider runs the model its route names; with none
1433+ // named (an Anthropic key), the tier's, as on g1t's models.
1434+ const named = own && session?.model ? session.model : null;
1435+ const model = named ?? routing.tiers[tier].model;
1436+ const modelName = named ?? routing.tiers[tier].modelName;
1437+ const reason = named ? `Used ${named}: the workspace's route for this work names it.` : routed.reason;
14081438 const ticket = await billingClient(this.env.BILLING).startRun({
14091439 workspace: repo.namespace,
14101440 repo,
14121442 task,
14131443 model: own ? `${modelName} (${session?.providerName ?? "own provider"})` : modelName,
14141444 billedTo: own ? "workspace" : "g1t",
1415− session: own ? null : (session?.id ?? direct ?? null),
1416− tier: own ? null : tier,
1445+ // On the workspace's own provider too: billing counts its tokens by
1446+ // it for the agent rate.
1447+ session: session?.id ?? direct ?? null,
1448+ tier: named ? null : tier,
14171449 });
14181450 if (!ticket.ok) return ticket;
14191451 const vars: Record<string, string> = session
14201452 ? {
1421− // On g1t's models, the tier's model, and the small tier's for the
1422− // harness's own small tasks.
1423− ...(own ? {} : tierVars(routing, tier)),
1453+ // The tier's model, and the small tier's for the harness's own
1454+ // small tasks; a route that names its model uses it for both.
1455+ ...tierVars(routing, tier),
14241456 ANTHROPIC_MODEL: model,
14251457 AGENT_MODEL_NAME: own ? `${modelName}, through ${session.providerName}` : modelName,
14261458 ANTHROPIC_BASE_URL: `${this.env.MODELS_URL!.replace(/\/+$/, "")}/anthropic`,
14281460 ANTHROPIC_API_KEY: session.token,
14291461 // An endpoint that names models its own way gets its model for
14301462 // the harness's small tasks too.
1431− ...(session.model ? { ANTHROPIC_SMALL_FAST_MODEL: session.model } : {}),
1463+ ...(named ? { ANTHROPIC_SMALL_FAST_MODEL: named, ANTHROPIC_DEFAULT_HAIKU_MODEL: named } : {}),
14321464 }
14331465 : modelEnv(this.env, routing, task, tier, direct ? { ...tags, session: direct } : tags);
1466+ // Why this model: shown on the run and at the top of its session.
1467+ vars.AGENT_MODEL_REASON = reason;
14341468 if (ticket.value) {
14351469 // How the sandbox says what the run cost. Kept from the agent.
14361470 vars.BILLING_RUN = ticket.value.runId;
15631597
15641598 /** The same, for a step g1t takes by itself: a refusal stops the step. */
15651599 private async modelEnvOrThrow(
1566− task: AgentTask,
1600+ task: JobKind,
15671601 repo: RepoPath,
15681602 pull: number,
15691603 requestedBy: string | null,
15741608 return vars.value;
15751609 }
15761610
1577− /**
1578− * Whether the latest run of the same work failed, so that this one is a
1579− * retry: the same kind of run on the same pull request, or for a plan,
1580− * the latest plan with the same brief. Read as the one the run is for;
1581− * unknown counts as not.
1582− */
1583− private async failedBefore(
1584− viewer: User,
1585− repo: RepoPath,
1586− kind: "review" | "update" | "plan",
1587− number: number | null,
1588− title?: string,
1589− ): Promise<boolean> {
1590− const runs = await agentsClient(this.env.WORK)
1591− .listRuns(viewer, { repo, kind, ...(number != null ? { number } : {}), limit: 1 })
1592− .catch(() => null);
1593− return runs?.ok ? lastAttemptFailed(runs.value, title) : false;
1594− }
1595−
15961611 /** Whether sandboxes have a way to reach a model at all. */
15971612 private modelsReachable(): boolean {
15981613 return Boolean(this.env.MODELS_URL) || canReachModel(this.env);
22662281 job.repo,
22672282 job.author,
22682283 ),
2269− ...(await this.modelEnvOrThrow("implement", job.repo, job.number, job.author.username)),
2284+ // Work handed over to the change: routed as revising it.
2285+ ...(await this.modelEnvOrThrow("revise", job.repo, job.number, job.author.username, {
2286+ labels: job.issue?.labels ?? [],
2287+ viewer: job.author,
2288+ })),
22702289 },
22712290 });
22722291 }
23192338 job.repo,
23202339 job.author,
23212340 ),
2322− ...(await this.modelEnvOrThrow("implement", job.repo, job.number, startedBy ?? job.author.username)),
2341+ ...(await this.modelEnvOrThrow("revise", job.repo, job.number, startedBy ?? job.author.username, {
2342+ labels: job.issue?.labels ?? [],
2343+ viewer: job.author,
2344+ // The first revision is the first time the change fell short;
2345+ // each after it is another failure in a row.
2346+ failures: Math.max(0, job.round - 1),
2347+ })),
23232348 },
23242349 });
23252350 }
25752600 repo,
25762601 actor,
25772602 ),
2578− ...(await this.modelEnvOrThrow("update", repo, number, actor.username, {
2579− retried: () => this.failedBefore(actor, repo, "update", number),
2580− })),
2603+ ...(await this.modelEnvOrThrow("update", repo, number, actor.username, { viewer: actor })),
25812604 },
25822605 });
25832606 }
26342657 const model = await this.modelEnv("review", repo, number, job.author.username, {
26352658 change: job.files?.length ? changeSize(job.files, job.sensitive ?? []) : null,
26362659 labels: job.issue?.labels ?? [],
2637− retried: () => this.failedBefore(job.author, repo, "review", number),
2660+ viewer: job.author,
26382661 });
26392662 if (!model.ok) {
26402663 await workClient(this.env.WORK).failReview(job.runId, job.token, model.error.message);
26952718 const started = await work.startPlan(actor, repo, brief);
26962719 if (!started.ok) return started;
26972720 const job = started.value;
2698− const model = await this.modelEnv("plan", repo, 0, actor.username, {
2699− retried: () => this.failedBefore(actor, repo, "plan", null, job.brief),
2700− });
2721+ const model = await this.modelEnv("plan", repo, 0, actor.username, { viewer: actor, title: job.brief });
27012722 if (!model.ok) {
27022723 await work.failPlan(job.planId, job.token, model.error.message);
27032724 return model;
28502871 // Opened without a branch, so it has a fork.
28512872 const fork = pull.fork!;
28522873
2853− const model = await this.modelEnv("implement", repo, pull.number, actor.username);
2874+ const model = await this.modelEnv("implement", repo, pull.number, actor.username, { labels: issue.labels, viewer: actor });
28542875 if (!model.ok) {
28552876 await work.closePull(actor, repo, pull.number);
28562877 return model;
29973018 ({ title, body } = found.value.issue);
29983019 comments = found.value.comments;
29993020 }
3000− const model = await this.modelEnv("implement", job.repo, job.number, job.actor.username);
3021+ // A question answered from the code: it changes nothing.
3022+ const model = await this.modelEnv("answer", job.repo, job.number, job.actor.username, { viewer: job.actor });
30013023 if (!model.ok) return model;
30023024 const source = job.pull?.source ?? job.repo;
30033025 // Reads the code; pushes nothing. Its answer is posted with its tools.
+134−10
99 canReachModel,
1010 changeSize,
1111 chooseTier,
12+ failuresInARow,
1213 gatewaySession,
1314 lastAttemptFailed,
15+ leftLowConfidence,
1416 modelEnv,
17+ outcomesOf,
1518 parseRouting,
19+ route,
20+ taskOf,
21+ tierOfModel,
1622 } from "./model-env.ts";
1723
1824 const routes: AgentRouting = {
2026 tiers: {
2127 small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" },
2228 large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" },
29+ frontier: { modelName: "Claude Opus 5.5", model: "claude-opus-5-5" },
2330 },
2431 };
2532 const tags = { repo: "acme/site", pull: 12 };
3744
3845 const small: ChangeSize = { files: 3, lines: 80, sensitive: [] };
3946
40−test("planning and catching up run on the small tier, making a change on the large", () => {
41− assert.equal(chooseTier("plan", {}, routes), "small");
47+test("Auto starts each kind of job on its tier: fast for catching up and answering, standard for changes, most capable for plans", () => {
4248 assert.equal(chooseTier("update", {}, routes), "small");
49+ assert.equal(chooseTier("answer", {}, routes), "small");
4350 assert.equal(chooseTier("implement", {}, routes), "large");
51+ assert.equal(chooseTier("revise", {}, routes), "large");
4452 assert.equal(chooseTier("implement", { change: small }, routes), "large");
53+ assert.equal(chooseTier("plan", {}, routes), "frontier");
54+});
55+
56+test("revising and answering go to the workspace's implement route and bill", () => {
57+ assert.equal(taskOf("revise"), "implement");
58+ assert.equal(taskOf("answer"), "implement");
59+ assert.equal(taskOf("review"), "review");
60+ assert.equal(taskOf("plan"), "plan");
4561 });
4662
47−test("a review is small only for a small change that touches nothing sensitive", () => {
63+test("a review is sized by its change: fast when small and safe, most capable when large", () => {
4864 assert.equal(chooseTier("review", { change: small }, routes), "small");
4965 assert.equal(chooseTier("review", { change: { ...small, lines: 200, files: 10 } }, routes), "small");
5066 assert.equal(chooseTier("review", { change: { ...small, lines: 201 } }, routes), "large");
5268 assert.equal(chooseTier("review", { change: { ...small, sensitive: ["CI workflows"] } }, routes), "large");
5369 assert.equal(chooseTier("review", { change: small, labels: ["Security"] }, routes), "large");
5470 assert.equal(chooseTier("review", { change: small, labels: ["docs"] }, routes), "small");
71+ assert.equal(chooseTier("review", { change: { ...small, files: 61 } }, routes), "frontier");
72+ assert.equal(chooseTier("review", { change: { ...small, lines: 3001 } }, routes), "frontier");
5573 });
5674
57−test("a review of a change g1t cannot size runs on the large tier", () => {
75+test("a review of a change g1t cannot size runs on the standard tier", () => {
5876 assert.equal(chooseTier("review", {}, routes), "large");
5977 assert.equal(chooseTier("review", { change: null }, routes), "large");
6078 assert.equal(chooseTier("review", { change: { files: 0, lines: 0, sensitive: [] } }, routes), "large");
6179 });
6280
63−test("a retry after a failed attempt goes up to the large tier", () => {
64− assert.equal(chooseTier("plan", { retry: true }, routes), "large");
81+test("labels move work: architecture to the most capable, documentation to the fast tier", () => {
82+ assert.equal(chooseTier("implement", { labels: ["Architecture"] }, routes), "frontier");
83+ assert.equal(chooseTier("review", { change: small, labels: ["architecture"] }, routes), "frontier");
84+ assert.equal(chooseTier("implement", { labels: ["docs"] }, routes), "small");
85+ assert.equal(chooseTier("answer", { labels: ["typo"] }, routes), "small");
86+ // A small label never takes a review or a plan down.
87+ assert.equal(chooseTier("plan", { labels: ["docs"] }, routes), "frontier");
88+ // Security outranks documentation.
89+ assert.equal(chooseTier("implement", { labels: ["docs", "security"] }, routes), "large");
90+});
91+
92+test("a failed attempt goes one tier up, and repeated failures to the most capable", () => {
6593 assert.equal(chooseTier("update", { retry: true }, routes), "large");
6694 assert.equal(chooseTier("review", { change: small, retry: true }, routes), "large");
95+ assert.equal(chooseTier("implement", { failures: 1 }, routes), "frontier");
96+ assert.equal(chooseTier("update", { failures: 2 }, routes), "frontier");
97+ assert.equal(chooseTier("plan", { failures: 1 }, routes), "frontier");
98+ const why = route("update", { failures: 2 }, routes).reason;
99+ assert.equal(why, "Used the most capable model (Claude Opus 5.5): the last 2 attempts at this work failed.");
100+});
101+
102+test("a change left at low confidence sends the next attempt one tier up", () => {
103+ assert.equal(chooseTier("revise", { lowConfidence: true }, routes), "frontier");
104+ assert.equal(chooseTier("update", { lowConfidence: true }, routes), "large");
105+});
106+
107+test("a tier the workspace chose wins over Auto", () => {
108+ const chosen = route("review", { change: small, failures: 3, chosen: "large" }, routes);
109+ assert.equal(chosen.tier, "large");
110+ assert.equal(chosen.reason, "Used the standard model (Claude Sonnet 5.5): the workspace chose the standard model for this work.");
111+});
112+
113+const ok = (tier: "small" | "large" | "frontier", n: number) => Array.from({ length: n }, () => ({ tier, ok: true }));
114+const bad = (tier: "small" | "large" | "frontier", n: number) => Array.from({ length: n }, () => ({ tier, ok: false }));
115+
116+test("a repository whose fast runs finish nearly always steps the work down", () => {
117+ const history = [...ok("small", 9), ...bad("small", 1)];
118+ const routed = route("implement", { history }, routes);
119+ assert.equal(routed.tier, "small");
120+ assert.equal(routed.reason, "Used a fast model (Claude Haiku 4.5): it finished 9 of its last 10 runs like this here.");
121+ // Too few runs to tell, or not quite enough of them finished: no change.
122+ assert.equal(chooseTier("implement", { history: ok("small", 4) }, routes), "large");
123+ assert.equal(chooseTier("implement", { history: [...ok("small", 8), ...bad("small", 2)] }, routes), "large");
124+ // Plans step down from the most capable model the same way.
125+ assert.equal(chooseTier("plan", { history: ok("large", 6) }, routes), "large");
126+});
127+
128+test("learning never steps down sensitive or labelled work, nor a retry", () => {
129+ const history = ok("small", 10);
130+ assert.equal(chooseTier("review", { change: { ...small, files: 20, sensitive: ["secrets"] }, history }, routes), "large");
131+ assert.equal(chooseTier("implement", { labels: ["security"], history }, routes), "large");
132+ assert.equal(chooseTier("implement", { failures: 1, history }, routes), "frontier");
133+});
134+
135+test("a tier that fails half its runs in a repository hands the work up", () => {
136+ const history = [...bad("large", 3), ...ok("large", 2)];
137+ const routed = route("implement", { history }, routes);
138+ assert.equal(routed.tier, "frontier");
139+ assert.equal(routed.reason, "Used the most capable model (Claude Opus 5.5): the standard model failed 3 of its last 5 runs like this here.");
140+});
141+
142+test("past runs are read by the tier they ran on and whether they did the work", () => {
143+ const runs = [
144+ { status: "running", model: "Claude Haiku 4.5" },
145+ { status: "succeeded", model: "Claude Haiku 4.5" },
146+ { status: "succeeded", model: "Claude Sonnet 5.5", confidence: { level: "low" } },
147+ { status: "failed", model: "Claude Opus 5.5" },
148+ { status: "stopped", halted: null, model: "Claude Haiku 4.5" },
149+ { status: "stopped", halted: "budget", model: "claude-haiku-4-5-20251001" },
150+ { status: "succeeded", model: "Claude Sonnet 5.5, through Acme" },
151+ ];
152+ assert.deepEqual(outcomesOf(runs, routes), [
153+ { tier: "small", ok: true },
154+ { tier: "large", ok: false },
155+ { tier: "frontier", ok: false },
156+ { tier: "small", ok: false },
157+ { tier: null, ok: true },
158+ ]);
159+ assert.equal(tierOfModel("Claude Opus 5.5", routes), "frontier");
160+ assert.equal(tierOfModel(null, routes), null);
161+});
162+
163+test("failures are counted in a row, newest first, and confidence is read from the last finished one", () => {
164+ assert.equal(failuresInARow([]), 0);
165+ assert.equal(failuresInARow([{ status: "failed" }, { status: "failed" }, { status: "succeeded" }, { status: "failed" }]), 2);
166+ assert.equal(failuresInARow([{ status: "stopped", halted: "time" }, { status: "failed" }]), 2);
167+ assert.equal(failuresInARow([{ status: "stopped", halted: null }, { status: "failed" }]), 0);
168+ assert.equal(failuresInARow([{ status: "failed", title: "Add search" }, { status: "failed", title: "Other" }], "Add search"), 1);
169+ assert.equal(leftLowConfidence([{ status: "succeeded", confidence: { level: "low" } }]), true);
170+ assert.equal(leftLowConfidence([{ status: "succeeded", confidence: { level: "medium" } }]), false);
171+ assert.equal(leftLowConfidence([]), false);
67172 });
68173
69174 test("a retry is the same work again after its latest attempt failed", () => {
77182 assert.equal(lastAttemptFailed([{ status: "failed", title: "Add search" }], "Add billing"), false);
78183 });
79184
80−test("the configuration decides the tiers, the rules and the limits", () => {
185+test("the configuration decides the catalogue, the rules and the limits", () => {
81186 const parsed = parseRouting(
82187 JSON.stringify({
83− tiers: { small: { modelName: "Small", model: "small-1" } },
84− tasks: { update: "large" },
188+ tiers: { small: { modelName: "Small", model: "small-1" }, frontier: { model: "" } },
189+ tasks: { update: "large", plan: "huge", answer: "change" },
85190 smallChange: { lines: 50 },
191+ frontierLabels: ["hard"],
192+ learning: { minRuns: 3 },
86193 }),
87194 );
88195 assert.deepEqual(parsed.tiers.small, { modelName: "Small", model: "small-1" });
89196 assert.deepEqual(parsed.tiers.large, DEFAULT_ROUTING.tiers.large);
197+ // A tier with no model keeps the default.
198+ assert.deepEqual(parsed.tiers.frontier, DEFAULT_ROUTING.tiers.frontier);
90199 assert.equal(chooseTier("update", {}, parsed), "large");
91− assert.equal(chooseTier("plan", {}, parsed), "small");
200+ // A rule that names no tier keeps the default.
201+ assert.equal(chooseTier("plan", {}, parsed), "frontier");
202+ assert.equal(parsed.tasks.answer, "change");
92203 assert.equal(chooseTier("review", { change: { ...small, lines: 51 } }, parsed), "large");
93204 assert.equal(chooseTier("review", { change: { ...small, files: 10, lines: 50 } }, parsed), "small");
205+ assert.equal(chooseTier("implement", { labels: ["hard"] }, parsed), "frontier");
206+ assert.equal(parsed.learning.minRuns, 3);
207+ assert.equal(parsed.learning.window, DEFAULT_ROUTING.learning.window);
94208 assert.deepEqual(parseRouting(undefined), DEFAULT_ROUTING);
95209 assert.deepEqual(parseRouting("not json"), DEFAULT_ROUTING);
210+ assert.deepEqual(parseRouting("[1]"), DEFAULT_ROUTING);
211+});
212+
213+test("the routing the runner ships with is the default, written as configuration", async () => {
214+ const { readFile } = await import("node:fs/promises");
215+ const config = await readFile(new URL("../wrangler.jsonc", import.meta.url), "utf8");
216+ const line = config.split("\n").find((l) => l.trim().startsWith('"AGENT_ROUTING"'));
217+ assert.ok(line, "wrangler.jsonc sets AGENT_ROUTING");
218+ const value = JSON.parse(line.trim().replace(/^"AGENT_ROUTING":\s*/, "").replace(/,$/, ""));
219+ assert.deepEqual(parseRouting(value), DEFAULT_ROUTING);
96220 });
97221
98222 test("a change's size is its files and the lines added and removed", () => {
+316−43
1−/** The kinds of work a g1t agent does. Each is routed on its own. */
1+/**
2+ * The kinds of work a g1t agent does, as model routes and billing name
3+ * them. A workspace routes each to a provider (Integrations → Models).
4+ */
25 export type AgentTask = "implement" | "review" | "update" | "plan";
36
47 /**
5− * How capable, and how costly, a model is. g1t's hosted models come in
6− * two: `small` for work a smaller model does as well, `large` for the rest.
8+ * What the router routes: every kind of agent job. Revising a change and
9+ * answering a question are routed on their own, and go to the workspace's
10+ * `implement` route (`taskOf`).
711 */
8−export type Tier = "small" | "large";
12+export type JobKind = AgentTask | "revise" | "answer";
913
14+/** The route and the bill a job goes on. */
15+export function taskOf(kind: JobKind): AgentTask {
16+ return kind === "revise" || kind === "answer" ? "implement" : kind;
17+}
18+
19+/**
20+ * How capable, and how costly, a model is: `small` (fast and cheap, for
21+ * work a smaller model does as well), `large` (the standard, most
22+ * changes) and `frontier` (the most capable, for hard work only).
23+ */
24+export type Tier = "small" | "large" | "frontier";
25+
26+/** Cheapest first. */
27+export const TIERS: Tier[] = ["small", "large", "frontier"];
28+
29+/** Per million tokens, in US dollars: what the provider lists. */
30+export type TokenPrice = { input: number; output: number; cacheRead: number; cacheWrite: number };
31+
1032 /** Where one kind of work goes: what people see, and what is sent. */
1133 export type ModelRoute = {
1234 /** The model's public name, e.g. `Claude Sonnet 5.5`. */
1335 modelName: string;
1436 /** The identifier sent to the provider. */
1537 model: string;
38+ /**
39+ * The provider's list price, for estimates (the savings report, routing
40+ * by cost). Never what anyone is charged: runs are charged what AI
41+ * Gateway priced them at.
42+ */
43+ price?: TokenPrice;
1644 };
1745
46+/** How a job's rule decides: a tier, or `change` to size the change it reads. */
47+export type TaskRule = Tier | "change";
48+
49+/**
50+ * Learning from a repository's own runs: of its last `window` runs of the
51+ * same kind, a cheaper tier that finished at least `stepDownAt` of at
52+ * least `minRuns` takes the work; a tier that finished less than
53+ * `stepUpAt` of at least `minRuns` hands it up.
54+ */
55+export type Learning = { window: number; minRuns: number; stepDownAt: number; stepUpAt: number };
56+
1857 /**
19− * g1t's routing policy for the runs it pays the model for. Nobody
20− * assigning an agent picks a model; the work decides, here, and the
21− * operator changes it in one place (`AGENT_ROUTING` in wrangler.jsonc).
58+ * g1t's routing policy. Nobody assigning an agent has to pick a model:
59+ * "Auto" decides here, by the work, and the operator changes the policy in
60+ * one place (`AGENT_ROUTING` in wrangler.jsonc), never in code.
2261 */
2362 export type AgentRouting = {
24− /** The model behind each tier. */
63+ /** The model behind each tier: the catalogue. */
2564 tiers: Record<Tier, ModelRoute>;
26− /**
27− * The tier each kind of work runs on. `change` decides by the change the
28− * work reads: small when it is small and touches nothing sensitive.
29− */
30− tasks: Record<AgentTask, Tier | "change">;
65+ /** The tier each kind of job starts from, or `change` to size it. */
66+ tasks: Record<JobKind, TaskRule>;
3167 /** The largest change `change` sends to the small tier. */
3268 smallChange: { files: number; lines: number };
33− /** Labels on the issue behind the work that send `change` to the large tier. */
69+ /** A change larger than this (either) is reviewed on the frontier tier. */
70+ largeChange: { files: number; lines: number };
71+ /** Issue labels that keep work off the small tier. */
3472 largeLabels: string[];
73+ /** Issue labels that send work to the frontier tier. */
74+ frontierLabels: string[];
75+ /** Issue labels that let changes and answers start on the small tier. */
76+ smallLabels: string[];
77+ /** Failed attempts at the same work, in a row, before the frontier tier. */
78+ frontierAfter: number;
79+ learning: Learning;
3580 };
3681
3782 /** What a change is, as far as routing cares. */
4691 sensitive: string[];
4792 };
4893
94+/** One past run of the same kind of job in the repository, for learning. */
95+export type PastOutcome = {
96+ /** The tier it ran on, when it ran on one of g1t's. */
97+ tier: Tier | null;
98+ /** It finished, and did not leave a change g1t had low confidence in. */
99+ ok: boolean;
100+};
101+
49102 /** What g1t knows about one piece of work when it routes it. */
50103 export type RouteSignals = {
51104 /** The change the work reads; null or absent when g1t does not know it. */
52105 change?: ChangeSize | null;
53106 /** Labels on the issue the work is for. */
54107 labels?: string[];
55− /** The last attempt at the same work failed. */
108+ /** The last attempt at the same work failed. Same as `failures: 1`. */
56109 retry?: boolean;
110+ /** Failed attempts at the same work, in a row, most recent last. */
111+ failures?: number;
112+ /** The last attempt finished, but left a change g1t has low confidence in. */
113+ lowConfidence?: boolean;
114+ /** Recent runs of the same kind in this repository, newest first. */
115+ history?: PastOutcome[];
116+ /** A tier the workspace chose for this work instead of Auto. */
117+ chosen?: Tier | null;
118+};
119+
120+/** The router's answer: the tier, and why, in one line people can read. */
121+export type Routed = {
122+ tier: Tier;
123+ /** E.g. `Used a fast model (Claude Haiku 4.5): small change, 3 files and 80 lines.` */
124+ reason: string;
57125 };
58126
59127 /** The routing g1t ships with, for whatever the configuration leaves out. */
60128 export const DEFAULT_ROUTING: AgentRouting = {
61129 tiers: {
62− small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" },
63− large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" },
130+ small: {
131+ modelName: "Claude Haiku 4.5",
132+ model: "claude-haiku-4-5-20251001",
133+ price: { input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 },
134+ },
135+ large: {
136+ modelName: "Claude Sonnet 5.5",
137+ model: "claude-sonnet-5-5",
138+ price: { input: 2, output: 10, cacheRead: 0.2, cacheWrite: 2.5 },
139+ },
140+ frontier: {
141+ modelName: "Claude Opus 5.5",
142+ model: "claude-opus-5-5",
143+ price: { input: 4, output: 20, cacheRead: 0.2, cacheWrite: 5 },
144+ },
64145 },
65− tasks: { implement: "large", review: "change", update: "small", plan: "small" },
146+ tasks: { implement: "large", revise: "large", answer: "small", review: "change", update: "small", plan: "frontier" },
66147 smallChange: { files: 10, lines: 200 },
148+ largeChange: { files: 60, lines: 3000 },
67149 largeLabels: ["security"],
150+ frontierLabels: ["architecture"],
151+ smallLabels: ["documentation", "docs", "typo"],
152+ frontierAfter: 2,
153+ learning: { window: 20, minRuns: 5, stepDownAt: 0.9, stepUpAt: 0.5 },
68154 };
69155
156+function isTier(value: unknown): value is Tier {
157+ return value === "small" || value === "large" || value === "frontier";
158+}
159+
70160 /**
71161 * The routing in `AGENT_ROUTING`, with anything it leaves out taken from
72− * `DEFAULT_ROUTING`. An unset or unreadable value is the default.
162+ * `DEFAULT_ROUTING`. An unset or unreadable value is the default, and so is
163+ * any single rule that names no tier.
73164 */
74165 export function parseRouting(json: string | undefined): AgentRouting {
75166 let given: Partial<AgentRouting> = {};
76167 try {
77− given = json ? (JSON.parse(json) as Partial<AgentRouting>) : {};
168+ const parsed: unknown = json ? JSON.parse(json) : {};
169+ given = parsed && typeof parsed === "object" && !Array.isArray(parsed) ? (parsed as Partial<AgentRouting>) : {};
78170 } catch {
79171 console.log("AGENT_ROUTING is not JSON; using the default routing");
80172 }
173+ const tasks = { ...DEFAULT_ROUTING.tasks };
174+ for (const [kind, rule] of Object.entries(given.tasks ?? {})) {
175+ if (kind in tasks && (isTier(rule) || rule === "change")) tasks[kind as JobKind] = rule;
176+ }
177+ const tiers = { ...DEFAULT_ROUTING.tiers };
178+ for (const tier of TIERS) {
179+ const route = given.tiers?.[tier];
180+ if (route && typeof route.model === "string" && route.model) {
181+ // A model named without a price has none: estimates leave it out
182+ // rather than price it as another model.
183+ tiers[tier] = { modelName: route.modelName || route.model, model: route.model, ...(route.price ? { price: route.price } : {}) };
184+ }
185+ }
186+ const labels = (list: unknown, fallback: string[]) =>
187+ Array.isArray(list) ? list.filter((label): label is string => typeof label === "string") : fallback;
81188 return {
82− tiers: { ...DEFAULT_ROUTING.tiers, ...given.tiers },
83− tasks: { ...DEFAULT_ROUTING.tasks, ...given.tasks },
189+ tiers,
190+ tasks,
84191 smallChange: { ...DEFAULT_ROUTING.smallChange, ...given.smallChange },
85− largeLabels: given.largeLabels ?? DEFAULT_ROUTING.largeLabels,
192+ largeChange: { ...DEFAULT_ROUTING.largeChange, ...given.largeChange },
193+ largeLabels: labels(given.largeLabels, DEFAULT_ROUTING.largeLabels),
194+ frontierLabels: labels(given.frontierLabels, DEFAULT_ROUTING.frontierLabels),
195+ smallLabels: labels(given.smallLabels, DEFAULT_ROUTING.smallLabels),
196+ frontierAfter: typeof given.frontierAfter === "number" && given.frontierAfter >= 1 ? given.frontierAfter : DEFAULT_ROUTING.frontierAfter,
197+ learning: { ...DEFAULT_ROUTING.learning, ...given.learning },
86198 };
87199 }
88200
201+/** How each tier is named to people. */
202+export const TIER_LABEL: Record<Tier, { noun: string; used: string }> = {
203+ small: { noun: "fast", used: "Used a fast model" },
204+ large: { noun: "standard", used: "Used the standard model" },
205+ frontier: { noun: "most capable", used: "Used the most capable model" },
206+};
207+
208+const up = (tier: Tier): Tier => TIERS[Math.min(TIERS.indexOf(tier) + 1, TIERS.length - 1)];
209+const down = (tier: Tier): Tier => TIERS[Math.max(TIERS.indexOf(tier) - 1, 0)];
210+const rank = (tier: Tier) => TIERS.indexOf(tier);
211+
212+function plural(n: number, one: string, many: string): string {
213+ return `${n} ${n === 1 ? one : many}`;
214+}
215+
216+/** How a tier did in the repository's recent runs of the same kind. */
217+export function record(history: PastOutcome[], tier: Tier, window: number): { runs: number; ok: number } {
218+ const recent = history.slice(0, window).filter((run) => run.tier === tier);
219+ return { runs: recent.length, ok: recent.filter((run) => run.ok).length };
220+}
221+
89222 /**
90− * The tier one piece of work runs on: the cheapest that can do it.
91− * Planning and catching up are small; making a change is large; a review
92− * is small for a small change that touches nothing sensitive, and large
93− * for anything else, including a change g1t does not know the size of.
94− * A retry after a failed attempt is always large, so that work the small
95− * tier could not finish goes up rather than failing again the same way.
223+ * Where one job runs, and why. In order:
224+ *
225+ * 1. A tier the workspace chose for this work is used as chosen.
226+ * 2. The job's rule gives the starting tier: a fixed tier, or for
227+ * `change`, the change's size (small and touching nothing sensitive:
228+ * small; larger than `largeChange`: frontier; unknown or anything
229+ * else: large). Issue labels move it: `frontierLabels` to the frontier,
230+ * `largeLabels` off the small tier, `smallLabels` let a change or an
231+ * answer start small.
232+ * 3. Escalation: `frontierAfter` failures in a row go to the frontier; one
233+ * failure, or a last attempt that left low confidence, one tier up.
234+ * 4. Otherwise, learning from the repository's own runs of the same kind:
235+ * one tier down when the cheaper tier finished nearly all of its recent
236+ * ones (never for sensitive or labelled work), one tier up when this
237+ * tier failed half of its own.
96238 */
97−export function chooseTier(task: AgentTask, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Tier {
98− if (signals.retry) return "large";
99− const rule = routing.tasks[task] ?? "large";
100− if (rule !== "change") return rule;
101− const change = signals.change;
102− if (!change || change.files === 0) return "large";
103− const large = new Set(routing.largeLabels.map((label) => label.toLowerCase()));
104− if ((signals.labels ?? []).some((label) => large.has(label.toLowerCase()))) return "large";
105− const small =
106− change.sensitive.length === 0 &&
107− change.files <= routing.smallChange.files &&
108− change.lines <= routing.smallChange.lines;
109− return small ? "small" : "large";
239+export function route(kind: JobKind, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Routed {
240+ const say = (tier: Tier, why: string): Routed => ({
241+ tier,
242+ reason: `${TIER_LABEL[tier].used} (${routing.tiers[tier].modelName}): ${why}.`,
243+ });
244+ if (signals.chosen && isTier(signals.chosen)) {
245+ return say(signals.chosen, `the workspace chose the ${TIER_LABEL[signals.chosen].noun} model for this work`);
246+ }
247+
248+ const labels = new Set((signals.labels ?? []).map((label) => label.toLowerCase()));
249+ const has = (list: string[]) => list.find((label) => labels.has(label.toLowerCase()));
250+ const rule = routing.tasks[kind] ?? "large";
251+ let tier: Tier;
252+ let why: string;
253+ // Sensitive or labelled work is never stepped down by learning.
254+ let pinned = false;
255+ if (rule === "change") {
256+ const change = signals.change;
257+ if (!change || change.files === 0) {
258+ [tier, why] = ["large", "the change's size is not known"];
259+ } else if (change.sensitive.length > 0) {
260+ [tier, why, pinned] = ["large", `it touches ${change.sensitive.join(", ")}`, true];
261+ } else if (change.files > routing.largeChange.files || change.lines > routing.largeChange.lines) {
262+ [tier, why] = ["frontier", `large change, ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`];
263+ } else if (change.files <= routing.smallChange.files && change.lines <= routing.smallChange.lines) {
264+ [tier, why] = ["small", `small change, ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`];
265+ } else {
266+ [tier, why] = ["large", `a change of ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`];
267+ }
268+ } else {
269+ tier = rule;
270+ why = DEFAULT_WHY[kind];
271+ }
272+ const frontierLabel = has(routing.frontierLabels);
273+ const largeLabel = has(routing.largeLabels);
274+ const smallLabel = has(routing.smallLabels);
275+ if (frontierLabel) {
276+ [tier, why, pinned] = ["frontier", `the issue is labelled ${frontierLabel}`, true];
277+ } else if (largeLabel && tier === "small") {
278+ [tier, why, pinned] = ["large", `the issue is labelled ${largeLabel}`, true];
279+ } else if (largeLabel) {
280+ pinned = true;
281+ } else if (smallLabel && tier === "large" && (kind === "implement" || kind === "revise" || kind === "answer")) {
282+ [tier, why] = ["small", `the issue is labelled ${smallLabel}`];
283+ }
284+
285+ const failures = Math.max(signals.failures ?? 0, signals.retry ? 1 : 0);
286+ if (failures >= routing.frontierAfter) {
287+ return say("frontier", `the last ${plural(failures, "attempt", "attempts")} at this work failed`);
288+ }
289+ if (failures > 0) {
290+ return tier === "frontier" ? say(tier, `${why}; the last attempt failed`) : say(up(tier), "the last attempt at this work failed");
291+ }
292+ if (signals.lowConfidence) {
293+ return tier === "frontier" ? say(tier, why) : say(up(tier), "the last attempt left a change g1t was not confident in");
294+ }
295+
296+ const history = signals.history ?? [];
297+ const { window, minRuns, stepDownAt, stepUpAt } = routing.learning;
298+ const here = record(history, tier, window);
299+ if (here.runs >= minRuns && here.ok / here.runs < stepUpAt && tier !== "frontier") {
300+ const failed = here.runs - here.ok;
301+ return say(up(tier), `the ${TIER_LABEL[tier].noun} model failed ${failed} of its last ${here.runs} runs like this here`);
302+ }
303+ if (!pinned && tier !== "small") {
304+ const cheaper = record(history, down(tier), window);
305+ if (cheaper.runs >= minRuns && cheaper.ok / cheaper.runs >= stepDownAt) {
306+ return say(down(tier), `it finished ${cheaper.ok} of its last ${cheaper.runs} runs like this here`);
307+ }
308+ }
309+ return say(tier, why);
110310 }
111311
312+/** Why each kind of job starts where it does, when nothing else decides. */
313+const DEFAULT_WHY: Record<JobKind, string> = {
314+ implement: "making a change",
315+ revise: "revising a change",
316+ answer: "answering a question",
317+ review: "reviewing a change",
318+ update: "catching up with the base branch",
319+ plan: "planning work",
320+};
321+
322+/**
323+ * The tier one piece of work runs on: `route`'s tier, for callers that
324+ * need no reason.
325+ */
326+export function chooseTier(kind: JobKind, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Tier {
327+ return route(kind, signals, routing).tier;
328+}
329+
330+/** The tier a model ran as, by its public name or id; null when none of g1t's. */
331+export function tierOfModel(model: string | null | undefined, routing: AgentRouting): Tier | null {
332+ if (!model) return null;
333+ return TIERS.find((tier) => routing.tiers[tier].modelName === model || routing.tiers[tier].model === model) ?? null;
334+}
335+
112336 /** The settings that decide where model requests go. */
113337 export type ModelRouting = {
114338 /**
198422 }
199423
200424 /** A past run of the same work, as the work service lists it. */
201−export type PastAttempt = { status: string; halted?: string | null; title?: string | null };
425+export type PastAttempt = {
426+ status: string;
427+ halted?: string | null;
428+ title?: string | null;
429+ /** The pull request or issue it was for. */
430+ number?: number | null;
431+ /** The model it ran on, by its public name. */
432+ model?: string | null;
433+ confidence?: { level: string } | null;
434+};
202435
203436 /**
204437 * Whether the latest attempt at the same work failed: it failed, or g1t
212445 if (title !== undefined && (last.title ?? "").trim() !== title.trim()) return false;
213446 return last.status === "failed" || (last.status === "stopped" && Boolean(last.halted));
214447 }
448+
449+/** Whether a past run failed: it failed, or g1t stopped it at a cap of its guardrails. */
450+function failed(run: PastAttempt): boolean {
451+ return run.status === "failed" || (run.status === "stopped" && Boolean(run.halted));
452+}
453+
454+/**
455+ * How many of the latest attempts at the same work failed in a row. A
456+ * finished one, or a person stopping one, ends the count. `title` narrows
457+ * it to the same plan, as for `lastAttemptFailed`.
458+ */
459+export function failuresInARow(newestFirst: PastAttempt[], title?: string): number {
460+ let count = 0;
461+ for (const run of newestFirst) {
462+ if (title !== undefined && (run.title ?? "").trim() !== title.trim()) break;
463+ if (!failed(run)) break;
464+ count += 1;
465+ }
466+ return count;
467+}
468+
469+/** Whether the latest attempt finished but left a change g1t was not confident in. */
470+export function leftLowConfidence(newestFirst: PastAttempt[]): boolean {
471+ const last = newestFirst[0];
472+ return Boolean(last && last.status === "succeeded" && last.confidence?.level === "low");
473+}
474+
475+/**
476+ * The repository's recent runs of one kind, as learning reads them: the
477+ * tier each ran on, and whether it did the work. Runs still going say
478+ * nothing yet, and a person stopping one is not the model's failure.
479+ */
480+export function outcomesOf(newestFirst: PastAttempt[], routing: AgentRouting): PastOutcome[] {
481+ return newestFirst
482+ .filter((run) => run.status === "succeeded" || failed(run))
483+ .map((run) => ({
484+ tier: tierOfModel(run.model, routing),
485+ ok: run.status === "succeeded" && run.confidence?.level !== "low",
486+ }));
487+}
+13−9
7575 // credit. While billing takes no real money, hosted models are open
7676 // only to these workspaces; once it does, to every workspace.
7777 "HOSTED_AGENT_WORKSPACES": "flagon-io",
78− // How g1t routes the work it pays the model for. Nobody assigning an
79− // agent chooses; this is g1t's policy (AgentRouting in src/model-env.ts).
80− // "tiers": the model behind each tier; "modelName" is shown to people in
81− // the session, "model" is sent to the provider. "tasks": the tier of each
82− // kind of work; "change" reviews a change of at most "smallChange" that
83− // touches nothing sensitive, for an issue without one of "largeLabels",
84− // on the small tier, and anything else on the large. A retry after a
85− // failed attempt always runs on the large tier.
86− "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\"},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"}},\"tasks\":{\"implement\":\"large\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"small\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeLabels\":[\"security\"]}",
78+ // How g1t routes agent work: "Auto" (AgentRouting and route in
79+ // src/model-env.ts). Nobody assigning an agent has to choose; a workspace
80+ // can still choose a tier per kind of work under Integrations. "tiers": the
81+ // catalogue, the model behind small (fast), large (standard) and frontier
82+ // (most capable); "modelName" is shown to people, "model" is sent to the
83+ // provider, "price" (dollars per million tokens) is for estimates only.
84+ // "tasks": the tier each kind of job starts on, or "change" to size the
85+ // change a review reads ("smallChange" or less and nothing sensitive: small;
86+ // more than "largeChange": frontier). "frontierLabels", "largeLabels" and
87+ // "smallLabels" move work by its issue's labels. A failed attempt goes one
88+ // tier up and "frontierAfter" failures in a row to frontier; "learning"
89+ // steps work down or up by the repository's own recent runs of the kind.
90+ "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\",\"price\":{\"input\":1,\"output\":5,\"cacheRead\":0.1,\"cacheWrite\":1.25}},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\",\"price\":{\"input\":2,\"output\":10,\"cacheRead\":0.2,\"cacheWrite\":2.5}},\"frontier\":{\"modelName\":\"Claude Opus 5.5\",\"model\":\"claude-opus-5-5\",\"price\":{\"input\":4,\"output\":20,\"cacheRead\":0.2,\"cacheWrite\":5}}},\"tasks\":{\"implement\":\"large\",\"revise\":\"large\",\"answer\":\"small\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"frontier\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeChange\":{\"files\":60,\"lines\":3000},\"largeLabels\":[\"security\"],\"frontierLabels\":[\"architecture\"],\"smallLabels\":[\"documentation\",\"docs\",\"typo\"],\"frontierAfter\":2,\"learning\":{\"window\":20,\"minRuns\":5,\"stepDownAt\":0.9,\"stepUpAt\":0.5}}",
8791 // Where sandboxes send model requests, with a token for their run.
8892 // The proxy holds the keys: g1t's gateway's, or the workspace's own.
8993 "MODELS_URL": "https://models.g1t.sh",