Runner: Auto routes each agent job to the cheapest model that can do it
The router (route in model-env.ts) now has three tiers, small (Claude Haiku 4.5), large (Claude Sonnet 5.5) and frontier (Claude Opus 5.5), and routes every kind of job on its own: catching up and answering questions start fast, changes and revisions standard, plans most capable, and a review is sized by its change (small and safe: fast; over 60 files or 3,000 lines: most capable). Issue labels move work (architecture up, security off the fast tier, docs and typos down). A failed attempt goes one tier up, two in a row to the most capable, and a change left at low confidence sends the next attempt up. A repository's own recent runs of the same kind step work down when the cheaper tier finished 9 in 10, and up when this tier failed half. The catalogue (models, list prices for estimates) and every rule are configuration (AGENT_ROUTING), tested to equal the defaults. A workspace can choose a tier per kind of work on g1t's models instead of Auto (model_routes, integrations), and a workspace's own Anthropic key with no model named is routed the same way. Each run says which model and why in one line, as its first step and at the top of its session (AGENT_MODEL_REASON). Own-provider runs keep their model session for the agent rate. The sandbox image needs rebuilding for the session line and the token report.
| 170 | 170 | /// The kinds of work a model is chosen for, and `default` for the rest. | |
| 171 | 171 | pub const MODEL_TASKS: [&str; 5] = ["default", "implement", "review", "plan", "update"]; | |
| 172 | 172 | ||
| 173 | + | /// The tiers g1t routes its hosted models' work to, cheapest first: `small` | |
| 174 | + | /// (fast), `large` (standard) and `frontier` (most capable). A route to | |
| 175 | + | /// g1t's models may name one instead of leaving the choice to Auto. | |
| 176 | + | pub const MODEL_TIERS: [&str; 3] = ["small", "large", "frontier"]; | |
| 177 | + | ||
| 173 | 178 | #[derive(Clone, Copy, Debug, PartialEq, Eq, Serialize, Deserialize)] | |
| 174 | 179 | #[serde(rename_all = "snake_case")] | |
| 175 | 180 | pub enum ProviderKind { | |
| ⋯ | |||
| 286 | 291 | /// The workspace's own model connection, or `None` for g1t's hosted | |
| 287 | 292 | /// models. | |
| 288 | 293 | pub connection_id: Option<String>, | |
| 289 | − | /// The model at that connection; its default model when `None`. | |
| 294 | + | /// The model at that connection; its default model when `None`. On | |
| 295 | + | /// g1t's hosted models, one of [`MODEL_TIERS`], or `None` for Auto. | |
| 290 | 296 | pub model: Option<String>, | |
| 291 | 297 | } | |
| 292 | 298 | ||
| ⋯ | |||
| 354 | 360 | /// The model to use instead of g1t's choice, if the connection names one. | |
| 355 | 361 | pub model: Option<String>, | |
| 356 | 362 | /// Names the run in AI Gateway's logs (`metadata.session`), so billing | |
| 357 | − | /// can charge each run what the gateway priced its requests at. Not a | |
| 358 | − | /// secret: it cannot be turned back into the token. | |
| 363 | + | /// can charge each run what the gateway priced its requests at, and its | |
| 364 | + | /// tokens in billing's count (the agent rate). Not a secret: it cannot | |
| 365 | + | /// be turned back into the token. | |
| 359 | 366 | #[serde(default)] | |
| 360 | 367 | pub id: String, | |
| 368 | + | /// The tier the workspace chose for this work on g1t's models instead | |
| 369 | + | /// of Auto: the run goes there. `None` for Auto and on its own | |
| 370 | + | /// providers. | |
| 371 | + | #[serde(default)] | |
| 372 | + | pub tier_choice: Option<String>, | |
| 361 | 373 | } | |
| 362 | 374 | ||
| 363 | 375 | /// What the model proxy needs to forward one run's requests. | |
| ⋯ | |||
| 565 | 577 | /// decides that; this service only follows the routes. | |
| 566 | 578 | #[serde(default = "yes")] | |
| 567 | 579 | pub hosted_open: bool, | |
| 568 | − | /// `small` or `large`: the tier the runner routed the run to on g1t's | |
| 569 | − | /// hosted models, tagged on its requests at the gateway. Kept only | |
| 570 | − | /// when the run goes to g1t's models. | |
| 580 | + | /// `small`, `large` or `frontier`: the tier the runner routed the run | |
| 581 | + | /// to on g1t's hosted models, tagged on its requests at the gateway. | |
| 582 | + | /// Kept only when the run goes to g1t's models, where a tier the | |
| 583 | + | /// workspace chose for the work takes its place. | |
| 571 | 584 | #[serde(default)] | |
| 572 | 585 | pub tier: Option<String>, | |
| 573 | 586 | /// The person the run is for, by username, so usage can be shown per | |
| 107 | 107 | let auth = auth_option(&env("G1T_USER")?, &env("G1T_TOKEN")?); | |
| 108 | 108 | let workdir = Path::new(WORKDIR); | |
| 109 | 109 | ||
| 110 | − | if let Ok(model) = std::env::var("AGENT_MODEL_NAME") { | |
| 110 | + | // Which model, and why: the router's one line names both. | |
| 111 | + | let reason = std::env::var("AGENT_MODEL_REASON").unwrap_or_default(); | |
| 112 | + | if !reason.trim().is_empty() { | |
| 113 | + | reporter.record(Entry::new("note", reason.trim())); | |
| 114 | + | } else if let Ok(model) = std::env::var("AGENT_MODEL_NAME") { | |
| 111 | 115 | reporter.record(Entry::new("note", &format!("Running on {model}."))); | |
| 112 | 116 | } | |
| 113 | 117 | reporter.record(Entry::new("prompt", &prompt)); |
| 117 | 117 | lastSeen: string; | |
| 118 | 118 | }; | |
| 119 | 119 | ||
| 120 | + | /** | |
| 121 | + | * How capable, and how costly, a model g1t routes to is: `small` (fast), | |
| 122 | + | * `large` (standard) or `frontier` (most capable). | |
| 123 | + | */ | |
| 124 | + | export type ModelTier = "small" | "large" | "frontier"; | |
| 125 | + | ||
| 126 | + | /** The tiers, cheapest first. */ | |
| 127 | + | export const MODEL_TIERS: ModelTier[] = ["small", "large", "frontier"]; | |
| 128 | + | ||
| 120 | 129 | export type ModelSession = { | |
| 121 | 130 | token: string; | |
| 122 | 131 | billedTo: "g1t" | "workspace"; | |
| 123 | 132 | providerName: string | null; | |
| 124 | 133 | model: string | null; | |
| 125 | − | /** Names the run in AI Gateway's logs (`metadata.session`), for billing. */ | |
| 134 | + | /** | |
| 135 | + | * Names the run in AI Gateway's logs (`metadata.session`) and its tokens | |
| 136 | + | * in billing's count, on g1t's models and the workspace's own alike. | |
| 137 | + | */ | |
| 126 | 138 | id: string; | |
| 139 | + | /** | |
| 140 | + | * The tier the workspace chose for this work on g1t's models, instead of | |
| 141 | + | * Auto; the run goes there. Null for Auto, and on its own providers. | |
| 142 | + | */ | |
| 143 | + | tierChoice?: ModelTier | null; | |
| 127 | 144 | }; | |
| 128 | 145 | ||
| 129 | 146 | export type ModelUpstream = { | |
| ⋯ | |||
| 143 | 160 | /** The session's id; see `ModelSession.id`. */ | |
| 144 | 161 | session: string; | |
| 145 | 162 | /** For `g1t`: the tier the run was routed to. */ | |
| 146 | − | tier?: "small" | "large" | null; | |
| 163 | + | tier?: ModelTier | null; | |
| 147 | 164 | /** The person the run is for, by username. Null when nobody asked; never `g1t`. */ | |
| 148 | 165 | requestedBy: string | null; | |
| 149 | 166 | baseUrl: string | null; | |
| ⋯ | |||
| 188 | 205 | task: string; | |
| 189 | 206 | hostedOpen: boolean; | |
| 190 | 207 | /** The tier the run is routed to on g1t's hosted models, for the gateway's logs. */ | |
| 191 | − | tier?: "small" | "large" | null; | |
| 208 | + | tier?: ModelTier | null; | |
| 192 | 209 | /** The person the run is for, by username, so its tokens show under them. */ | |
| 193 | 210 | requestedBy?: string | null; | |
| 194 | 211 | }): Promise<Result<ModelSession>>; | |
| 130 | 130 | (!name.is_empty() && name != g1t_contracts::identity::AGENT_NAME).then_some(name) | |
| 131 | 131 | } | |
| 132 | 132 | ||
| 133 | + | /// The tier a workspace chose on g1t's models for `task`: its own route's, | |
| 134 | + | /// else the `default` route's, when that route is to g1t's models and names | |
| 135 | + | /// one. `None` is Auto. | |
| 136 | + | fn hosted_choice(routes: &[RouteRow], task: &str) -> Option<String> { | |
| 137 | + | let chosen = routes.iter().find(|route| route.task == task).or_else(|| routes.iter().find(|route| route.task == "default"))?; | |
| 138 | + | if chosen.connection_id.is_some() { | |
| 139 | + | return None; | |
| 140 | + | } | |
| 141 | + | chosen.model.as_deref().map(str::trim).filter(|model| MODEL_TIERS.contains(model)).map(str::to_owned) | |
| 142 | + | } | |
| 143 | + | ||
| 133 | 144 | /// A model session's public id: the start of its token's hash. | |
| 134 | 145 | fn session_id(token_hash: &str) -> String { | |
| 135 | 146 | format!("ms_{}", &token_hash[..token_hash.len().min(24)]) | |
| ⋯ | |||
| 1159 | 1170 | } | |
| 1160 | 1171 | seen.push(route.task.clone()); | |
| 1161 | 1172 | let model = route.model.as_deref().map(str::trim).filter(|model| !model.is_empty()); | |
| 1173 | + | // On g1t's models a route names a tier, or nothing for Auto. | |
| 1174 | + | if route.connection_id.is_none() | |
| 1175 | + | && let Some(model) = model | |
| 1176 | + | && !MODEL_TIERS.contains(&model) | |
| 1177 | + | { | |
| 1178 | + | return Ok(fail(FailureCode::Invalid, "On g1t's models, choose Auto, fast, standard or most capable.")); | |
| 1179 | + | } | |
| 1162 | 1180 | if let Some(id) = &route.connection_id { | |
| 1163 | 1181 | let Some(row) = rows.iter().find(|row| &row.id == id && row.provider().kind() == ProviderKind::Models) else { | |
| 1164 | 1182 | return Ok(fail(FailureCode::NotFound, "A route names a model provider this workspace does not have.")); | |
| ⋯ | |||
| 1233 | 1251 | Some((row, model)) => (Some(row), model.clone()), | |
| 1234 | 1252 | None => (None, None), | |
| 1235 | 1253 | }; | |
| 1254 | + | // On g1t's models, a tier the workspace chose for this work replaces | |
| 1255 | + | // the runner's (Auto's). | |
| 1256 | + | let choice = if connection.is_none() { hosted_choice(&self.route_rows(&workspace).await?, &a.task) } else { None }; | |
| 1257 | + | let tier = choice.clone().or_else(|| a.tier.clone()); | |
| 1236 | 1258 | self.db | |
| 1237 | 1259 | .batch(vec![ | |
| 1238 | 1260 | self.db | |
| ⋯ | |||
| 1253 | 1275 | rfc3339(now + MODEL_SESSION_SECONDS * 1000).into(), | |
| 1254 | 1276 | optional(model.as_deref()), | |
| 1255 | 1277 | // The tier is g1t's routing; it means nothing on the workspace's own provider. | |
| 1256 | − | optional( | |
| 1257 | − | a.tier | |
| 1258 | − | .as_deref() | |
| 1259 | − | .filter(|tier| connection.is_none() && matches!(*tier, "small" | "large")), | |
| 1260 | − | ), | |
| 1278 | + | optional(tier.as_deref().filter(|tier| connection.is_none() && MODEL_TIERS.contains(tier))), | |
| 1261 | 1279 | optional(requester(a.requested_by.as_deref()).as_deref()), | |
| 1262 | 1280 | ])?, | |
| 1263 | 1281 | ]) | |
| ⋯ | |||
| 1269 | 1287 | provider_name: connection.map(|row| row.name.clone()), | |
| 1270 | 1288 | model, | |
| 1271 | 1289 | id, | |
| 1290 | + | tier_choice: choice, | |
| 1272 | 1291 | })) | |
| 1273 | 1292 | } | |
| 1274 | 1293 | ||
| ⋯ | |||
| 1523 | 1542 | assert_eq!(closable_hashes(&many).len(), 20); | |
| 1524 | 1543 | } | |
| 1525 | 1544 | } | |
| 1545 | + | ||
| 1546 | + | #[cfg(test)] | |
| 1547 | + | mod choice_tests { | |
| 1548 | + | use super::{RouteRow, hosted_choice}; | |
| 1549 | + | ||
| 1550 | + | fn route(task: &str, connection: Option<&str>, model: Option<&str>) -> RouteRow { | |
| 1551 | + | RouteRow { task: task.into(), connection_id: connection.map(Into::into), model: model.map(Into::into) } | |
| 1552 | + | } | |
| 1553 | + | ||
| 1554 | + | #[test] | |
| 1555 | + | fn a_tier_chosen_on_g1ts_models_is_the_works_own_or_the_defaults() { | |
| 1556 | + | let routes = vec![route("default", None, Some("large")), route("review", None, Some("frontier")), route("plan", None, None)]; | |
| 1557 | + | assert_eq!(hosted_choice(&routes, "review").as_deref(), Some("frontier")); | |
| 1558 | + | // No route of its own: the default's. | |
| 1559 | + | assert_eq!(hosted_choice(&routes, "implement").as_deref(), Some("large")); | |
| 1560 | + | // Its own route to g1t's models on Auto: Auto, not the default's. | |
| 1561 | + | assert_eq!(hosted_choice(&routes, "plan"), None); | |
| 1562 | + | // No routes at all: Auto. | |
| 1563 | + | assert_eq!(hosted_choice(&[], "implement"), None); | |
| 1564 | + | } | |
| 1565 | + | ||
| 1566 | + | #[test] | |
| 1567 | + | fn a_route_to_the_workspaces_own_provider_chooses_no_tier() { | |
| 1568 | + | let routes = vec![route("default", Some("con_1"), Some("claude-sonnet-5-5")), route("update", None, Some("huge"))]; | |
| 1569 | + | assert_eq!(hosted_choice(&routes, "implement"), None); | |
| 1570 | + | // Not a tier: Auto. | |
| 1571 | + | assert_eq!(hosted_choice(&routes, "update"), None); | |
| 1572 | + | } | |
| 1573 | + | } | |
| 12 | 12 | session: string; | |
| 13 | 13 | person: string | null; | |
| 14 | 14 | model: string; | |
| 15 | − | tier: "small" | "large" | null; | |
| 15 | + | tier: "small" | "large" | "frontier" | null; | |
| 16 | 16 | input: number; | |
| 17 | 17 | output: number; | |
| 18 | 18 | cacheRead: number; |
| 62 | 62 | } from "@g1t/contracts"; | |
| 63 | 63 | ||
| 64 | 64 | import { | |
| 65 | − | type AgentTask, | |
| 65 | + | type JobKind, | |
| 66 | + | type PastAttempt, | |
| 66 | 67 | type RouteSignals, | |
| 67 | − | type Tier, | |
| 68 | 68 | canReachModel, | |
| 69 | 69 | changeSize, | |
| 70 | − | chooseTier, | |
| 70 | + | failuresInARow, | |
| 71 | 71 | gatewaySession, | |
| 72 | − | lastAttemptFailed, | |
| 72 | + | leftLowConfidence, | |
| 73 | 73 | modelEnv, | |
| 74 | + | outcomesOf, | |
| 74 | 75 | parseRouting, | |
| 76 | + | route, | |
| 77 | + | taskOf, | |
| 75 | 78 | tierVars, | |
| 76 | 79 | } from "./model-env"; | |
| 77 | 80 | import { hubContext } from "./hub"; | |
| ⋯ | |||
| 151 | 154 | */ | |
| 152 | 155 | HOSTED_AGENT_WORKSPACES: string; | |
| 153 | 156 | /** | |
| 154 | − | * How g1t routes the work it pays the model for, as JSON (`AgentRouting` | |
| 155 | − | * in model-env.ts): `tiers`, the model behind `small` and `large`, each | |
| 156 | − | * `{ modelName, model }`; `tasks`, the tier of each kind of work, or | |
| 157 | − | * `change` to decide a review by its change; `smallChange`, the largest | |
| 158 | − | * change reviewed on the small tier; `largeLabels`, issue labels that | |
| 159 | − | * keep a review large. Anything left out takes the default. | |
| 157 | + | * How g1t routes agent work ("Auto"), as JSON (`AgentRouting` in | |
| 158 | + | * model-env.ts): `tiers`, the catalogue (the model behind `small`, | |
| 159 | + | * `large` and `frontier`, each `{ modelName, model, price }`); `tasks`, | |
| 160 | + | * the tier each kind of job starts on, or `change` to size the change; | |
| 161 | + | * `smallChange` and `largeChange`, the bounds of a small and a large | |
| 162 | + | * change; `largeLabels`, `frontierLabels` and `smallLabels`, issue | |
| 163 | + | * labels that move work; `frontierAfter`, failures in a row before the | |
| 164 | + | * frontier tier; `learning`, how a repository's own runs move it. | |
| 165 | + | * Anything left out takes the default. | |
| 160 | 166 | */ | |
| 161 | 167 | AGENT_ROUTING?: string; | |
| 162 | 168 | /** | |
| ⋯ | |||
| 191 | 197 | } | |
| 192 | 198 | ||
| 193 | 199 | /** | |
| 194 | − | * What routing knows about one piece of work, and how to ask whether it is | |
| 195 | − | * a retry (asked only when the answer matters). | |
| 200 | + | * What routing knows about one piece of work. With `viewer`, the | |
| 201 | + | * repository's recent runs of the same kind are read as them, for | |
| 202 | + | * retries, confidence and learning; `title` narrows the same work to one | |
| 203 | + | * plan's brief, since plans have no pull request. | |
| 196 | 204 | */ | |
| 197 | − | type RouteInput = RouteSignals & { retried?: () => Promise<boolean> }; | |
| 205 | + | type RouteInput = RouteSignals & { viewer?: User; title?: string }; | |
| 198 | 206 | ||
| 199 | 207 | /** A run that takes longer than this has its token expire under it. */ | |
| 200 | 208 | const TOKEN_TTL_SECONDS = 2 * 60 * 60; | |
| ⋯ | |||
| 561 | 569 | return null; | |
| 562 | 570 | } | |
| 563 | 571 | await this.ctx.storage.put("agentRun", opened.value); | |
| 572 | + | // Which model it runs on, and why, as the run's first step. | |
| 573 | + | if (envVars.AGENT_MODEL_REASON) { | |
| 574 | + | await reportRun(this.env.WORK, opened.value, { steps: [envVars.AGENT_MODEL_REASON] }).catch(() => undefined); | |
| 575 | + | } | |
| 564 | 576 | return opened.value; | |
| 565 | 577 | } | |
| 566 | 578 | ||
| ⋯ | |||
| 1354 | 1366 | } | |
| 1355 | 1367 | ||
| 1356 | 1368 | /** | |
| 1357 | − | * What a sandbox needs to reach the model routed for `task`, having | |
| 1358 | − | * opened the run the repository's workspace will be charged for. Refused | |
| 1359 | − | * when that workspace has no credit. | |
| 1369 | + | * What a sandbox needs to reach the model routed for `kind`, having | |
| 1370 | + | * opened the run the repository's workspace will be charged for. | |
| 1371 | + | * Refused when that workspace has no credit. | |
| 1360 | 1372 | * | |
| 1361 | − | * On g1t's hosted models the work goes to the cheapest tier that can do | |
| 1362 | − | * it (`chooseTier`), by what `route` says about it. Whether this is a | |
| 1363 | − | * retry is asked only when it would change the answer: when the work | |
| 1364 | − | * would otherwise go to the small tier. | |
| 1373 | + | * "Auto" (`route` in model-env.ts) picks the cheapest tier that can do | |
| 1374 | + | * the work, from what `input` says about it and the repository's own | |
| 1375 | + | * recent runs of the same kind (read once, only for a person who can see | |
| 1376 | + | * them), unless the workspace chose a tier for this work. The choice and | |
| 1377 | + | * why go to the sandbox (`AGENT_MODEL_REASON`), which records them on | |
| 1378 | + | * the run and in its session. A workspace's own Anthropic key with no | |
| 1379 | + | * model of its own named is routed the same way. | |
| 1365 | 1380 | * | |
| 1366 | 1381 | * `requestedBy` is the person the run is for, by username, so the run's | |
| 1367 | 1382 | * tokens are counted under them. | |
| 1368 | 1383 | */ | |
| 1369 | 1384 | private async modelEnv( | |
| 1370 | − | task: AgentTask, | |
| 1385 | + | kind: JobKind, | |
| 1371 | 1386 | repo: RepoPath, | |
| 1372 | 1387 | pull: number, | |
| 1373 | 1388 | requestedBy: string | null, | |
| 1374 | − | route: RouteInput = {}, | |
| 1389 | + | input: RouteInput = {}, | |
| 1375 | 1390 | ): Promise<Result<Record<string, string>>> { | |
| 1376 | 1391 | const routing = parseRouting(this.env.AGENT_ROUTING); | |
| 1377 | − | let tier: Tier = chooseTier(task, route, routing); | |
| 1378 | − | if (tier === "small" && route.retried && (await route.retried().catch(() => false))) { | |
| 1379 | − | tier = chooseTier(task, { ...route, retry: true }, routing); | |
| 1392 | + | const task = taskOf(kind); | |
| 1393 | + | const signals: RouteSignals = { ...input }; | |
| 1394 | + | if (input.viewer) { | |
| 1395 | + | // One read: the repository's recent runs of this kind, newest first. | |
| 1396 | + | // The same work's attempts are among them. | |
| 1397 | + | const recent = await agentsClient(this.env.WORK) | |
| 1398 | + | .listRuns(input.viewer, { repo, kind, limit: routing.learning.window }) | |
| 1399 | + | .catch(() => null); | |
| 1400 | + | const runs: PastAttempt[] = recent?.ok ? recent.value : []; | |
| 1401 | + | const same = input.title !== undefined ? runs : runs.filter((run) => pull > 0 && run.number === pull); | |
| 1402 | + | signals.failures = Math.max(signals.failures ?? 0, failuresInARow(same, input.title)); | |
| 1403 | + | signals.lowConfidence = signals.lowConfidence ?? leftLowConfidence(same); | |
| 1404 | + | signals.history = outcomesOf(runs, routing); | |
| 1380 | 1405 | } | |
| 1406 | + | let routed = route(kind, signals, routing); | |
| 1381 | 1407 | const tags = { repo: `${repo.namespace}/${repo.name}`, pull }; | |
| 1382 | 1408 | // Where the run's model requests go, by the workspace's routes: g1t's | |
| 1383 | 1409 | // hosted models, or one of its own providers. | |
| ⋯ | |||
| 1389 | 1415 | number: pull, | |
| 1390 | 1416 | task, | |
| 1391 | 1417 | hostedOpen: (await this.modelAccess(repo.namespace)).hosted, | |
| 1392 | − | tier, | |
| 1418 | + | tier: routed.tier, | |
| 1393 | 1419 | requestedBy, | |
| 1394 | 1420 | }); | |
| 1395 | 1421 | if (!opened.ok) return opened; | |
| 1396 | 1422 | session = opened.value; | |
| 1423 | + | // The workspace chose a tier for this work instead of Auto. | |
| 1424 | + | if (session.tierChoice) routed = route(kind, { chosen: session.tierChoice }, routing); | |
| 1397 | 1425 | } | |
| 1398 | 1426 | const own = session?.billedTo === "workspace"; | |
| 1427 | + | const tier = routed.tier; | |
| 1399 | 1428 | // Straight to the gateway, without the proxy: the run still gets a | |
| 1400 | 1429 | // session there, so billing settles it to what the gateway priced it | |
| 1401 | 1430 | // at instead of leaving the sandbox's own figure. | |
| 1402 | 1431 | const direct = !session && this.env.AI_GATEWAY_ID ? gatewaySession() : undefined; | |
| 1403 | − | // A workspace's own provider is not routed by tier: it runs the model | |
| 1404 | − | // its route names, or for an Anthropic provider, the large tier's. | |
| 1405 | − | const routed = routing.tiers[own ? "large" : tier]; | |
| 1406 | − | const model = session?.model ?? routed.model; | |
| 1407 | − | const modelName = session?.model ?? routed.modelName; | |
| 1432 | + | // A workspace's own provider runs the model its route names; with none | |
| 1433 | + | // named (an Anthropic key), the tier's, as on g1t's models. | |
| 1434 | + | const named = own && session?.model ? session.model : null; | |
| 1435 | + | const model = named ?? routing.tiers[tier].model; | |
| 1436 | + | const modelName = named ?? routing.tiers[tier].modelName; | |
| 1437 | + | const reason = named ? `Used ${named}: the workspace's route for this work names it.` : routed.reason; | |
| 1408 | 1438 | const ticket = await billingClient(this.env.BILLING).startRun({ | |
| 1409 | 1439 | workspace: repo.namespace, | |
| 1410 | 1440 | repo, | |
| ⋯ | |||
| 1412 | 1442 | task, | |
| 1413 | 1443 | model: own ? `${modelName} (${session?.providerName ?? "own provider"})` : modelName, | |
| 1414 | 1444 | billedTo: own ? "workspace" : "g1t", | |
| 1415 | − | session: own ? null : (session?.id ?? direct ?? null), | |
| 1416 | − | tier: own ? null : tier, | |
| 1445 | + | // On the workspace's own provider too: billing counts its tokens by | |
| 1446 | + | // it for the agent rate. | |
| 1447 | + | session: session?.id ?? direct ?? null, | |
| 1448 | + | tier: named ? null : tier, | |
| 1417 | 1449 | }); | |
| 1418 | 1450 | if (!ticket.ok) return ticket; | |
| 1419 | 1451 | const vars: Record<string, string> = session | |
| 1420 | 1452 | ? { | |
| 1421 | − | // On g1t's models, the tier's model, and the small tier's for the | |
| 1422 | − | // harness's own small tasks. | |
| 1423 | − | ...(own ? {} : tierVars(routing, tier)), | |
| 1453 | + | // The tier's model, and the small tier's for the harness's own | |
| 1454 | + | // small tasks; a route that names its model uses it for both. | |
| 1455 | + | ...tierVars(routing, tier), | |
| 1424 | 1456 | ANTHROPIC_MODEL: model, | |
| 1425 | 1457 | AGENT_MODEL_NAME: own ? `${modelName}, through ${session.providerName}` : modelName, | |
| 1426 | 1458 | ANTHROPIC_BASE_URL: `${this.env.MODELS_URL!.replace(/\/+$/, "")}/anthropic`, | |
| ⋯ | |||
| 1428 | 1460 | ANTHROPIC_API_KEY: session.token, | |
| 1429 | 1461 | // An endpoint that names models its own way gets its model for | |
| 1430 | 1462 | // the harness's small tasks too. | |
| 1431 | − | ...(session.model ? { ANTHROPIC_SMALL_FAST_MODEL: session.model } : {}), | |
| 1463 | + | ...(named ? { ANTHROPIC_SMALL_FAST_MODEL: named, ANTHROPIC_DEFAULT_HAIKU_MODEL: named } : {}), | |
| 1432 | 1464 | } | |
| 1433 | 1465 | : modelEnv(this.env, routing, task, tier, direct ? { ...tags, session: direct } : tags); | |
| 1466 | + | // Why this model: shown on the run and at the top of its session. | |
| 1467 | + | vars.AGENT_MODEL_REASON = reason; | |
| 1434 | 1468 | if (ticket.value) { | |
| 1435 | 1469 | // How the sandbox says what the run cost. Kept from the agent. | |
| 1436 | 1470 | vars.BILLING_RUN = ticket.value.runId; | |
| ⋯ | |||
| 1563 | 1597 | ||
| 1564 | 1598 | /** The same, for a step g1t takes by itself: a refusal stops the step. */ | |
| 1565 | 1599 | private async modelEnvOrThrow( | |
| 1566 | − | task: AgentTask, | |
| 1600 | + | task: JobKind, | |
| 1567 | 1601 | repo: RepoPath, | |
| 1568 | 1602 | pull: number, | |
| 1569 | 1603 | requestedBy: string | null, | |
| ⋯ | |||
| 1574 | 1608 | return vars.value; | |
| 1575 | 1609 | } | |
| 1576 | 1610 | ||
| 1577 | − | /** | |
| 1578 | − | * Whether the latest run of the same work failed, so that this one is a | |
| 1579 | − | * retry: the same kind of run on the same pull request, or for a plan, | |
| 1580 | − | * the latest plan with the same brief. Read as the one the run is for; | |
| 1581 | − | * unknown counts as not. | |
| 1582 | − | */ | |
| 1583 | − | private async failedBefore( | |
| 1584 | − | viewer: User, | |
| 1585 | − | repo: RepoPath, | |
| 1586 | − | kind: "review" | "update" | "plan", | |
| 1587 | − | number: number | null, | |
| 1588 | − | title?: string, | |
| 1589 | − | ): Promise<boolean> { | |
| 1590 | − | const runs = await agentsClient(this.env.WORK) | |
| 1591 | − | .listRuns(viewer, { repo, kind, ...(number != null ? { number } : {}), limit: 1 }) | |
| 1592 | − | .catch(() => null); | |
| 1593 | − | return runs?.ok ? lastAttemptFailed(runs.value, title) : false; | |
| 1594 | − | } | |
| 1595 | − | ||
| 1596 | 1611 | /** Whether sandboxes have a way to reach a model at all. */ | |
| 1597 | 1612 | private modelsReachable(): boolean { | |
| 1598 | 1613 | return Boolean(this.env.MODELS_URL) || canReachModel(this.env); | |
| ⋯ | |||
| 2266 | 2281 | job.repo, | |
| 2267 | 2282 | job.author, | |
| 2268 | 2283 | ), | |
| 2269 | − | ...(await this.modelEnvOrThrow("implement", job.repo, job.number, job.author.username)), | |
| 2284 | + | // Work handed over to the change: routed as revising it. | |
| 2285 | + | ...(await this.modelEnvOrThrow("revise", job.repo, job.number, job.author.username, { | |
| 2286 | + | labels: job.issue?.labels ?? [], | |
| 2287 | + | viewer: job.author, | |
| 2288 | + | })), | |
| 2270 | 2289 | }, | |
| 2271 | 2290 | }); | |
| 2272 | 2291 | } | |
| ⋯ | |||
| 2319 | 2338 | job.repo, | |
| 2320 | 2339 | job.author, | |
| 2321 | 2340 | ), | |
| 2322 | − | ...(await this.modelEnvOrThrow("implement", job.repo, job.number, startedBy ?? job.author.username)), | |
| 2341 | + | ...(await this.modelEnvOrThrow("revise", job.repo, job.number, startedBy ?? job.author.username, { | |
| 2342 | + | labels: job.issue?.labels ?? [], | |
| 2343 | + | viewer: job.author, | |
| 2344 | + | // The first revision is the first time the change fell short; | |
| 2345 | + | // each after it is another failure in a row. | |
| 2346 | + | failures: Math.max(0, job.round - 1), | |
| 2347 | + | })), | |
| 2323 | 2348 | }, | |
| 2324 | 2349 | }); | |
| 2325 | 2350 | } | |
| ⋯ | |||
| 2575 | 2600 | repo, | |
| 2576 | 2601 | actor, | |
| 2577 | 2602 | ), | |
| 2578 | − | ...(await this.modelEnvOrThrow("update", repo, number, actor.username, { | |
| 2579 | − | retried: () => this.failedBefore(actor, repo, "update", number), | |
| 2580 | − | })), | |
| 2603 | + | ...(await this.modelEnvOrThrow("update", repo, number, actor.username, { viewer: actor })), | |
| 2581 | 2604 | }, | |
| 2582 | 2605 | }); | |
| 2583 | 2606 | } | |
| ⋯ | |||
| 2634 | 2657 | const model = await this.modelEnv("review", repo, number, job.author.username, { | |
| 2635 | 2658 | change: job.files?.length ? changeSize(job.files, job.sensitive ?? []) : null, | |
| 2636 | 2659 | labels: job.issue?.labels ?? [], | |
| 2637 | − | retried: () => this.failedBefore(job.author, repo, "review", number), | |
| 2660 | + | viewer: job.author, | |
| 2638 | 2661 | }); | |
| 2639 | 2662 | if (!model.ok) { | |
| 2640 | 2663 | await workClient(this.env.WORK).failReview(job.runId, job.token, model.error.message); | |
| ⋯ | |||
| 2695 | 2718 | const started = await work.startPlan(actor, repo, brief); | |
| 2696 | 2719 | if (!started.ok) return started; | |
| 2697 | 2720 | const job = started.value; | |
| 2698 | − | const model = await this.modelEnv("plan", repo, 0, actor.username, { | |
| 2699 | − | retried: () => this.failedBefore(actor, repo, "plan", null, job.brief), | |
| 2700 | − | }); | |
| 2721 | + | const model = await this.modelEnv("plan", repo, 0, actor.username, { viewer: actor, title: job.brief }); | |
| 2701 | 2722 | if (!model.ok) { | |
| 2702 | 2723 | await work.failPlan(job.planId, job.token, model.error.message); | |
| 2703 | 2724 | return model; | |
| ⋯ | |||
| 2850 | 2871 | // Opened without a branch, so it has a fork. | |
| 2851 | 2872 | const fork = pull.fork!; | |
| 2852 | 2873 | ||
| 2853 | − | const model = await this.modelEnv("implement", repo, pull.number, actor.username); | |
| 2874 | + | const model = await this.modelEnv("implement", repo, pull.number, actor.username, { labels: issue.labels, viewer: actor }); | |
| 2854 | 2875 | if (!model.ok) { | |
| 2855 | 2876 | await work.closePull(actor, repo, pull.number); | |
| 2856 | 2877 | return model; | |
| ⋯ | |||
| 2997 | 3018 | ({ title, body } = found.value.issue); | |
| 2998 | 3019 | comments = found.value.comments; | |
| 2999 | 3020 | } | |
| 3000 | − | const model = await this.modelEnv("implement", job.repo, job.number, job.actor.username); | |
| 3021 | + | // A question answered from the code: it changes nothing. | |
| 3022 | + | const model = await this.modelEnv("answer", job.repo, job.number, job.actor.username, { viewer: job.actor }); | |
| 3001 | 3023 | if (!model.ok) return model; | |
| 3002 | 3024 | const source = job.pull?.source ?? job.repo; | |
| 3003 | 3025 | // Reads the code; pushes nothing. Its answer is posted with its tools. | |
| 9 | 9 | canReachModel, | |
| 10 | 10 | changeSize, | |
| 11 | 11 | chooseTier, | |
| 12 | + | failuresInARow, | |
| 12 | 13 | gatewaySession, | |
| 13 | 14 | lastAttemptFailed, | |
| 15 | + | leftLowConfidence, | |
| 14 | 16 | modelEnv, | |
| 17 | + | outcomesOf, | |
| 15 | 18 | parseRouting, | |
| 19 | + | route, | |
| 20 | + | taskOf, | |
| 21 | + | tierOfModel, | |
| 16 | 22 | } from "./model-env.ts"; | |
| 17 | 23 | ||
| 18 | 24 | const routes: AgentRouting = { | |
| ⋯ | |||
| 20 | 26 | tiers: { | |
| 21 | 27 | small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" }, | |
| 22 | 28 | large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" }, | |
| 29 | + | frontier: { modelName: "Claude Opus 5.5", model: "claude-opus-5-5" }, | |
| 23 | 30 | }, | |
| 24 | 31 | }; | |
| 25 | 32 | const tags = { repo: "acme/site", pull: 12 }; | |
| ⋯ | |||
| 37 | 44 | ||
| 38 | 45 | const small: ChangeSize = { files: 3, lines: 80, sensitive: [] }; | |
| 39 | 46 | ||
| 40 | − | test("planning and catching up run on the small tier, making a change on the large", () => { | |
| 41 | − | assert.equal(chooseTier("plan", {}, routes), "small"); | |
| 47 | + | test("Auto starts each kind of job on its tier: fast for catching up and answering, standard for changes, most capable for plans", () => { | |
| 42 | 48 | assert.equal(chooseTier("update", {}, routes), "small"); | |
| 49 | + | assert.equal(chooseTier("answer", {}, routes), "small"); | |
| 43 | 50 | assert.equal(chooseTier("implement", {}, routes), "large"); | |
| 51 | + | assert.equal(chooseTier("revise", {}, routes), "large"); | |
| 44 | 52 | assert.equal(chooseTier("implement", { change: small }, routes), "large"); | |
| 53 | + | assert.equal(chooseTier("plan", {}, routes), "frontier"); | |
| 54 | + | }); | |
| 55 | + | ||
| 56 | + | test("revising and answering go to the workspace's implement route and bill", () => { | |
| 57 | + | assert.equal(taskOf("revise"), "implement"); | |
| 58 | + | assert.equal(taskOf("answer"), "implement"); | |
| 59 | + | assert.equal(taskOf("review"), "review"); | |
| 60 | + | assert.equal(taskOf("plan"), "plan"); | |
| 45 | 61 | }); | |
| 46 | 62 | ||
| 47 | − | test("a review is small only for a small change that touches nothing sensitive", () => { | |
| 63 | + | test("a review is sized by its change: fast when small and safe, most capable when large", () => { | |
| 48 | 64 | assert.equal(chooseTier("review", { change: small }, routes), "small"); | |
| 49 | 65 | assert.equal(chooseTier("review", { change: { ...small, lines: 200, files: 10 } }, routes), "small"); | |
| 50 | 66 | assert.equal(chooseTier("review", { change: { ...small, lines: 201 } }, routes), "large"); | |
| ⋯ | |||
| 52 | 68 | assert.equal(chooseTier("review", { change: { ...small, sensitive: ["CI workflows"] } }, routes), "large"); | |
| 53 | 69 | assert.equal(chooseTier("review", { change: small, labels: ["Security"] }, routes), "large"); | |
| 54 | 70 | assert.equal(chooseTier("review", { change: small, labels: ["docs"] }, routes), "small"); | |
| 71 | + | assert.equal(chooseTier("review", { change: { ...small, files: 61 } }, routes), "frontier"); | |
| 72 | + | assert.equal(chooseTier("review", { change: { ...small, lines: 3001 } }, routes), "frontier"); | |
| 55 | 73 | }); | |
| 56 | 74 | ||
| 57 | − | test("a review of a change g1t cannot size runs on the large tier", () => { | |
| 75 | + | test("a review of a change g1t cannot size runs on the standard tier", () => { | |
| 58 | 76 | assert.equal(chooseTier("review", {}, routes), "large"); | |
| 59 | 77 | assert.equal(chooseTier("review", { change: null }, routes), "large"); | |
| 60 | 78 | assert.equal(chooseTier("review", { change: { files: 0, lines: 0, sensitive: [] } }, routes), "large"); | |
| 61 | 79 | }); | |
| 62 | 80 | ||
| 63 | − | test("a retry after a failed attempt goes up to the large tier", () => { | |
| 64 | − | assert.equal(chooseTier("plan", { retry: true }, routes), "large"); | |
| 81 | + | test("labels move work: architecture to the most capable, documentation to the fast tier", () => { | |
| 82 | + | assert.equal(chooseTier("implement", { labels: ["Architecture"] }, routes), "frontier"); | |
| 83 | + | assert.equal(chooseTier("review", { change: small, labels: ["architecture"] }, routes), "frontier"); | |
| 84 | + | assert.equal(chooseTier("implement", { labels: ["docs"] }, routes), "small"); | |
| 85 | + | assert.equal(chooseTier("answer", { labels: ["typo"] }, routes), "small"); | |
| 86 | + | // A small label never takes a review or a plan down. | |
| 87 | + | assert.equal(chooseTier("plan", { labels: ["docs"] }, routes), "frontier"); | |
| 88 | + | // Security outranks documentation. | |
| 89 | + | assert.equal(chooseTier("implement", { labels: ["docs", "security"] }, routes), "large"); | |
| 90 | + | }); | |
| 91 | + | ||
| 92 | + | test("a failed attempt goes one tier up, and repeated failures to the most capable", () => { | |
| 65 | 93 | assert.equal(chooseTier("update", { retry: true }, routes), "large"); | |
| 66 | 94 | assert.equal(chooseTier("review", { change: small, retry: true }, routes), "large"); | |
| 95 | + | assert.equal(chooseTier("implement", { failures: 1 }, routes), "frontier"); | |
| 96 | + | assert.equal(chooseTier("update", { failures: 2 }, routes), "frontier"); | |
| 97 | + | assert.equal(chooseTier("plan", { failures: 1 }, routes), "frontier"); | |
| 98 | + | const why = route("update", { failures: 2 }, routes).reason; | |
| 99 | + | assert.equal(why, "Used the most capable model (Claude Opus 5.5): the last 2 attempts at this work failed."); | |
| 100 | + | }); | |
| 101 | + | ||
| 102 | + | test("a change left at low confidence sends the next attempt one tier up", () => { | |
| 103 | + | assert.equal(chooseTier("revise", { lowConfidence: true }, routes), "frontier"); | |
| 104 | + | assert.equal(chooseTier("update", { lowConfidence: true }, routes), "large"); | |
| 105 | + | }); | |
| 106 | + | ||
| 107 | + | test("a tier the workspace chose wins over Auto", () => { | |
| 108 | + | const chosen = route("review", { change: small, failures: 3, chosen: "large" }, routes); | |
| 109 | + | assert.equal(chosen.tier, "large"); | |
| 110 | + | assert.equal(chosen.reason, "Used the standard model (Claude Sonnet 5.5): the workspace chose the standard model for this work."); | |
| 111 | + | }); | |
| 112 | + | ||
| 113 | + | const ok = (tier: "small" | "large" | "frontier", n: number) => Array.from({ length: n }, () => ({ tier, ok: true })); | |
| 114 | + | const bad = (tier: "small" | "large" | "frontier", n: number) => Array.from({ length: n }, () => ({ tier, ok: false })); | |
| 115 | + | ||
| 116 | + | test("a repository whose fast runs finish nearly always steps the work down", () => { | |
| 117 | + | const history = [...ok("small", 9), ...bad("small", 1)]; | |
| 118 | + | const routed = route("implement", { history }, routes); | |
| 119 | + | assert.equal(routed.tier, "small"); | |
| 120 | + | assert.equal(routed.reason, "Used a fast model (Claude Haiku 4.5): it finished 9 of its last 10 runs like this here."); | |
| 121 | + | // Too few runs to tell, or not quite enough of them finished: no change. | |
| 122 | + | assert.equal(chooseTier("implement", { history: ok("small", 4) }, routes), "large"); | |
| 123 | + | assert.equal(chooseTier("implement", { history: [...ok("small", 8), ...bad("small", 2)] }, routes), "large"); | |
| 124 | + | // Plans step down from the most capable model the same way. | |
| 125 | + | assert.equal(chooseTier("plan", { history: ok("large", 6) }, routes), "large"); | |
| 126 | + | }); | |
| 127 | + | ||
| 128 | + | test("learning never steps down sensitive or labelled work, nor a retry", () => { | |
| 129 | + | const history = ok("small", 10); | |
| 130 | + | assert.equal(chooseTier("review", { change: { ...small, files: 20, sensitive: ["secrets"] }, history }, routes), "large"); | |
| 131 | + | assert.equal(chooseTier("implement", { labels: ["security"], history }, routes), "large"); | |
| 132 | + | assert.equal(chooseTier("implement", { failures: 1, history }, routes), "frontier"); | |
| 133 | + | }); | |
| 134 | + | ||
| 135 | + | test("a tier that fails half its runs in a repository hands the work up", () => { | |
| 136 | + | const history = [...bad("large", 3), ...ok("large", 2)]; | |
| 137 | + | const routed = route("implement", { history }, routes); | |
| 138 | + | assert.equal(routed.tier, "frontier"); | |
| 139 | + | assert.equal(routed.reason, "Used the most capable model (Claude Opus 5.5): the standard model failed 3 of its last 5 runs like this here."); | |
| 140 | + | }); | |
| 141 | + | ||
| 142 | + | test("past runs are read by the tier they ran on and whether they did the work", () => { | |
| 143 | + | const runs = [ | |
| 144 | + | { status: "running", model: "Claude Haiku 4.5" }, | |
| 145 | + | { status: "succeeded", model: "Claude Haiku 4.5" }, | |
| 146 | + | { status: "succeeded", model: "Claude Sonnet 5.5", confidence: { level: "low" } }, | |
| 147 | + | { status: "failed", model: "Claude Opus 5.5" }, | |
| 148 | + | { status: "stopped", halted: null, model: "Claude Haiku 4.5" }, | |
| 149 | + | { status: "stopped", halted: "budget", model: "claude-haiku-4-5-20251001" }, | |
| 150 | + | { status: "succeeded", model: "Claude Sonnet 5.5, through Acme" }, | |
| 151 | + | ]; | |
| 152 | + | assert.deepEqual(outcomesOf(runs, routes), [ | |
| 153 | + | { tier: "small", ok: true }, | |
| 154 | + | { tier: "large", ok: false }, | |
| 155 | + | { tier: "frontier", ok: false }, | |
| 156 | + | { tier: "small", ok: false }, | |
| 157 | + | { tier: null, ok: true }, | |
| 158 | + | ]); | |
| 159 | + | assert.equal(tierOfModel("Claude Opus 5.5", routes), "frontier"); | |
| 160 | + | assert.equal(tierOfModel(null, routes), null); | |
| 161 | + | }); | |
| 162 | + | ||
| 163 | + | test("failures are counted in a row, newest first, and confidence is read from the last finished one", () => { | |
| 164 | + | assert.equal(failuresInARow([]), 0); | |
| 165 | + | assert.equal(failuresInARow([{ status: "failed" }, { status: "failed" }, { status: "succeeded" }, { status: "failed" }]), 2); | |
| 166 | + | assert.equal(failuresInARow([{ status: "stopped", halted: "time" }, { status: "failed" }]), 2); | |
| 167 | + | assert.equal(failuresInARow([{ status: "stopped", halted: null }, { status: "failed" }]), 0); | |
| 168 | + | assert.equal(failuresInARow([{ status: "failed", title: "Add search" }, { status: "failed", title: "Other" }], "Add search"), 1); | |
| 169 | + | assert.equal(leftLowConfidence([{ status: "succeeded", confidence: { level: "low" } }]), true); | |
| 170 | + | assert.equal(leftLowConfidence([{ status: "succeeded", confidence: { level: "medium" } }]), false); | |
| 171 | + | assert.equal(leftLowConfidence([]), false); | |
| 67 | 172 | }); | |
| 68 | 173 | ||
| 69 | 174 | test("a retry is the same work again after its latest attempt failed", () => { | |
| ⋯ | |||
| 77 | 182 | assert.equal(lastAttemptFailed([{ status: "failed", title: "Add search" }], "Add billing"), false); | |
| 78 | 183 | }); | |
| 79 | 184 | ||
| 80 | − | test("the configuration decides the tiers, the rules and the limits", () => { | |
| 185 | + | test("the configuration decides the catalogue, the rules and the limits", () => { | |
| 81 | 186 | const parsed = parseRouting( | |
| 82 | 187 | JSON.stringify({ | |
| 83 | − | tiers: { small: { modelName: "Small", model: "small-1" } }, | |
| 84 | − | tasks: { update: "large" }, | |
| 188 | + | tiers: { small: { modelName: "Small", model: "small-1" }, frontier: { model: "" } }, | |
| 189 | + | tasks: { update: "large", plan: "huge", answer: "change" }, | |
| 85 | 190 | smallChange: { lines: 50 }, | |
| 191 | + | frontierLabels: ["hard"], | |
| 192 | + | learning: { minRuns: 3 }, | |
| 86 | 193 | }), | |
| 87 | 194 | ); | |
| 88 | 195 | assert.deepEqual(parsed.tiers.small, { modelName: "Small", model: "small-1" }); | |
| 89 | 196 | assert.deepEqual(parsed.tiers.large, DEFAULT_ROUTING.tiers.large); | |
| 197 | + | // A tier with no model keeps the default. | |
| 198 | + | assert.deepEqual(parsed.tiers.frontier, DEFAULT_ROUTING.tiers.frontier); | |
| 90 | 199 | assert.equal(chooseTier("update", {}, parsed), "large"); | |
| 91 | − | assert.equal(chooseTier("plan", {}, parsed), "small"); | |
| 200 | + | // A rule that names no tier keeps the default. | |
| 201 | + | assert.equal(chooseTier("plan", {}, parsed), "frontier"); | |
| 202 | + | assert.equal(parsed.tasks.answer, "change"); | |
| 92 | 203 | assert.equal(chooseTier("review", { change: { ...small, lines: 51 } }, parsed), "large"); | |
| 93 | 204 | assert.equal(chooseTier("review", { change: { ...small, files: 10, lines: 50 } }, parsed), "small"); | |
| 205 | + | assert.equal(chooseTier("implement", { labels: ["hard"] }, parsed), "frontier"); | |
| 206 | + | assert.equal(parsed.learning.minRuns, 3); | |
| 207 | + | assert.equal(parsed.learning.window, DEFAULT_ROUTING.learning.window); | |
| 94 | 208 | assert.deepEqual(parseRouting(undefined), DEFAULT_ROUTING); | |
| 95 | 209 | assert.deepEqual(parseRouting("not json"), DEFAULT_ROUTING); | |
| 210 | + | assert.deepEqual(parseRouting("[1]"), DEFAULT_ROUTING); | |
| 211 | + | }); | |
| 212 | + | ||
| 213 | + | test("the routing the runner ships with is the default, written as configuration", async () => { | |
| 214 | + | const { readFile } = await import("node:fs/promises"); | |
| 215 | + | const config = await readFile(new URL("../wrangler.jsonc", import.meta.url), "utf8"); | |
| 216 | + | const line = config.split("\n").find((l) => l.trim().startsWith('"AGENT_ROUTING"')); | |
| 217 | + | assert.ok(line, "wrangler.jsonc sets AGENT_ROUTING"); | |
| 218 | + | const value = JSON.parse(line.trim().replace(/^"AGENT_ROUTING":\s*/, "").replace(/,$/, "")); | |
| 219 | + | assert.deepEqual(parseRouting(value), DEFAULT_ROUTING); | |
| 96 | 220 | }); | |
| 97 | 221 | ||
| 98 | 222 | test("a change's size is its files and the lines added and removed", () => { | |
| 1 | − | /** The kinds of work a g1t agent does. Each is routed on its own. */ | |
| 1 | + | /** | |
| 2 | + | * The kinds of work a g1t agent does, as model routes and billing name | |
| 3 | + | * them. A workspace routes each to a provider (Integrations → Models). | |
| 4 | + | */ | |
| 2 | 5 | export type AgentTask = "implement" | "review" | "update" | "plan"; | |
| 3 | 6 | ||
| 4 | 7 | /** | |
| 5 | − | * How capable, and how costly, a model is. g1t's hosted models come in | |
| 6 | − | * two: `small` for work a smaller model does as well, `large` for the rest. | |
| 8 | + | * What the router routes: every kind of agent job. Revising a change and | |
| 9 | + | * answering a question are routed on their own, and go to the workspace's | |
| 10 | + | * `implement` route (`taskOf`). | |
| 7 | 11 | */ | |
| 8 | − | export type Tier = "small" | "large"; | |
| 12 | + | export type JobKind = AgentTask | "revise" | "answer"; | |
| 9 | 13 | ||
| 14 | + | /** The route and the bill a job goes on. */ | |
| 15 | + | export function taskOf(kind: JobKind): AgentTask { | |
| 16 | + | return kind === "revise" || kind === "answer" ? "implement" : kind; | |
| 17 | + | } | |
| 18 | + | ||
| 19 | + | /** | |
| 20 | + | * How capable, and how costly, a model is: `small` (fast and cheap, for | |
| 21 | + | * work a smaller model does as well), `large` (the standard, most | |
| 22 | + | * changes) and `frontier` (the most capable, for hard work only). | |
| 23 | + | */ | |
| 24 | + | export type Tier = "small" | "large" | "frontier"; | |
| 25 | + | ||
| 26 | + | /** Cheapest first. */ | |
| 27 | + | export const TIERS: Tier[] = ["small", "large", "frontier"]; | |
| 28 | + | ||
| 29 | + | /** Per million tokens, in US dollars: what the provider lists. */ | |
| 30 | + | export type TokenPrice = { input: number; output: number; cacheRead: number; cacheWrite: number }; | |
| 31 | + | ||
| 10 | 32 | /** Where one kind of work goes: what people see, and what is sent. */ | |
| 11 | 33 | export type ModelRoute = { | |
| 12 | 34 | /** The model's public name, e.g. `Claude Sonnet 5.5`. */ | |
| 13 | 35 | modelName: string; | |
| 14 | 36 | /** The identifier sent to the provider. */ | |
| 15 | 37 | model: string; | |
| 38 | + | /** | |
| 39 | + | * The provider's list price, for estimates (the savings report, routing | |
| 40 | + | * by cost). Never what anyone is charged: runs are charged what AI | |
| 41 | + | * Gateway priced them at. | |
| 42 | + | */ | |
| 43 | + | price?: TokenPrice; | |
| 16 | 44 | }; | |
| 17 | 45 | ||
| 46 | + | /** How a job's rule decides: a tier, or `change` to size the change it reads. */ | |
| 47 | + | export type TaskRule = Tier | "change"; | |
| 48 | + | ||
| 49 | + | /** | |
| 50 | + | * Learning from a repository's own runs: of its last `window` runs of the | |
| 51 | + | * same kind, a cheaper tier that finished at least `stepDownAt` of at | |
| 52 | + | * least `minRuns` takes the work; a tier that finished less than | |
| 53 | + | * `stepUpAt` of at least `minRuns` hands it up. | |
| 54 | + | */ | |
| 55 | + | export type Learning = { window: number; minRuns: number; stepDownAt: number; stepUpAt: number }; | |
| 56 | + | ||
| 18 | 57 | /** | |
| 19 | − | * g1t's routing policy for the runs it pays the model for. Nobody | |
| 20 | − | * assigning an agent picks a model; the work decides, here, and the | |
| 21 | − | * operator changes it in one place (`AGENT_ROUTING` in wrangler.jsonc). | |
| 58 | + | * g1t's routing policy. Nobody assigning an agent has to pick a model: | |
| 59 | + | * "Auto" decides here, by the work, and the operator changes the policy in | |
| 60 | + | * one place (`AGENT_ROUTING` in wrangler.jsonc), never in code. | |
| 22 | 61 | */ | |
| 23 | 62 | export type AgentRouting = { | |
| 24 | − | /** The model behind each tier. */ | |
| 63 | + | /** The model behind each tier: the catalogue. */ | |
| 25 | 64 | tiers: Record<Tier, ModelRoute>; | |
| 26 | − | /** | |
| 27 | − | * The tier each kind of work runs on. `change` decides by the change the | |
| 28 | − | * work reads: small when it is small and touches nothing sensitive. | |
| 29 | − | */ | |
| 30 | − | tasks: Record<AgentTask, Tier | "change">; | |
| 65 | + | /** The tier each kind of job starts from, or `change` to size it. */ | |
| 66 | + | tasks: Record<JobKind, TaskRule>; | |
| 31 | 67 | /** The largest change `change` sends to the small tier. */ | |
| 32 | 68 | smallChange: { files: number; lines: number }; | |
| 33 | − | /** Labels on the issue behind the work that send `change` to the large tier. */ | |
| 69 | + | /** A change larger than this (either) is reviewed on the frontier tier. */ | |
| 70 | + | largeChange: { files: number; lines: number }; | |
| 71 | + | /** Issue labels that keep work off the small tier. */ | |
| 34 | 72 | largeLabels: string[]; | |
| 73 | + | /** Issue labels that send work to the frontier tier. */ | |
| 74 | + | frontierLabels: string[]; | |
| 75 | + | /** Issue labels that let changes and answers start on the small tier. */ | |
| 76 | + | smallLabels: string[]; | |
| 77 | + | /** Failed attempts at the same work, in a row, before the frontier tier. */ | |
| 78 | + | frontierAfter: number; | |
| 79 | + | learning: Learning; | |
| 35 | 80 | }; | |
| 36 | 81 | ||
| 37 | 82 | /** What a change is, as far as routing cares. */ | |
| ⋯ | |||
| 46 | 91 | sensitive: string[]; | |
| 47 | 92 | }; | |
| 48 | 93 | ||
| 94 | + | /** One past run of the same kind of job in the repository, for learning. */ | |
| 95 | + | export type PastOutcome = { | |
| 96 | + | /** The tier it ran on, when it ran on one of g1t's. */ | |
| 97 | + | tier: Tier | null; | |
| 98 | + | /** It finished, and did not leave a change g1t had low confidence in. */ | |
| 99 | + | ok: boolean; | |
| 100 | + | }; | |
| 101 | + | ||
| 49 | 102 | /** What g1t knows about one piece of work when it routes it. */ | |
| 50 | 103 | export type RouteSignals = { | |
| 51 | 104 | /** The change the work reads; null or absent when g1t does not know it. */ | |
| 52 | 105 | change?: ChangeSize | null; | |
| 53 | 106 | /** Labels on the issue the work is for. */ | |
| 54 | 107 | labels?: string[]; | |
| 55 | − | /** The last attempt at the same work failed. */ | |
| 108 | + | /** The last attempt at the same work failed. Same as `failures: 1`. */ | |
| 56 | 109 | retry?: boolean; | |
| 110 | + | /** Failed attempts at the same work, in a row, most recent last. */ | |
| 111 | + | failures?: number; | |
| 112 | + | /** The last attempt finished, but left a change g1t has low confidence in. */ | |
| 113 | + | lowConfidence?: boolean; | |
| 114 | + | /** Recent runs of the same kind in this repository, newest first. */ | |
| 115 | + | history?: PastOutcome[]; | |
| 116 | + | /** A tier the workspace chose for this work instead of Auto. */ | |
| 117 | + | chosen?: Tier | null; | |
| 118 | + | }; | |
| 119 | + | ||
| 120 | + | /** The router's answer: the tier, and why, in one line people can read. */ | |
| 121 | + | export type Routed = { | |
| 122 | + | tier: Tier; | |
| 123 | + | /** E.g. `Used a fast model (Claude Haiku 4.5): small change, 3 files and 80 lines.` */ | |
| 124 | + | reason: string; | |
| 57 | 125 | }; | |
| 58 | 126 | ||
| 59 | 127 | /** The routing g1t ships with, for whatever the configuration leaves out. */ | |
| 60 | 128 | export const DEFAULT_ROUTING: AgentRouting = { | |
| 61 | 129 | tiers: { | |
| 62 | − | small: { modelName: "Claude Haiku 4.5", model: "claude-haiku-4-5-20251001" }, | |
| 63 | − | large: { modelName: "Claude Sonnet 5.5", model: "claude-sonnet-5-5" }, | |
| 130 | + | small: { | |
| 131 | + | modelName: "Claude Haiku 4.5", | |
| 132 | + | model: "claude-haiku-4-5-20251001", | |
| 133 | + | price: { input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 }, | |
| 134 | + | }, | |
| 135 | + | large: { | |
| 136 | + | modelName: "Claude Sonnet 5.5", | |
| 137 | + | model: "claude-sonnet-5-5", | |
| 138 | + | price: { input: 2, output: 10, cacheRead: 0.2, cacheWrite: 2.5 }, | |
| 139 | + | }, | |
| 140 | + | frontier: { | |
| 141 | + | modelName: "Claude Opus 5.5", | |
| 142 | + | model: "claude-opus-5-5", | |
| 143 | + | price: { input: 4, output: 20, cacheRead: 0.2, cacheWrite: 5 }, | |
| 144 | + | }, | |
| 64 | 145 | }, | |
| 65 | − | tasks: { implement: "large", review: "change", update: "small", plan: "small" }, | |
| 146 | + | tasks: { implement: "large", revise: "large", answer: "small", review: "change", update: "small", plan: "frontier" }, | |
| 66 | 147 | smallChange: { files: 10, lines: 200 }, | |
| 148 | + | largeChange: { files: 60, lines: 3000 }, | |
| 67 | 149 | largeLabels: ["security"], | |
| 150 | + | frontierLabels: ["architecture"], | |
| 151 | + | smallLabels: ["documentation", "docs", "typo"], | |
| 152 | + | frontierAfter: 2, | |
| 153 | + | learning: { window: 20, minRuns: 5, stepDownAt: 0.9, stepUpAt: 0.5 }, | |
| 68 | 154 | }; | |
| 69 | 155 | ||
| 156 | + | function isTier(value: unknown): value is Tier { | |
| 157 | + | return value === "small" || value === "large" || value === "frontier"; | |
| 158 | + | } | |
| 159 | + | ||
| 70 | 160 | /** | |
| 71 | 161 | * The routing in `AGENT_ROUTING`, with anything it leaves out taken from | |
| 72 | − | * `DEFAULT_ROUTING`. An unset or unreadable value is the default. | |
| 162 | + | * `DEFAULT_ROUTING`. An unset or unreadable value is the default, and so is | |
| 163 | + | * any single rule that names no tier. | |
| 73 | 164 | */ | |
| 74 | 165 | export function parseRouting(json: string | undefined): AgentRouting { | |
| 75 | 166 | let given: Partial<AgentRouting> = {}; | |
| 76 | 167 | try { | |
| 77 | − | given = json ? (JSON.parse(json) as Partial<AgentRouting>) : {}; | |
| 168 | + | const parsed: unknown = json ? JSON.parse(json) : {}; | |
| 169 | + | given = parsed && typeof parsed === "object" && !Array.isArray(parsed) ? (parsed as Partial<AgentRouting>) : {}; | |
| 78 | 170 | } catch { | |
| 79 | 171 | console.log("AGENT_ROUTING is not JSON; using the default routing"); | |
| 80 | 172 | } | |
| 173 | + | const tasks = { ...DEFAULT_ROUTING.tasks }; | |
| 174 | + | for (const [kind, rule] of Object.entries(given.tasks ?? {})) { | |
| 175 | + | if (kind in tasks && (isTier(rule) || rule === "change")) tasks[kind as JobKind] = rule; | |
| 176 | + | } | |
| 177 | + | const tiers = { ...DEFAULT_ROUTING.tiers }; | |
| 178 | + | for (const tier of TIERS) { | |
| 179 | + | const route = given.tiers?.[tier]; | |
| 180 | + | if (route && typeof route.model === "string" && route.model) { | |
| 181 | + | // A model named without a price has none: estimates leave it out | |
| 182 | + | // rather than price it as another model. | |
| 183 | + | tiers[tier] = { modelName: route.modelName || route.model, model: route.model, ...(route.price ? { price: route.price } : {}) }; | |
| 184 | + | } | |
| 185 | + | } | |
| 186 | + | const labels = (list: unknown, fallback: string[]) => | |
| 187 | + | Array.isArray(list) ? list.filter((label): label is string => typeof label === "string") : fallback; | |
| 81 | 188 | return { | |
| 82 | − | tiers: { ...DEFAULT_ROUTING.tiers, ...given.tiers }, | |
| 83 | − | tasks: { ...DEFAULT_ROUTING.tasks, ...given.tasks }, | |
| 189 | + | tiers, | |
| 190 | + | tasks, | |
| 84 | 191 | smallChange: { ...DEFAULT_ROUTING.smallChange, ...given.smallChange }, | |
| 85 | − | largeLabels: given.largeLabels ?? DEFAULT_ROUTING.largeLabels, | |
| 192 | + | largeChange: { ...DEFAULT_ROUTING.largeChange, ...given.largeChange }, | |
| 193 | + | largeLabels: labels(given.largeLabels, DEFAULT_ROUTING.largeLabels), | |
| 194 | + | frontierLabels: labels(given.frontierLabels, DEFAULT_ROUTING.frontierLabels), | |
| 195 | + | smallLabels: labels(given.smallLabels, DEFAULT_ROUTING.smallLabels), | |
| 196 | + | frontierAfter: typeof given.frontierAfter === "number" && given.frontierAfter >= 1 ? given.frontierAfter : DEFAULT_ROUTING.frontierAfter, | |
| 197 | + | learning: { ...DEFAULT_ROUTING.learning, ...given.learning }, | |
| 86 | 198 | }; | |
| 87 | 199 | } | |
| 88 | 200 | ||
| 201 | + | /** How each tier is named to people. */ | |
| 202 | + | export const TIER_LABEL: Record<Tier, { noun: string; used: string }> = { | |
| 203 | + | small: { noun: "fast", used: "Used a fast model" }, | |
| 204 | + | large: { noun: "standard", used: "Used the standard model" }, | |
| 205 | + | frontier: { noun: "most capable", used: "Used the most capable model" }, | |
| 206 | + | }; | |
| 207 | + | ||
| 208 | + | const up = (tier: Tier): Tier => TIERS[Math.min(TIERS.indexOf(tier) + 1, TIERS.length - 1)]; | |
| 209 | + | const down = (tier: Tier): Tier => TIERS[Math.max(TIERS.indexOf(tier) - 1, 0)]; | |
| 210 | + | const rank = (tier: Tier) => TIERS.indexOf(tier); | |
| 211 | + | ||
| 212 | + | function plural(n: number, one: string, many: string): string { | |
| 213 | + | return `${n} ${n === 1 ? one : many}`; | |
| 214 | + | } | |
| 215 | + | ||
| 216 | + | /** How a tier did in the repository's recent runs of the same kind. */ | |
| 217 | + | export function record(history: PastOutcome[], tier: Tier, window: number): { runs: number; ok: number } { | |
| 218 | + | const recent = history.slice(0, window).filter((run) => run.tier === tier); | |
| 219 | + | return { runs: recent.length, ok: recent.filter((run) => run.ok).length }; | |
| 220 | + | } | |
| 221 | + | ||
| 89 | 222 | /** | |
| 90 | − | * The tier one piece of work runs on: the cheapest that can do it. | |
| 91 | − | * Planning and catching up are small; making a change is large; a review | |
| 92 | − | * is small for a small change that touches nothing sensitive, and large | |
| 93 | − | * for anything else, including a change g1t does not know the size of. | |
| 94 | − | * A retry after a failed attempt is always large, so that work the small | |
| 95 | − | * tier could not finish goes up rather than failing again the same way. | |
| 223 | + | * Where one job runs, and why. In order: | |
| 224 | + | * | |
| 225 | + | * 1. A tier the workspace chose for this work is used as chosen. | |
| 226 | + | * 2. The job's rule gives the starting tier: a fixed tier, or for | |
| 227 | + | * `change`, the change's size (small and touching nothing sensitive: | |
| 228 | + | * small; larger than `largeChange`: frontier; unknown or anything | |
| 229 | + | * else: large). Issue labels move it: `frontierLabels` to the frontier, | |
| 230 | + | * `largeLabels` off the small tier, `smallLabels` let a change or an | |
| 231 | + | * answer start small. | |
| 232 | + | * 3. Escalation: `frontierAfter` failures in a row go to the frontier; one | |
| 233 | + | * failure, or a last attempt that left low confidence, one tier up. | |
| 234 | + | * 4. Otherwise, learning from the repository's own runs of the same kind: | |
| 235 | + | * one tier down when the cheaper tier finished nearly all of its recent | |
| 236 | + | * ones (never for sensitive or labelled work), one tier up when this | |
| 237 | + | * tier failed half of its own. | |
| 96 | 238 | */ | |
| 97 | − | export function chooseTier(task: AgentTask, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Tier { | |
| 98 | − | if (signals.retry) return "large"; | |
| 99 | − | const rule = routing.tasks[task] ?? "large"; | |
| 100 | − | if (rule !== "change") return rule; | |
| 101 | − | const change = signals.change; | |
| 102 | − | if (!change || change.files === 0) return "large"; | |
| 103 | − | const large = new Set(routing.largeLabels.map((label) => label.toLowerCase())); | |
| 104 | − | if ((signals.labels ?? []).some((label) => large.has(label.toLowerCase()))) return "large"; | |
| 105 | − | const small = | |
| 106 | − | change.sensitive.length === 0 && | |
| 107 | − | change.files <= routing.smallChange.files && | |
| 108 | − | change.lines <= routing.smallChange.lines; | |
| 109 | − | return small ? "small" : "large"; | |
| 239 | + | export function route(kind: JobKind, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Routed { | |
| 240 | + | const say = (tier: Tier, why: string): Routed => ({ | |
| 241 | + | tier, | |
| 242 | + | reason: `${TIER_LABEL[tier].used} (${routing.tiers[tier].modelName}): ${why}.`, | |
| 243 | + | }); | |
| 244 | + | if (signals.chosen && isTier(signals.chosen)) { | |
| 245 | + | return say(signals.chosen, `the workspace chose the ${TIER_LABEL[signals.chosen].noun} model for this work`); | |
| 246 | + | } | |
| 247 | + | ||
| 248 | + | const labels = new Set((signals.labels ?? []).map((label) => label.toLowerCase())); | |
| 249 | + | const has = (list: string[]) => list.find((label) => labels.has(label.toLowerCase())); | |
| 250 | + | const rule = routing.tasks[kind] ?? "large"; | |
| 251 | + | let tier: Tier; | |
| 252 | + | let why: string; | |
| 253 | + | // Sensitive or labelled work is never stepped down by learning. | |
| 254 | + | let pinned = false; | |
| 255 | + | if (rule === "change") { | |
| 256 | + | const change = signals.change; | |
| 257 | + | if (!change || change.files === 0) { | |
| 258 | + | [tier, why] = ["large", "the change's size is not known"]; | |
| 259 | + | } else if (change.sensitive.length > 0) { | |
| 260 | + | [tier, why, pinned] = ["large", `it touches ${change.sensitive.join(", ")}`, true]; | |
| 261 | + | } else if (change.files > routing.largeChange.files || change.lines > routing.largeChange.lines) { | |
| 262 | + | [tier, why] = ["frontier", `large change, ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`]; | |
| 263 | + | } else if (change.files <= routing.smallChange.files && change.lines <= routing.smallChange.lines) { | |
| 264 | + | [tier, why] = ["small", `small change, ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`]; | |
| 265 | + | } else { | |
| 266 | + | [tier, why] = ["large", `a change of ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`]; | |
| 267 | + | } | |
| 268 | + | } else { | |
| 269 | + | tier = rule; | |
| 270 | + | why = DEFAULT_WHY[kind]; | |
| 271 | + | } | |
| 272 | + | const frontierLabel = has(routing.frontierLabels); | |
| 273 | + | const largeLabel = has(routing.largeLabels); | |
| 274 | + | const smallLabel = has(routing.smallLabels); | |
| 275 | + | if (frontierLabel) { | |
| 276 | + | [tier, why, pinned] = ["frontier", `the issue is labelled ${frontierLabel}`, true]; | |
| 277 | + | } else if (largeLabel && tier === "small") { | |
| 278 | + | [tier, why, pinned] = ["large", `the issue is labelled ${largeLabel}`, true]; | |
| 279 | + | } else if (largeLabel) { | |
| 280 | + | pinned = true; | |
| 281 | + | } else if (smallLabel && tier === "large" && (kind === "implement" || kind === "revise" || kind === "answer")) { | |
| 282 | + | [tier, why] = ["small", `the issue is labelled ${smallLabel}`]; | |
| 283 | + | } | |
| 284 | + | ||
| 285 | + | const failures = Math.max(signals.failures ?? 0, signals.retry ? 1 : 0); | |
| 286 | + | if (failures >= routing.frontierAfter) { | |
| 287 | + | return say("frontier", `the last ${plural(failures, "attempt", "attempts")} at this work failed`); | |
| 288 | + | } | |
| 289 | + | if (failures > 0) { | |
| 290 | + | return tier === "frontier" ? say(tier, `${why}; the last attempt failed`) : say(up(tier), "the last attempt at this work failed"); | |
| 291 | + | } | |
| 292 | + | if (signals.lowConfidence) { | |
| 293 | + | return tier === "frontier" ? say(tier, why) : say(up(tier), "the last attempt left a change g1t was not confident in"); | |
| 294 | + | } | |
| 295 | + | ||
| 296 | + | const history = signals.history ?? []; | |
| 297 | + | const { window, minRuns, stepDownAt, stepUpAt } = routing.learning; | |
| 298 | + | const here = record(history, tier, window); | |
| 299 | + | if (here.runs >= minRuns && here.ok / here.runs < stepUpAt && tier !== "frontier") { | |
| 300 | + | const failed = here.runs - here.ok; | |
| 301 | + | return say(up(tier), `the ${TIER_LABEL[tier].noun} model failed ${failed} of its last ${here.runs} runs like this here`); | |
| 302 | + | } | |
| 303 | + | if (!pinned && tier !== "small") { | |
| 304 | + | const cheaper = record(history, down(tier), window); | |
| 305 | + | if (cheaper.runs >= minRuns && cheaper.ok / cheaper.runs >= stepDownAt) { | |
| 306 | + | return say(down(tier), `it finished ${cheaper.ok} of its last ${cheaper.runs} runs like this here`); | |
| 307 | + | } | |
| 308 | + | } | |
| 309 | + | return say(tier, why); | |
| 110 | 310 | } | |
| 111 | 311 | ||
| 312 | + | /** Why each kind of job starts where it does, when nothing else decides. */ | |
| 313 | + | const DEFAULT_WHY: Record<JobKind, string> = { | |
| 314 | + | implement: "making a change", | |
| 315 | + | revise: "revising a change", | |
| 316 | + | answer: "answering a question", | |
| 317 | + | review: "reviewing a change", | |
| 318 | + | update: "catching up with the base branch", | |
| 319 | + | plan: "planning work", | |
| 320 | + | }; | |
| 321 | + | ||
| 322 | + | /** | |
| 323 | + | * The tier one piece of work runs on: `route`'s tier, for callers that | |
| 324 | + | * need no reason. | |
| 325 | + | */ | |
| 326 | + | export function chooseTier(kind: JobKind, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Tier { | |
| 327 | + | return route(kind, signals, routing).tier; | |
| 328 | + | } | |
| 329 | + | ||
| 330 | + | /** The tier a model ran as, by its public name or id; null when none of g1t's. */ | |
| 331 | + | export function tierOfModel(model: string | null | undefined, routing: AgentRouting): Tier | null { | |
| 332 | + | if (!model) return null; | |
| 333 | + | return TIERS.find((tier) => routing.tiers[tier].modelName === model || routing.tiers[tier].model === model) ?? null; | |
| 334 | + | } | |
| 335 | + | ||
| 112 | 336 | /** The settings that decide where model requests go. */ | |
| 113 | 337 | export type ModelRouting = { | |
| 114 | 338 | /** | |
| ⋯ | |||
| 198 | 422 | } | |
| 199 | 423 | ||
| 200 | 424 | /** A past run of the same work, as the work service lists it. */ | |
| 201 | − | export type PastAttempt = { status: string; halted?: string | null; title?: string | null }; | |
| 425 | + | export type PastAttempt = { | |
| 426 | + | status: string; | |
| 427 | + | halted?: string | null; | |
| 428 | + | title?: string | null; | |
| 429 | + | /** The pull request or issue it was for. */ | |
| 430 | + | number?: number | null; | |
| 431 | + | /** The model it ran on, by its public name. */ | |
| 432 | + | model?: string | null; | |
| 433 | + | confidence?: { level: string } | null; | |
| 434 | + | }; | |
| 202 | 435 | ||
| 203 | 436 | /** | |
| 204 | 437 | * Whether the latest attempt at the same work failed: it failed, or g1t | |
| ⋯ | |||
| 212 | 445 | if (title !== undefined && (last.title ?? "").trim() !== title.trim()) return false; | |
| 213 | 446 | return last.status === "failed" || (last.status === "stopped" && Boolean(last.halted)); | |
| 214 | 447 | } | |
| 448 | + | ||
| 449 | + | /** Whether a past run failed: it failed, or g1t stopped it at a cap of its guardrails. */ | |
| 450 | + | function failed(run: PastAttempt): boolean { | |
| 451 | + | return run.status === "failed" || (run.status === "stopped" && Boolean(run.halted)); | |
| 452 | + | } | |
| 453 | + | ||
| 454 | + | /** | |
| 455 | + | * How many of the latest attempts at the same work failed in a row. A | |
| 456 | + | * finished one, or a person stopping one, ends the count. `title` narrows | |
| 457 | + | * it to the same plan, as for `lastAttemptFailed`. | |
| 458 | + | */ | |
| 459 | + | export function failuresInARow(newestFirst: PastAttempt[], title?: string): number { | |
| 460 | + | let count = 0; | |
| 461 | + | for (const run of newestFirst) { | |
| 462 | + | if (title !== undefined && (run.title ?? "").trim() !== title.trim()) break; | |
| 463 | + | if (!failed(run)) break; | |
| 464 | + | count += 1; | |
| 465 | + | } | |
| 466 | + | return count; | |
| 467 | + | } | |
| 468 | + | ||
| 469 | + | /** Whether the latest attempt finished but left a change g1t was not confident in. */ | |
| 470 | + | export function leftLowConfidence(newestFirst: PastAttempt[]): boolean { | |
| 471 | + | const last = newestFirst[0]; | |
| 472 | + | return Boolean(last && last.status === "succeeded" && last.confidence?.level === "low"); | |
| 473 | + | } | |
| 474 | + | ||
| 475 | + | /** | |
| 476 | + | * The repository's recent runs of one kind, as learning reads them: the | |
| 477 | + | * tier each ran on, and whether it did the work. Runs still going say | |
| 478 | + | * nothing yet, and a person stopping one is not the model's failure. | |
| 479 | + | */ | |
| 480 | + | export function outcomesOf(newestFirst: PastAttempt[], routing: AgentRouting): PastOutcome[] { | |
| 481 | + | return newestFirst | |
| 482 | + | .filter((run) => run.status === "succeeded" || failed(run)) | |
| 483 | + | .map((run) => ({ | |
| 484 | + | tier: tierOfModel(run.model, routing), | |
| 485 | + | ok: run.status === "succeeded" && run.confidence?.level !== "low", | |
| 486 | + | })); | |
| 487 | + | } | |
| 75 | 75 | // credit. While billing takes no real money, hosted models are open | |
| 76 | 76 | // only to these workspaces; once it does, to every workspace. | |
| 77 | 77 | "HOSTED_AGENT_WORKSPACES": "flagon-io", | |
| 78 | − | // How g1t routes the work it pays the model for. Nobody assigning an | |
| 79 | − | // agent chooses; this is g1t's policy (AgentRouting in src/model-env.ts). | |
| 80 | − | // "tiers": the model behind each tier; "modelName" is shown to people in | |
| 81 | − | // the session, "model" is sent to the provider. "tasks": the tier of each | |
| 82 | − | // kind of work; "change" reviews a change of at most "smallChange" that | |
| 83 | − | // touches nothing sensitive, for an issue without one of "largeLabels", | |
| 84 | − | // on the small tier, and anything else on the large. A retry after a | |
| 85 | − | // failed attempt always runs on the large tier. | |
| 86 | − | "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\"},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\"}},\"tasks\":{\"implement\":\"large\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"small\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeLabels\":[\"security\"]}", | |
| 78 | + | // How g1t routes agent work: "Auto" (AgentRouting and route in | |
| 79 | + | // src/model-env.ts). Nobody assigning an agent has to choose; a workspace | |
| 80 | + | // can still choose a tier per kind of work under Integrations. "tiers": the | |
| 81 | + | // catalogue, the model behind small (fast), large (standard) and frontier | |
| 82 | + | // (most capable); "modelName" is shown to people, "model" is sent to the | |
| 83 | + | // provider, "price" (dollars per million tokens) is for estimates only. | |
| 84 | + | // "tasks": the tier each kind of job starts on, or "change" to size the | |
| 85 | + | // change a review reads ("smallChange" or less and nothing sensitive: small; | |
| 86 | + | // more than "largeChange": frontier). "frontierLabels", "largeLabels" and | |
| 87 | + | // "smallLabels" move work by its issue's labels. A failed attempt goes one | |
| 88 | + | // tier up and "frontierAfter" failures in a row to frontier; "learning" | |
| 89 | + | // steps work down or up by the repository's own recent runs of the kind. | |
| 90 | + | "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\",\"price\":{\"input\":1,\"output\":5,\"cacheRead\":0.1,\"cacheWrite\":1.25}},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\",\"price\":{\"input\":2,\"output\":10,\"cacheRead\":0.2,\"cacheWrite\":2.5}},\"frontier\":{\"modelName\":\"Claude Opus 5.5\",\"model\":\"claude-opus-5-5\",\"price\":{\"input\":4,\"output\":20,\"cacheRead\":0.2,\"cacheWrite\":5}}},\"tasks\":{\"implement\":\"large\",\"revise\":\"large\",\"answer\":\"small\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"frontier\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeChange\":{\"files\":60,\"lines\":3000},\"largeLabels\":[\"security\"],\"frontierLabels\":[\"architecture\"],\"smallLabels\":[\"documentation\",\"docs\",\"typo\"],\"frontierAfter\":2,\"learning\":{\"window\":20,\"minRuns\":5,\"stepDownAt\":0.9,\"stepUpAt\":0.5}}", | |
| 87 | 91 | // Where sandboxes send model requests, with a token for their run. | |
| 88 | 92 | // The proxy holds the keys: g1t's gateway's, or the workspace's own. | |
| 89 | 93 | "MODELS_URL": "https://models.g1t.sh", |