Skip to content

g1t/services/runner/src/model-env.ts

487 lines20,379 bytesCodeBlame

Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.

Merge branch 'model-routing'1/**
2 * The kinds of work a g1t agent does, as model routes and billing name
3 * them. A workspace routes each to a provider (Integrations → Models).
4 */
Agents as a team: lifecycle, merge queue, billing and a new shell5export type AgentTask = "implement" | "review" | "update" | "plan";
Members can read a private repository's pull request forks6
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier7/**
Merge branch 'model-routing'8 * What the router routes: every kind of agent job. Revising a change and
9 * answering a question are routed on their own, and go to the workspace's
10 * `implement` route (`taskOf`).
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier11 */
Merge branch 'model-routing'12export type JobKind = AgentTask | "revise" | "answer";
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier13
Merge branch 'model-routing'14/** The route and the bill a job goes on. */
15export function taskOf(kind: JobKind): AgentTask {
16 return kind === "revise" || kind === "answer" ? "implement" : kind;
17}
18
19/**
20 * How capable, and how costly, a model is: `small` (fast and cheap, for
21 * work a smaller model does as well), `large` (the standard, most
22 * changes) and `frontier` (the most capable, for hard work only).
23 */
24export type Tier = "small" | "large" | "frontier";
25
26/** Cheapest first. */
27export const TIERS: Tier[] = ["small", "large", "frontier"];
28
29/** Per million tokens, in US dollars: what the provider lists. */
30export type TokenPrice = { input: number; output: number; cacheRead: number; cacheWrite: number };
31
Agents as a team: lifecycle, merge queue, billing and a new shell32/** Where one kind of work goes: what people see, and what is sent. */
33export type ModelRoute = {
34 /** The model's public name, e.g. `Claude Sonnet 5.5`. */
35 modelName: string;
36 /** The identifier sent to the provider. */
37 model: string;
Merge branch 'model-routing'38 /**
39 * The provider's list price, for estimates (the savings report, routing
40 * by cost). Never what anyone is charged: runs are charged what AI
41 * Gateway priced them at.
42 */
43 price?: TokenPrice;
Agents as a team: lifecycle, merge queue, billing and a new shell44};
Members can read a private repository's pull request forks45
Merge branch 'model-routing'46/** How a job's rule decides: a tier, or `change` to size the change it reads. */
47export type TaskRule = Tier | "change";
48
49/**
50 * Learning from a repository's own runs: of its last `window` runs of the
51 * same kind, a cheaper tier that finished at least `stepDownAt` of at
52 * least `minRuns` takes the work; a tier that finished less than
53 * `stepUpAt` of at least `minRuns` hands it up.
54 */
55export type Learning = { window: number; minRuns: number; stepDownAt: number; stepUpAt: number };
56
Agents as a team: lifecycle, merge queue, billing and a new shell57/**
Merge branch 'model-routing'58 * g1t's routing policy. Nobody assigning an agent has to pick a model:
59 * "Auto" decides here, by the work, and the operator changes the policy in
60 * one place (`AGENT_ROUTING` in wrangler.jsonc), never in code.
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier61 */
62export type AgentRouting = {
Merge branch 'model-routing'63 /** The model behind each tier: the catalogue. */
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier64 tiers: Record<Tier, ModelRoute>;
Merge branch 'model-routing'65 /** The tier each kind of job starts from, or `change` to size it. */
66 tasks: Record<JobKind, TaskRule>;
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier67 /** The largest change `change` sends to the small tier. */
68 smallChange: { files: number; lines: number };
Merge branch 'model-routing'69 /** A change larger than this (either) is reviewed on the frontier tier. */
70 largeChange: { files: number; lines: number };
71 /** Issue labels that keep work off the small tier. */
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier72 largeLabels: string[];
Merge branch 'model-routing'73 /** Issue labels that send work to the frontier tier. */
74 frontierLabels: string[];
75 /** Issue labels that let changes and answers start on the small tier. */
76 smallLabels: string[];
77 /** Failed attempts at the same work, in a row, before the frontier tier. */
78 frontierAfter: number;
79 learning: Learning;
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier80};
81
82/** What a change is, as far as routing cares. */
83export type ChangeSize = {
84 files: number;
85 /** Lines added and removed. */
86 lines: number;
87 /**
88 * What it touches that runs, configures or guards things: CI, secrets,
89 * infrastructure, ownership (work's confidence.rs `sensitive`).
90 */
91 sensitive: string[];
92};
93
Merge branch 'model-routing'94/** One past run of the same kind of job in the repository, for learning. */
95export type PastOutcome = {
96 /** The tier it ran on, when it ran on one of g1t's. */
97 tier: Tier | null;
98 /** It finished, and did not leave a change g1t had low confidence in. */
99 ok: boolean;
100};
101
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier102/** What g1t knows about one piece of work when it routes it. */
103export type RouteSignals = {
104 /** The change the work reads; null or absent when g1t does not know it. */
105 change?: ChangeSize | null;
106 /** Labels on the issue the work is for. */
107 labels?: string[];
Merge branch 'model-routing'108 /** The last attempt at the same work failed. Same as `failures: 1`. */
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier109 retry?: boolean;
Merge branch 'model-routing'110 /** Failed attempts at the same work, in a row, most recent last. */
111 failures?: number;
112 /** The last attempt finished, but left a change g1t has low confidence in. */
113 lowConfidence?: boolean;
114 /** Recent runs of the same kind in this repository, newest first. */
115 history?: PastOutcome[];
116 /** A tier the workspace chose for this work instead of Auto. */
117 chosen?: Tier | null;
118};
119
120/** The router's answer: the tier, and why, in one line people can read. */
121export type Routed = {
122 tier: Tier;
123 /** E.g. `Used a fast model (Claude Haiku 4.5): small change, 3 files and 80 lines.` */
124 reason: string;
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier125};
126
127/** The routing g1t ships with, for whatever the configuration leaves out. */
128export const DEFAULT_ROUTING: AgentRouting = {
129 tiers: {
Merge branch 'model-routing'130 small: {
131 modelName: "Claude Haiku 4.5",
132 model: "claude-haiku-4-5-20251001",
133 price: { input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 },
134 },
135 large: {
136 modelName: "Claude Sonnet 5.5",
137 model: "claude-sonnet-5-5",
138 price: { input: 2, output: 10, cacheRead: 0.2, cacheWrite: 2.5 },
139 },
140 frontier: {
141 modelName: "Claude Opus 5.5",
142 model: "claude-opus-5-5",
143 price: { input: 4, output: 20, cacheRead: 0.2, cacheWrite: 5 },
144 },
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier145 },
Merge branch 'model-routing'146 tasks: { implement: "large", revise: "large", answer: "small", review: "change", update: "small", plan: "large" },
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier147 smallChange: { files: 10, lines: 200 },
Merge branch 'model-routing'148 largeChange: { files: 60, lines: 3000 },
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier149 largeLabels: ["security"],
Merge branch 'model-routing'150 frontierLabels: ["architecture"],
151 smallLabels: ["documentation", "docs", "typo"],
152 frontierAfter: 2,
153 learning: { window: 20, minRuns: 5, stepDownAt: 0.9, stepUpAt: 0.5 },
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier154};
155
Merge branch 'model-routing'156function isTier(value: unknown): value is Tier {
157 return value === "small" || value === "large" || value === "frontier";
158}
159
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier160/**
161 * The routing in `AGENT_ROUTING`, with anything it leaves out taken from
Merge branch 'model-routing'162 * `DEFAULT_ROUTING`. An unset or unreadable value is the default, and so is
163 * any single rule that names no tier.
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier164 */
165export function parseRouting(json: string | undefined): AgentRouting {
166 let given: Partial<AgentRouting> = {};
167 try {
Merge branch 'model-routing'168 const parsed: unknown = json ? JSON.parse(json) : {};
169 given = parsed && typeof parsed === "object" && !Array.isArray(parsed) ? (parsed as Partial<AgentRouting>) : {};
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier170 } catch {
171 console.log("AGENT_ROUTING is not JSON; using the default routing");
172 }
Merge branch 'model-routing'173 const tasks = { ...DEFAULT_ROUTING.tasks };
174 for (const [kind, rule] of Object.entries(given.tasks ?? {})) {
175 if (kind in tasks && (isTier(rule) || rule === "change")) tasks[kind as JobKind] = rule;
176 }
177 const tiers = { ...DEFAULT_ROUTING.tiers };
178 for (const tier of TIERS) {
179 const route = given.tiers?.[tier];
180 if (route && typeof route.model === "string" && route.model) {
181 // A model named without a price has none: estimates leave it out
182 // rather than price it as another model.
183 tiers[tier] = { modelName: route.modelName || route.model, model: route.model, ...(route.price ? { price: route.price } : {}) };
184 }
185 }
186 const labels = (list: unknown, fallback: string[]) =>
187 Array.isArray(list) ? list.filter((label): label is string => typeof label === "string") : fallback;
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier188 return {
Merge branch 'model-routing'189 tiers,
190 tasks,
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier191 smallChange: { ...DEFAULT_ROUTING.smallChange, ...given.smallChange },
Merge branch 'model-routing'192 largeChange: { ...DEFAULT_ROUTING.largeChange, ...given.largeChange },
193 largeLabels: labels(given.largeLabels, DEFAULT_ROUTING.largeLabels),
194 frontierLabels: labels(given.frontierLabels, DEFAULT_ROUTING.frontierLabels),
195 smallLabels: labels(given.smallLabels, DEFAULT_ROUTING.smallLabels),
196 frontierAfter: typeof given.frontierAfter === "number" && given.frontierAfter >= 1 ? given.frontierAfter : DEFAULT_ROUTING.frontierAfter,
197 learning: { ...DEFAULT_ROUTING.learning, ...given.learning },
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier198 };
199}
200
Merge branch 'model-routing'201/** How each tier is named to people. */
202export const TIER_LABEL: Record<Tier, { noun: string; used: string }> = {
203 small: { noun: "fast", used: "Used a fast model" },
204 large: { noun: "standard", used: "Used the standard model" },
205 frontier: { noun: "most capable", used: "Used the most capable model" },
206};
207
208const up = (tier: Tier): Tier => TIERS[Math.min(TIERS.indexOf(tier) + 1, TIERS.length - 1)];
209const down = (tier: Tier): Tier => TIERS[Math.max(TIERS.indexOf(tier) - 1, 0)];
210const rank = (tier: Tier) => TIERS.indexOf(tier);
211
212function plural(n: number, one: string, many: string): string {
213 return `${n} ${n === 1 ? one : many}`;
214}
215
216/** How a tier did in the repository's recent runs of the same kind. */
217export function record(history: PastOutcome[], tier: Tier, window: number): { runs: number; ok: number } {
218 const recent = history.slice(0, window).filter((run) => run.tier === tier);
219 return { runs: recent.length, ok: recent.filter((run) => run.ok).length };
220}
221
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier222/**
Merge branch 'model-routing'223 * Where one job runs, and why. In order:
224 *
225 * 1. A tier the workspace chose for this work is used as chosen.
226 * 2. The job's rule gives the starting tier: a fixed tier, or for
227 * `change`, the change's size (small and touching nothing sensitive:
228 * small; larger than `largeChange`: frontier; unknown or anything
229 * else: large). Issue labels move it: `frontierLabels` to the frontier,
230 * `largeLabels` off the small tier, `smallLabels` let a change or an
231 * answer start small.
232 * 3. Escalation: `frontierAfter` failures in a row go to the frontier; one
233 * failure, or a last attempt that left low confidence, one tier up.
234 * 4. Otherwise, learning from the repository's own runs of the same kind:
235 * one tier down when the cheaper tier finished nearly all of its recent
236 * ones (never for sensitive or labelled work), one tier up when this
237 * tier failed half of its own.
Agents as a team: lifecycle, merge queue, billing and a new shell238 */
Merge branch 'model-routing'239export function route(kind: JobKind, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Routed {
240 const say = (tier: Tier, why: string): Routed => ({
241 tier,
242 reason: `${TIER_LABEL[tier].used} (${routing.tiers[tier].modelName}): ${why}.`,
243 });
244 if (signals.chosen && isTier(signals.chosen)) {
245 return say(signals.chosen, `the workspace chose the ${TIER_LABEL[signals.chosen].noun} model for this work`);
246 }
247
248 const labels = new Set((signals.labels ?? []).map((label) => label.toLowerCase()));
249 const has = (list: string[]) => list.find((label) => labels.has(label.toLowerCase()));
250 const rule = routing.tasks[kind] ?? "large";
251 let tier: Tier;
252 let why: string;
253 // Sensitive or labelled work is never stepped down by learning.
254 let pinned = false;
255 if (rule === "change") {
256 const change = signals.change;
257 if (!change || change.files === 0) {
258 [tier, why] = ["large", "the change's size is not known"];
259 } else if (change.sensitive.length > 0) {
260 [tier, why, pinned] = ["large", `it touches ${change.sensitive.join(", ")}`, true];
261 } else if (change.files > routing.largeChange.files || change.lines > routing.largeChange.lines) {
262 [tier, why] = ["frontier", `large change, ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`];
263 } else if (change.files <= routing.smallChange.files && change.lines <= routing.smallChange.lines) {
264 [tier, why] = ["small", `small change, ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`];
265 } else {
266 [tier, why] = ["large", `a change of ${plural(change.files, "file", "files")} and ${plural(change.lines, "line", "lines")}`];
267 }
268 } else {
269 tier = rule;
270 why = DEFAULT_WHY[kind];
271 }
272 const frontierLabel = has(routing.frontierLabels);
273 const largeLabel = has(routing.largeLabels);
274 const smallLabel = has(routing.smallLabels);
275 if (frontierLabel) {
276 [tier, why, pinned] = ["frontier", `the issue is labelled ${frontierLabel}`, true];
277 } else if (largeLabel && tier === "small") {
278 [tier, why, pinned] = ["large", `the issue is labelled ${largeLabel}`, true];
279 } else if (largeLabel) {
280 pinned = true;
281 } else if (smallLabel && tier === "large" && (kind === "implement" || kind === "revise" || kind === "answer")) {
282 [tier, why] = ["small", `the issue is labelled ${smallLabel}`];
283 }
284
285 const failures = Math.max(signals.failures ?? 0, signals.retry ? 1 : 0);
286 if (failures >= routing.frontierAfter) {
287 return say("frontier", `the last ${plural(failures, "attempt", "attempts")} at this work failed`);
288 }
289 if (failures > 0) {
290 return tier === "frontier" ? say(tier, `${why}; the last attempt failed`) : say(up(tier), "the last attempt at this work failed");
291 }
292 if (signals.lowConfidence) {
293 return tier === "frontier" ? say(tier, why) : say(up(tier), "the last attempt left a change g1t was not confident in");
294 }
295
296 const history = signals.history ?? [];
297 const { window, minRuns, stepDownAt, stepUpAt } = routing.learning;
298 const here = record(history, tier, window);
299 if (here.runs >= minRuns && here.ok / here.runs < stepUpAt && tier !== "frontier") {
300 const failed = here.runs - here.ok;
301 return say(up(tier), `the ${TIER_LABEL[tier].noun} model failed ${failed} of its last ${here.runs} runs like this here`);
302 }
303 if (!pinned && tier !== "small") {
304 const cheaper = record(history, down(tier), window);
305 if (cheaper.runs >= minRuns && cheaper.ok / cheaper.runs >= stepDownAt) {
306 return say(down(tier), `it finished ${cheaper.ok} of its last ${cheaper.runs} runs like this here`);
307 }
308 }
309 return say(tier, why);
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier310}
Agents as a team: lifecycle, merge queue, billing and a new shell311
Merge branch 'model-routing'312/** Why each kind of job starts where it does, when nothing else decides. */
313const DEFAULT_WHY: Record<JobKind, string> = {
314 implement: "making a change",
315 revise: "revising a change",
316 answer: "answering a question",
317 review: "reviewing a change",
318 update: "catching up with the base branch",
319 plan: "planning work",
320};
321
322/**
323 * The tier one piece of work runs on: `route`'s tier, for callers that
324 * need no reason.
325 */
326export function chooseTier(kind: JobKind, signals: RouteSignals, routing: AgentRouting = DEFAULT_ROUTING): Tier {
327 return route(kind, signals, routing).tier;
328}
329
330/** The tier a model ran as, by its public name or id; null when none of g1t's. */
331export function tierOfModel(model: string | null | undefined, routing: AgentRouting): Tier | null {
332 if (!model) return null;
333 return TIERS.find((tier) => routing.tiers[tier].modelName === model || routing.tiers[tier].model === model) ?? null;
334}
335
Members can read a private repository's pull request forks336/** The settings that decide where model requests go. */
337export type ModelRouting = {
Agents as a team: lifecycle, merge queue, billing and a new shell338 /**
339 * The provider's key. Not needed when the gateway holds it and requests
340 * authenticate to the gateway instead.
341 */
Members can read a private repository's pull request forks342 ANTHROPIC_API_KEY?: string;
343 /** A Cloudflare AI Gateway id; empty sends requests to the provider directly. */
344 AI_GATEWAY_ID: string;
345 CLOUDFLARE_ACCOUNT_ID: string;
Agents as a team: lifecycle, merge queue, billing and a new shell346 /** Authenticates to the gateway, if it requires it. */
Members can read a private repository's pull request forks347 AI_GATEWAY_TOKEN?: string;
348};
349
Merge branch 'worktree-agent-a633ac0f7f66d419d'350/**
351 * What a run is for, attached to each of its requests at the gateway.
352 * `session` is the run's id there: billing finds the run's requests by it
353 * and settles the run to what the gateway priced them at.
354 */
355export type RunTags = { repo: string; pull: number; session?: string };
Agents as a team: lifecycle, merge queue, billing and a new shell356
Merge branch 'worktree-agent-a633ac0f7f66d419d'357/**
358 * A session id for a run that goes straight to the gateway (no model
359 * proxy): `rs_` and 24 hex characters, which billing's log filter needs
360 * no escaping for.
361 */
362export function gatewaySession(): string {
363 const bytes = crypto.getRandomValues(new Uint8Array(12));
364 return `rs_${Array.from(bytes, (b) => b.toString(16).padStart(2, "0")).join("")}`;
365}
366
Agents as a team: lifecycle, merge queue, billing and a new shell367/** Whether there is a way to reach a model at all. */
368export function canReachModel(env: ModelRouting): boolean {
369 return Boolean(env.ANTHROPIC_API_KEY || (env.AI_GATEWAY_ID && env.AI_GATEWAY_TOKEN));
370}
371
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier372/**
373 * The model variables of a run on g1t's hosted models: the tier's model
374 * for the work, and the small tier's for the harness's own small tasks.
375 */
376export function tierVars(routing: AgentRouting, tier: Tier): Record<string, string> {
377 const route = routing.tiers[tier];
378 return {
379 ANTHROPIC_MODEL: route.model,
380 // Recorded at the top of the session, so anyone can see what ran.
381 AGENT_MODEL_NAME: route.modelName,
382 ANTHROPIC_DEFAULT_HAIKU_MODEL: routing.tiers.small.model,
383 ANTHROPIC_SMALL_FAST_MODEL: routing.tiers.small.model,
384 };
385}
386
Members can read a private repository's pull request forks387/** Where the sandbox sends model requests, and what it sends with them. */
Agents as a team: lifecycle, merge queue, billing and a new shell388export function modelEnv(
389 env: ModelRouting,
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier390 routing: AgentRouting,
Agents as a team: lifecycle, merge queue, billing and a new shell391 task: AgentTask,
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier392 tier: Tier,
Agents as a team: lifecycle, merge queue, billing and a new shell393 tags: RunTags,
394): Record<string, string> {
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier395 const vars = tierVars(routing, tier);
Agents as a team: lifecycle, merge queue, billing and a new shell396 if (env.ANTHROPIC_API_KEY) vars.ANTHROPIC_API_KEY = env.ANTHROPIC_API_KEY;
397 if (!env.AI_GATEWAY_ID) return vars;
398
399 vars.ANTHROPIC_BASE_URL = `https://gateway.ai.cloudflare.com/v1/${env.CLOUDFLARE_ACCOUNT_ID}/${env.AI_GATEWAY_ID}/anthropic`;
400 // The gateway logs these with every request, so spend and failures can
Merge branch 'worktree-agent-a633ac0f7f66d419d'401 // be read per kind of work, tier, repository and pull request; and by
402 // the run's session, which billing settles the run's charge by.
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier403 const headers = [`cf-aig-metadata: ${JSON.stringify({ task, tier, ...tags })}`];
Agents as a team: lifecycle, merge queue, billing and a new shell404 if (env.AI_GATEWAY_TOKEN) {
405 vars.AI_GATEWAY_TOKEN = env.AI_GATEWAY_TOKEN;
406 headers.push(`cf-aig-authorization: Bearer ${env.AI_GATEWAY_TOKEN}`);
407 // With the provider's key stored in the gateway, the sandbox never
408 // holds it. The harness still wants the variable set.
409 vars.ANTHROPIC_API_KEY ??= env.AI_GATEWAY_TOKEN;
Members can read a private repository's pull request forks410 }
Agents as a team: lifecycle, merge queue, billing and a new shell411 vars.ANTHROPIC_CUSTOM_HEADERS = headers.join("\n");
Members can read a private repository's pull request forks412 return vars;
413}
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier414
415/** Lines added and removed across a change's files. */
416export function changeSize(files: { additions: number; deletions: number }[], sensitive: string[]): ChangeSize {
417 return {
418 files: files.length,
419 lines: files.reduce((sum, file) => sum + file.additions + file.deletions, 0),
420 sensitive,
421 };
422}
423
424/** A past run of the same work, as the work service lists it. */
Merge branch 'model-routing'425export type PastAttempt = {
426 status: string;
427 halted?: string | null;
428 title?: string | null;
429 /** The pull request or issue it was for. */
430 number?: number | null;
431 /** The model it ran on, by its public name. */
432 model?: string | null;
433 confidence?: { level: string } | null;
434};
Auto model routing: the cheapest tier that can do each piece of work, a retry goes up a tier, and each run records its tier435
436/**
437 * Whether the latest attempt at the same work failed: it failed, or g1t
438 * stopped it at a cap of its guardrails. A person stopping it is not a
439 * failure. `title` narrows it to the same plan, whose runs have no pull
440 * request to tell them apart.
441 */
442export function lastAttemptFailed(newestFirst: PastAttempt[], title?: string): boolean {
443 const last = newestFirst[0];
444 if (!last) return false;
445 if (title !== undefined && (last.title ?? "").trim() !== title.trim()) return false;
446 return last.status === "failed" || (last.status === "stopped" && Boolean(last.halted));
447}
Merge branch 'model-routing'448
449/** Whether a past run failed: it failed, or g1t stopped it at a cap of its guardrails. */
450function failed(run: PastAttempt): boolean {
451 return run.status === "failed" || (run.status === "stopped" && Boolean(run.halted));
452}
453
454/**
455 * How many of the latest attempts at the same work failed in a row. A
456 * finished one, or a person stopping one, ends the count. `title` narrows
457 * it to the same plan, as for `lastAttemptFailed`.
458 */
459export function failuresInARow(newestFirst: PastAttempt[], title?: string): number {
460 let count = 0;
461 for (const run of newestFirst) {
462 if (title !== undefined && (run.title ?? "").trim() !== title.trim()) break;
463 if (!failed(run)) break;
464 count += 1;
465 }
466 return count;
467}
468
469/** Whether the latest attempt finished but left a change g1t was not confident in. */
470export function leftLowConfidence(newestFirst: PastAttempt[]): boolean {
471 const last = newestFirst[0];
472 return Boolean(last && last.status === "succeeded" && last.confidence?.level === "low");
473}
474
475/**
476 * The repository's recent runs of one kind, as learning reads them: the
477 * tier each ran on, and whether it did the work. Runs still going say
478 * nothing yet, and a person stopping one is not the model's failure.
479 */
480export function outcomesOf(newestFirst: PastAttempt[], routing: AgentRouting): PastOutcome[] {
481 return newestFirst
482 .filter((run) => run.status === "succeeded" || failed(run))
483 .map((run) => ({
484 tier: tierOfModel(run.model, routing),
485 ok: run.status === "succeeded" && run.confidence?.level !== "low",
486 }));
487}

This file's history is long; its oldest lines are credited to the oldest commit read.