Runner: Claude Haiku 5.5 is the fast tier, plans start there, and effort by job
Haiku 5.5 (released 2026-10-07) costs a tenth of Haiku 4.5 up to 100,000-token prompts and is the first Haiku that takes an effort level. Auto's fast tier is now Haiku 5.5, and plans start on it at high effort; a failed or unsure plan goes up a tier as any work does. `effort` in AGENT_ROUTING sets how hard each kind of job thinks (plan high, answer medium, update low; the rest the harness's default). It is sent as CLAUDE_CODE_EFFORT_LEVEL on g1t's tiers only, and the run's "Used a ... model" line names it. Sonnet 5.5's estimate price now has its 0.05x cache reads. The models guide, the routing tables and PLAN.md say so.
7 files+73−260/7 viewed
| 25 | 25 | ||
| 26 | 26 | On g1t's models you do not have to pick a model. **Auto**, the default, | |
| 27 | 27 | sends each job to the least costly model that can do it, from three tiers: | |
| 28 | − | **fast** (Claude Haiku 4.5 today), **standard** (Claude Sonnet 5.5) and | |
| 28 | + | **fast** (Claude Haiku 5.5 today), **standard** (Claude Sonnet 5.5) and | |
| 29 | 29 | **most capable** (Claude Opus 5.5). It decides by the kind of job, the | |
| 30 | 30 | size of the change it reads, the issue's labels, whether the last attempt | |
| 31 | 31 | at the same work failed, and what has worked in the repository before: | |
| 32 | 32 | ||
| 33 | − | - Catching up, answering a question and reviewing a small change that | |
| 34 | − | touches no sensitive path start on the fast model. | |
| 35 | − | - Making and revising changes, planning, and most reviews start on the | |
| 36 | − | standard model. | |
| 33 | + | - Catching up, answering a question, planning, and reviewing a small | |
| 34 | + | change that touches no sensitive path start on the fast model. | |
| 35 | + | - Making and revising changes, and most reviews, start on the standard | |
| 36 | + | model. | |
| 37 | + | - Planning runs at high effort (the model thinks longer before it | |
| 38 | + | answers), answering at medium and catching up at low, on models that | |
| 39 | + | take an effort level. | |
| 37 | 40 | - A review of a very large change, work on an issue labelled | |
| 38 | 41 | `architecture`, and work that failed twice in a row go to the most | |
| 39 | 42 | capable model. One failure moves the next attempt up one tier. | |
| 42 | 45 | keeps failing, up. | |
| 43 | 46 | ||
| 44 | 47 | Every run says which model it used and why, in one line on its run and in | |
| 45 | − | its pull request's session, such as *Used a fast model (Claude Haiku 4.5): | |
| 48 | + | its pull request's session, such as *Used a fast model (Claude Haiku 5.5): | |
| 46 | 49 | small change, 3 files and 80 lines.* The full rules are in | |
| 47 | 50 | [which model runs](/guides/working-with-g1t/#which-model-runs). | |
| 48 | 51 |
| 446 | 446 | ||
| 447 | 447 | | Tier | Model today | For | | |
| 448 | 448 | | --- | --- | --- | | |
| 449 | − | | Fast | Claude Haiku 4.5 | Small, well-bounded work | | |
| 449 | + | | Fast | Claude Haiku 5.5 | Small, well-bounded work, and plans | | |
| 450 | 450 | | Standard | Claude Sonnet 5.5 | Most changes and reviews | | |
| 451 | 451 | | Most capable | Claude Opus 5.5 | Hard work, and work that failed on the standard model | | |
| 452 | 452 | ||
| 460 | 460 | | Reviewing a pull request that changes more than 60 files or 3,000 lines | Most capable | | |
| 461 | 461 | | Reviewing any other pull request, or one whose changed files g1t does not know yet | Standard | | |
| 462 | 462 | | Catching up with the base branch and resolving conflicts | Fast | | |
| 463 | − | | Planning an outcome | Standard | | |
| 463 | + | | Planning an outcome | Fast, at high effort | | |
| 464 | 464 | ||
| 465 | 465 | Then, in this order: | |
| 466 | 466 | ||
| 486 | 486 | ||
| 487 | 487 | The agent's own small background steps run on the fast tier. | |
| 488 | 488 | ||
| 489 | + | Some work also has an effort level, how long the model thinks before it | |
| 490 | + | answers: planning runs at high effort, answering a question at medium and | |
| 491 | + | catching up at low, whichever tier the work lands on. Other work runs at | |
| 492 | + | the model's default. A workspace's own provider is sent no effort level. | |
| 493 | + | ||
| 489 | 494 | Every run says which model it used and why, in one line: as the first | |
| 490 | 495 | step on its run, and at the top of its pull request's session. For | |
| 491 | − | example, *Used a fast model (Claude Haiku 4.5): small change, 3 files and | |
| 496 | + | example, *Used a fast model (Claude Haiku 5.5): small change, 3 files and | |
| 492 | 497 | 80 lines.* An agent's review also says which model wrote it. When a better | |
| 493 | 498 | model for a tier appears, g1t changes the route and nothing you have set | |
| 494 | 499 | up needs to change. |
| 1139 | 1139 | is cost per merged change, not cost per request: a cheap attempt that | |
| 1140 | 1140 | fails and is retried on the same model costs more than one that finishes. | |
| 1141 | 1141 | ||
| 1142 | − | - **Tiers and catalogue.** `small` (Claude Haiku 4.5, $1/$5 per million | |
| 1143 | − | input/output), `large` (Claude Sonnet 5.5, $2/$10) and `frontier` | |
| 1142 | + | - **Tiers and catalogue.** `small` (Claude Haiku 5.5 since 2026-10-08, | |
| 1143 | + | $0.10/$0.50 per million input/output up to 100k-token prompts, five | |
| 1144 | + | times that above; it was Haiku 4.5 at $1/$5), `large` (Claude Sonnet 5.5, $2/$10) and `frontier` | |
| 1144 | 1145 | (Claude Opus 5.5, $4/$20). Models, names and list prices are | |
| 1145 | 1146 | configuration (`AGENT_ROUTING`), never code; prices there are for | |
| 1146 | 1147 | estimates only, runs are charged what AI Gateway priced them at. | |
| 1147 | 1148 | - **Starting tier by job.** Catch-up, answering a question, and reviews of | |
| 1148 | 1149 | at most 10 files and 200 lines touching no sensitive path: small. | |
| 1149 | − | Changes, revisions, plans and other reviews: large. Reviews over 60 files | |
| 1150 | + | Plans: small at high effort (Haiku 5.5 takes an effort level). | |
| 1151 | + | Changes, revisions and other reviews: large. Reviews over 60 files | |
| 1150 | 1152 | or 3,000 lines: frontier. Labels: `architecture` frontier, `security` off | |
| 1151 | 1153 | small, `docs`/`documentation`/`typo` let changes and answers start small. | |
| 1152 | 1154 | - **Escalation.** A failed (or guardrail-stopped) attempt at the same work | |
| 1158 | 1160 | when this tier failed at least half of at least 5. No new tables: it | |
| 1159 | 1161 | reads work's `agent_runs` (model, status, confidence). | |
| 1160 | 1162 | - **Explained.** Every run's first step and session note is one line: | |
| 1161 | − | *Used a fast model (Claude Haiku 4.5): small change, 3 files and 80 | |
| 1162 | − | lines.* | |
| 1163 | + | *Used a fast model (Claude Haiku 5.5): small change, 3 files and 80 | |
| 1164 | + | lines.* Effort per kind of job (`effort` in `AGENT_ROUTING`: plan high, | |
| 1165 | + | answer medium, update low) is sent as `CLAUDE_CODE_EFFORT_LEVEL` on | |
| 1166 | + | g1t's tiers and named in that line. | |
| 1163 | 1167 | - **Chosen instead.** `model_routes` rows to g1t's models name `small`, | |
| 1164 | 1168 | `large` or `frontier`, or nothing for Auto (Integrations → Models). | |
| 1165 | 1169 | A workspace's own Anthropic key with no model named is routed by Auto |
| 1463 | 1463 | ...(named ? { ANTHROPIC_SMALL_FAST_MODEL: named, ANTHROPIC_DEFAULT_HAIKU_MODEL: named } : {}), | |
| 1464 | 1464 | } | |
| 1465 | 1465 | : modelEnv(this.env, routing, task, tier, direct ? { ...tags, session: direct } : tags); | |
| 1466 | + | // How hard it thinks, by the kind of work, on g1t's tiers. | |
| 1467 | + | const effort = named ? undefined : routing.effort[kind]; | |
| 1468 | + | if (effort) vars.CLAUDE_CODE_EFFORT_LEVEL = effort; | |
| 1466 | 1469 | // Why this model: shown on the run and at the top of its session. | |
| 1467 | − | vars.AGENT_MODEL_REASON = reason; | |
| 1470 | + | vars.AGENT_MODEL_REASON = effort ? `${reason.replace(/\.$/, "")}, at ${effort} effort.` : reason; | |
| 1468 | 1471 | if (ticket.value) { | |
| 1469 | 1472 | // How the sandbox says what the run cost. Kept from the agent. | |
| 1470 | 1473 | vars.BILLING_RUN = ticket.value.runId; |
| 44 | 44 | ||
| 45 | 45 | const small: ChangeSize = { files: 3, lines: 80, sensitive: [] }; | |
| 46 | 46 | ||
| 47 | − | test("Auto starts each kind of job on its tier: fast for catching up and answering, standard for changes and plans", () => { | |
| 47 | + | test("Auto starts each kind of job on its tier: fast for catching up, answering and plans, standard for changes", () => { | |
| 48 | 48 | assert.equal(chooseTier("update", {}, routes), "small"); | |
| 49 | 49 | assert.equal(chooseTier("answer", {}, routes), "small"); | |
| 50 | 50 | assert.equal(chooseTier("implement", {}, routes), "large"); | |
| 51 | 51 | assert.equal(chooseTier("revise", {}, routes), "large"); | |
| 52 | 52 | assert.equal(chooseTier("implement", { change: small }, routes), "large"); | |
| 53 | − | assert.equal(chooseTier("plan", {}, routes), "large"); | |
| 53 | + | assert.equal(chooseTier("plan", {}, routes), "small"); | |
| 54 | 54 | assert.equal(chooseTier("plan", { labels: ["architecture"] }, routes), "frontier"); | |
| 55 | 55 | }); | |
| 56 | 56 | ||
| 85 | 85 | assert.equal(chooseTier("implement", { labels: ["docs"] }, routes), "small"); | |
| 86 | 86 | assert.equal(chooseTier("answer", { labels: ["typo"] }, routes), "small"); | |
| 87 | 87 | // A small label never takes a review or a plan down. | |
| 88 | − | assert.equal(chooseTier("plan", { labels: ["docs"] }, routes), "large"); | |
| 88 | + | assert.equal(chooseTier("plan", { labels: ["docs"] }, routes), "small"); | |
| 89 | + | assert.equal(chooseTier("plan", { labels: ["security"] }, routes), "large"); | |
| 89 | 90 | // Security outranks documentation. | |
| 90 | 91 | assert.equal(chooseTier("implement", { labels: ["docs", "security"] }, routes), "large"); | |
| 91 | 92 | }); | |
| 95 | 96 | assert.equal(chooseTier("review", { change: small, retry: true }, routes), "large"); | |
| 96 | 97 | assert.equal(chooseTier("implement", { failures: 1 }, routes), "frontier"); | |
| 97 | 98 | assert.equal(chooseTier("update", { failures: 2 }, routes), "frontier"); | |
| 98 | − | assert.equal(chooseTier("plan", { failures: 1 }, routes), "frontier"); | |
| 99 | + | assert.equal(chooseTier("plan", { failures: 1 }, routes), "large"); | |
| 99 | 100 | const why = route("update", { failures: 2 }, routes).reason; | |
| 100 | 101 | assert.equal(why, "Used the most capable model (Claude Opus 5.5): the last 2 attempts at this work failed."); | |
| 101 | 102 | }); | |
| 199 | 200 | assert.deepEqual(parsed.tiers.frontier, DEFAULT_ROUTING.tiers.frontier); | |
| 200 | 201 | assert.equal(chooseTier("update", {}, parsed), "large"); | |
| 201 | 202 | // A rule that names no tier keeps the default. | |
| 202 | − | assert.equal(chooseTier("plan", {}, parsed), "large"); | |
| 203 | + | assert.equal(chooseTier("plan", {}, parsed), "small"); | |
| 203 | 204 | assert.equal(parsed.tasks.answer, "change"); | |
| 204 | 205 | assert.equal(chooseTier("review", { change: { ...small, lines: 51 } }, parsed), "large"); | |
| 205 | 206 | assert.equal(chooseTier("review", { change: { ...small, files: 10, lines: 50 } }, parsed), "small"); | |
| 316 | 317 | model: "claude-sonnet-5-5", | |
| 317 | 318 | }); | |
| 318 | 319 | }); | |
| 320 | + | ||
| 321 | + | test("Each kind of job has its effort: plans think hard, answers less, catching up least", () => { | |
| 322 | + | assert.deepEqual(DEFAULT_ROUTING.effort, { plan: "high", answer: "medium", update: "low" }); | |
| 323 | + | const parsed = parseRouting(JSON.stringify({ effort: { implement: "xhigh", plan: "enormous", nonsense: "low" } })); | |
| 324 | + | assert.equal(parsed.effort.implement, "xhigh"); | |
| 325 | + | assert.equal(parsed.effort.plan, "high"); | |
| 326 | + | assert.equal((parsed.effort as Record<string, string>).nonsense, undefined); | |
| 327 | + | }); |
| 43 | 43 | price?: TokenPrice; | |
| 44 | 44 | }; | |
| 45 | 45 | ||
| 46 | + | /** | |
| 47 | + | * How hard the model thinks before it answers, on models that take it | |
| 48 | + | * (Claude Haiku 5.5 and later): more effort, more thinking tokens. | |
| 49 | + | */ | |
| 50 | + | export type Effort = "low" | "medium" | "high" | "xhigh" | "max"; | |
| 51 | + | ||
| 52 | + | export const EFFORTS: Effort[] = ["low", "medium", "high", "xhigh", "max"]; | |
| 53 | + | ||
| 46 | 54 | /** How a job's rule decides: a tier, or `change` to size the change it reads. */ | |
| 47 | 55 | export type TaskRule = Tier | "change"; | |
| 48 | 56 | ||
| 64 | 72 | tiers: Record<Tier, ModelRoute>; | |
| 65 | 73 | /** The tier each kind of job starts from, or `change` to size it. */ | |
| 66 | 74 | tasks: Record<JobKind, TaskRule>; | |
| 75 | + | /** | |
| 76 | + | * The effort each kind of job runs at, whatever tier it lands on; left | |
| 77 | + | * out, the harness's own default. Only on g1t's tiers: a route that | |
| 78 | + | * names its own model is sent as it is. | |
| 79 | + | */ | |
| 80 | + | effort: Partial<Record<JobKind, Effort>>; | |
| 67 | 81 | /** The largest change `change` sends to the small tier. */ | |
| 68 | 82 | smallChange: { files: number; lines: number }; | |
| 69 | 83 | /** A change larger than this (either) is reviewed on the frontier tier. */ | |
| 127 | 141 | /** The routing g1t ships with, for whatever the configuration leaves out. */ | |
| 128 | 142 | export const DEFAULT_ROUTING: AgentRouting = { | |
| 129 | 143 | tiers: { | |
| 144 | + | // Prompts up to 100,000 tokens; past that, five times as much. | |
| 130 | 145 | small: { | |
| 131 | − | modelName: "Claude Haiku 4.5", | |
| 132 | − | model: "claude-haiku-4-5-20251001", | |
| 133 | − | price: { input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 }, | |
| 146 | + | modelName: "Claude Haiku 5.5", | |
| 147 | + | model: "claude-haiku-5-5", | |
| 148 | + | price: { input: 0.1, output: 0.5, cacheRead: 0.01, cacheWrite: 0.125 }, | |
| 134 | 149 | }, | |
| 135 | 150 | large: { | |
| 136 | 151 | modelName: "Claude Sonnet 5.5", | |
| 137 | 152 | model: "claude-sonnet-5-5", | |
| 138 | − | price: { input: 2, output: 10, cacheRead: 0.2, cacheWrite: 2.5 }, | |
| 153 | + | price: { input: 2, output: 10, cacheRead: 0.1, cacheWrite: 2.5 }, | |
| 139 | 154 | }, | |
| 140 | 155 | frontier: { | |
| 141 | 156 | modelName: "Claude Opus 5.5", | |
| 143 | 158 | price: { input: 4, output: 20, cacheRead: 0.2, cacheWrite: 5 }, | |
| 144 | 159 | }, | |
| 145 | 160 | }, | |
| 146 | − | tasks: { implement: "large", revise: "large", answer: "small", review: "change", update: "small", plan: "large" }, | |
| 161 | + | // Plans start on the fast model thinking hard, and go up a tier when | |
| 162 | + | // one fails or leaves low confidence, as any work does. | |
| 163 | + | tasks: { implement: "large", revise: "large", answer: "small", review: "change", update: "small", plan: "small" }, | |
| 164 | + | effort: { plan: "high", answer: "medium", update: "low" }, | |
| 147 | 165 | smallChange: { files: 10, lines: 200 }, | |
| 148 | 166 | largeChange: { files: 60, lines: 3000 }, | |
| 149 | 167 | largeLabels: ["security"], | |
| 174 | 192 | for (const [kind, rule] of Object.entries(given.tasks ?? {})) { | |
| 175 | 193 | if (kind in tasks && (isTier(rule) || rule === "change")) tasks[kind as JobKind] = rule; | |
| 176 | 194 | } | |
| 195 | + | const effort = { ...DEFAULT_ROUTING.effort }; | |
| 196 | + | for (const [kind, level] of Object.entries(given.effort ?? {})) { | |
| 197 | + | if (kind in tasks && EFFORTS.includes(level as Effort)) effort[kind as JobKind] = level as Effort; | |
| 198 | + | } | |
| 177 | 199 | const tiers = { ...DEFAULT_ROUTING.tiers }; | |
| 178 | 200 | for (const tier of TIERS) { | |
| 179 | 201 | const route = given.tiers?.[tier]; | |
| 188 | 210 | return { | |
| 189 | 211 | tiers, | |
| 190 | 212 | tasks, | |
| 213 | + | effort, | |
| 191 | 214 | smallChange: { ...DEFAULT_ROUTING.smallChange, ...given.smallChange }, | |
| 192 | 215 | largeChange: { ...DEFAULT_ROUTING.largeChange, ...given.largeChange }, | |
| 193 | 216 | largeLabels: labels(given.largeLabels, DEFAULT_ROUTING.largeLabels), |
| 87 | 87 | // "smallLabels" move work by its issue's labels. A failed attempt goes one | |
| 88 | 88 | // tier up and "frontierAfter" failures in a row to frontier; "learning" | |
| 89 | 89 | // steps work down or up by the repository's own recent runs of the kind. | |
| 90 | − | "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\",\"price\":{\"input\":1,\"output\":5,\"cacheRead\":0.1,\"cacheWrite\":1.25}},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\",\"price\":{\"input\":2,\"output\":10,\"cacheRead\":0.2,\"cacheWrite\":2.5}},\"frontier\":{\"modelName\":\"Claude Opus 5.5\",\"model\":\"claude-opus-5-5\",\"price\":{\"input\":4,\"output\":20,\"cacheRead\":0.2,\"cacheWrite\":5}}},\"tasks\":{\"implement\":\"large\",\"revise\":\"large\",\"answer\":\"small\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"large\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeChange\":{\"files\":60,\"lines\":3000},\"largeLabels\":[\"security\"],\"frontierLabels\":[\"architecture\"],\"smallLabels\":[\"documentation\",\"docs\",\"typo\"],\"frontierAfter\":2,\"learning\":{\"window\":20,\"minRuns\":5,\"stepDownAt\":0.9,\"stepUpAt\":0.5}}", | |
| 90 | + | "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 5.5\",\"model\":\"claude-haiku-5-5\",\"price\":{\"input\":0.1,\"output\":0.5,\"cacheRead\":0.01,\"cacheWrite\":0.125}},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\",\"price\":{\"input\":2,\"output\":10,\"cacheRead\":0.1,\"cacheWrite\":2.5}},\"frontier\":{\"modelName\":\"Claude Opus 5.5\",\"model\":\"claude-opus-5-5\",\"price\":{\"input\":4,\"output\":20,\"cacheRead\":0.2,\"cacheWrite\":5}}},\"tasks\":{\"implement\":\"large\",\"revise\":\"large\",\"answer\":\"small\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"small\"},\"effort\":{\"plan\":\"high\",\"answer\":\"medium\",\"update\":\"low\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeChange\":{\"files\":60,\"lines\":3000},\"largeLabels\":[\"security\"],\"frontierLabels\":[\"architecture\"],\"smallLabels\":[\"documentation\",\"docs\",\"typo\"],\"frontierAfter\":2,\"learning\":{\"window\":20,\"minRuns\":5,\"stepDownAt\":0.9,\"stepUpAt\":0.5}}", | |
| 91 | 91 | // Where sandboxes send model requests, with a token for their run. | |
| 92 | 92 | // The proxy holds the keys: g1t's gateway's, or the workspace's own. | |
| 93 | 93 | "MODELS_URL": "https://models.g1t.sh", |