Skip to content

Commit

Runner: Claude Haiku 5.5 is the fast tier, plans start there, and effort by job

Haiku 5.5 (released 2026-10-07) costs a tenth of Haiku 4.5 up to 100,000-token prompts and is the first Haiku that takes an effort level. Auto's fast tier is now Haiku 5.5, and plans start on it at high effort; a failed or unsure plan goes up a tier as any work does. `effort` in AGENT_ROUTING sets how hard each kind of job thinks (plan high, answer medium, update low; the rest the harness's default). It is sent as CLAUDE_CODE_EFFORT_LEVEL on g1t's tiers only, and the run's "Used a ... model" line names it. Sonnet 5.5's estimate price now has its 0.05x cache reads. The models guide, the routing tables and PLAN.md say so.

syntaqxcommitted Parent7650f59Browse files
7 files+73−260/7 viewed
+9−6
2525
2626 On g1t's models you do not have to pick a model. **Auto**, the default,
2727 sends each job to the least costly model that can do it, from three tiers:
28−**fast** (Claude Haiku 4.5 today), **standard** (Claude Sonnet 5.5) and
28+**fast** (Claude Haiku 5.5 today), **standard** (Claude Sonnet 5.5) and
2929 **most capable** (Claude Opus 5.5). It decides by the kind of job, the
3030 size of the change it reads, the issue's labels, whether the last attempt
3131 at the same work failed, and what has worked in the repository before:
3232
33−- Catching up, answering a question and reviewing a small change that
34− touches no sensitive path start on the fast model.
35−- Making and revising changes, planning, and most reviews start on the
36− standard model.
33+- Catching up, answering a question, planning, and reviewing a small
34+ change that touches no sensitive path start on the fast model.
35+- Making and revising changes, and most reviews, start on the standard
36+ model.
37+- Planning runs at high effort (the model thinks longer before it
38+ answers), answering at medium and catching up at low, on models that
39+ take an effort level.
3740 - A review of a very large change, work on an issue labelled
3841 `architecture`, and work that failed twice in a row go to the most
3942 capable model. One failure moves the next attempt up one tier.
4245 keeps failing, up.
4346
4447 Every run says which model it used and why, in one line on its run and in
45−its pull request's session, such as *Used a fast model (Claude Haiku 4.5):
48+its pull request's session, such as *Used a fast model (Claude Haiku 5.5):
4649 small change, 3 files and 80 lines.* The full rules are in
4750 [which model runs](/guides/working-with-g1t/#which-model-runs).
4851
+8−3
446446
447447 | Tier | Model today | For |
448448 | --- | --- | --- |
449−| Fast | Claude Haiku 4.5 | Small, well-bounded work |
449+| Fast | Claude Haiku 5.5 | Small, well-bounded work, and plans |
450450 | Standard | Claude Sonnet 5.5 | Most changes and reviews |
451451 | Most capable | Claude Opus 5.5 | Hard work, and work that failed on the standard model |
452452
460460 | Reviewing a pull request that changes more than 60 files or 3,000 lines | Most capable |
461461 | Reviewing any other pull request, or one whose changed files g1t does not know yet | Standard |
462462 | Catching up with the base branch and resolving conflicts | Fast |
463−| Planning an outcome | Standard |
463+| Planning an outcome | Fast, at high effort |
464464
465465 Then, in this order:
466466
486486
487487 The agent's own small background steps run on the fast tier.
488488
489+Some work also has an effort level, how long the model thinks before it
490+answers: planning runs at high effort, answering a question at medium and
491+catching up at low, whichever tier the work lands on. Other work runs at
492+the model's default. A workspace's own provider is sent no effort level.
493+
489494 Every run says which model it used and why, in one line: as the first
490495 step on its run, and at the top of its pull request's session. For
491−example, *Used a fast model (Claude Haiku 4.5): small change, 3 files and
496+example, *Used a fast model (Claude Haiku 5.5): small change, 3 files and
492497 80 lines.* An agent's review also says which model wrote it. When a better
493498 model for a tier appears, g1t changes the route and nothing you have set
494499 up needs to change.
+9−5
11391139 is cost per merged change, not cost per request: a cheap attempt that
11401140 fails and is retried on the same model costs more than one that finishes.
11411141
1142−- **Tiers and catalogue.** `small` (Claude Haiku 4.5, $1/$5 per million
1143− input/output), `large` (Claude Sonnet 5.5, $2/$10) and `frontier`
1142+- **Tiers and catalogue.** `small` (Claude Haiku 5.5 since 2026-10-08,
1143+ $0.10/$0.50 per million input/output up to 100k-token prompts, five
1144+ times that above; it was Haiku 4.5 at $1/$5), `large` (Claude Sonnet 5.5, $2/$10) and `frontier`
11441145 (Claude Opus 5.5, $4/$20). Models, names and list prices are
11451146 configuration (`AGENT_ROUTING`), never code; prices there are for
11461147 estimates only, runs are charged what AI Gateway priced them at.
11471148 - **Starting tier by job.** Catch-up, answering a question, and reviews of
11481149 at most 10 files and 200 lines touching no sensitive path: small.
1149− Changes, revisions, plans and other reviews: large. Reviews over 60 files
1150+ Plans: small at high effort (Haiku 5.5 takes an effort level).
1151+ Changes, revisions and other reviews: large. Reviews over 60 files
11501152 or 3,000 lines: frontier. Labels: `architecture` frontier, `security` off
11511153 small, `docs`/`documentation`/`typo` let changes and answers start small.
11521154 - **Escalation.** A failed (or guardrail-stopped) attempt at the same work
11581160 when this tier failed at least half of at least 5. No new tables: it
11591161 reads work's `agent_runs` (model, status, confidence).
11601162 - **Explained.** Every run's first step and session note is one line:
1161− *Used a fast model (Claude Haiku 4.5): small change, 3 files and 80
1162− lines.*
1163+ *Used a fast model (Claude Haiku 5.5): small change, 3 files and 80
1164+ lines.* Effort per kind of job (`effort` in `AGENT_ROUTING`: plan high,
1165+ answer medium, update low) is sent as `CLAUDE_CODE_EFFORT_LEVEL` on
1166+ g1t's tiers and named in that line.
11631167 - **Chosen instead.** `model_routes` rows to g1t's models name `small`,
11641168 `large` or `frontier`, or nothing for Auto (Integrations → Models).
11651169 A workspace's own Anthropic key with no model named is routed by Auto
+4−1
14631463 ...(named ? { ANTHROPIC_SMALL_FAST_MODEL: named, ANTHROPIC_DEFAULT_HAIKU_MODEL: named } : {}),
14641464 }
14651465 : modelEnv(this.env, routing, task, tier, direct ? { ...tags, session: direct } : tags);
1466+ // How hard it thinks, by the kind of work, on g1t's tiers.
1467+ const effort = named ? undefined : routing.effort[kind];
1468+ if (effort) vars.CLAUDE_CODE_EFFORT_LEVEL = effort;
14661469 // Why this model: shown on the run and at the top of its session.
1467− vars.AGENT_MODEL_REASON = reason;
1470+ vars.AGENT_MODEL_REASON = effort ? `${reason.replace(/\.$/, "")}, at ${effort} effort.` : reason;
14681471 if (ticket.value) {
14691472 // How the sandbox says what the run cost. Kept from the agent.
14701473 vars.BILLING_RUN = ticket.value.runId;
+14−5
4444
4545 const small: ChangeSize = { files: 3, lines: 80, sensitive: [] };
4646
47−test("Auto starts each kind of job on its tier: fast for catching up and answering, standard for changes and plans", () => {
47+test("Auto starts each kind of job on its tier: fast for catching up, answering and plans, standard for changes", () => {
4848 assert.equal(chooseTier("update", {}, routes), "small");
4949 assert.equal(chooseTier("answer", {}, routes), "small");
5050 assert.equal(chooseTier("implement", {}, routes), "large");
5151 assert.equal(chooseTier("revise", {}, routes), "large");
5252 assert.equal(chooseTier("implement", { change: small }, routes), "large");
53− assert.equal(chooseTier("plan", {}, routes), "large");
53+ assert.equal(chooseTier("plan", {}, routes), "small");
5454 assert.equal(chooseTier("plan", { labels: ["architecture"] }, routes), "frontier");
5555 });
5656
8585 assert.equal(chooseTier("implement", { labels: ["docs"] }, routes), "small");
8686 assert.equal(chooseTier("answer", { labels: ["typo"] }, routes), "small");
8787 // A small label never takes a review or a plan down.
88− assert.equal(chooseTier("plan", { labels: ["docs"] }, routes), "large");
88+ assert.equal(chooseTier("plan", { labels: ["docs"] }, routes), "small");
89+ assert.equal(chooseTier("plan", { labels: ["security"] }, routes), "large");
8990 // Security outranks documentation.
9091 assert.equal(chooseTier("implement", { labels: ["docs", "security"] }, routes), "large");
9192 });
9596 assert.equal(chooseTier("review", { change: small, retry: true }, routes), "large");
9697 assert.equal(chooseTier("implement", { failures: 1 }, routes), "frontier");
9798 assert.equal(chooseTier("update", { failures: 2 }, routes), "frontier");
98− assert.equal(chooseTier("plan", { failures: 1 }, routes), "frontier");
99+ assert.equal(chooseTier("plan", { failures: 1 }, routes), "large");
99100 const why = route("update", { failures: 2 }, routes).reason;
100101 assert.equal(why, "Used the most capable model (Claude Opus 5.5): the last 2 attempts at this work failed.");
101102 });
199200 assert.deepEqual(parsed.tiers.frontier, DEFAULT_ROUTING.tiers.frontier);
200201 assert.equal(chooseTier("update", {}, parsed), "large");
201202 // A rule that names no tier keeps the default.
202− assert.equal(chooseTier("plan", {}, parsed), "large");
203+ assert.equal(chooseTier("plan", {}, parsed), "small");
203204 assert.equal(parsed.tasks.answer, "change");
204205 assert.equal(chooseTier("review", { change: { ...small, lines: 51 } }, parsed), "large");
205206 assert.equal(chooseTier("review", { change: { ...small, files: 10, lines: 50 } }, parsed), "small");
316317 model: "claude-sonnet-5-5",
317318 });
318319 });
320+
321+test("Each kind of job has its effort: plans think hard, answers less, catching up least", () => {
322+ assert.deepEqual(DEFAULT_ROUTING.effort, { plan: "high", answer: "medium", update: "low" });
323+ const parsed = parseRouting(JSON.stringify({ effort: { implement: "xhigh", plan: "enormous", nonsense: "low" } }));
324+ assert.equal(parsed.effort.implement, "xhigh");
325+ assert.equal(parsed.effort.plan, "high");
326+ assert.equal((parsed.effort as Record<string, string>).nonsense, undefined);
327+});
+28−5
4343 price?: TokenPrice;
4444 };
4545
46+/**
47+ * How hard the model thinks before it answers, on models that take it
48+ * (Claude Haiku 5.5 and later): more effort, more thinking tokens.
49+ */
50+export type Effort = "low" | "medium" | "high" | "xhigh" | "max";
51+
52+export const EFFORTS: Effort[] = ["low", "medium", "high", "xhigh", "max"];
53+
4654 /** How a job's rule decides: a tier, or `change` to size the change it reads. */
4755 export type TaskRule = Tier | "change";
4856
6472 tiers: Record<Tier, ModelRoute>;
6573 /** The tier each kind of job starts from, or `change` to size it. */
6674 tasks: Record<JobKind, TaskRule>;
75+ /**
76+ * The effort each kind of job runs at, whatever tier it lands on; left
77+ * out, the harness's own default. Only on g1t's tiers: a route that
78+ * names its own model is sent as it is.
79+ */
80+ effort: Partial<Record<JobKind, Effort>>;
6781 /** The largest change `change` sends to the small tier. */
6882 smallChange: { files: number; lines: number };
6983 /** A change larger than this (either) is reviewed on the frontier tier. */
127141 /** The routing g1t ships with, for whatever the configuration leaves out. */
128142 export const DEFAULT_ROUTING: AgentRouting = {
129143 tiers: {
144+ // Prompts up to 100,000 tokens; past that, five times as much.
130145 small: {
131− modelName: "Claude Haiku 4.5",
132− model: "claude-haiku-4-5-20251001",
133− price: { input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 },
146+ modelName: "Claude Haiku 5.5",
147+ model: "claude-haiku-5-5",
148+ price: { input: 0.1, output: 0.5, cacheRead: 0.01, cacheWrite: 0.125 },
134149 },
135150 large: {
136151 modelName: "Claude Sonnet 5.5",
137152 model: "claude-sonnet-5-5",
138− price: { input: 2, output: 10, cacheRead: 0.2, cacheWrite: 2.5 },
153+ price: { input: 2, output: 10, cacheRead: 0.1, cacheWrite: 2.5 },
139154 },
140155 frontier: {
141156 modelName: "Claude Opus 5.5",
143158 price: { input: 4, output: 20, cacheRead: 0.2, cacheWrite: 5 },
144159 },
145160 },
146− tasks: { implement: "large", revise: "large", answer: "small", review: "change", update: "small", plan: "large" },
161+ // Plans start on the fast model thinking hard, and go up a tier when
162+ // one fails or leaves low confidence, as any work does.
163+ tasks: { implement: "large", revise: "large", answer: "small", review: "change", update: "small", plan: "small" },
164+ effort: { plan: "high", answer: "medium", update: "low" },
147165 smallChange: { files: 10, lines: 200 },
148166 largeChange: { files: 60, lines: 3000 },
149167 largeLabels: ["security"],
174192 for (const [kind, rule] of Object.entries(given.tasks ?? {})) {
175193 if (kind in tasks && (isTier(rule) || rule === "change")) tasks[kind as JobKind] = rule;
176194 }
195+ const effort = { ...DEFAULT_ROUTING.effort };
196+ for (const [kind, level] of Object.entries(given.effort ?? {})) {
197+ if (kind in tasks && EFFORTS.includes(level as Effort)) effort[kind as JobKind] = level as Effort;
198+ }
177199 const tiers = { ...DEFAULT_ROUTING.tiers };
178200 for (const tier of TIERS) {
179201 const route = given.tiers?.[tier];
188210 return {
189211 tiers,
190212 tasks,
213+ effort,
191214 smallChange: { ...DEFAULT_ROUTING.smallChange, ...given.smallChange },
192215 largeChange: { ...DEFAULT_ROUTING.largeChange, ...given.largeChange },
193216 largeLabels: labels(given.largeLabels, DEFAULT_ROUTING.largeLabels),
+1−1
8787 // "smallLabels" move work by its issue's labels. A failed attempt goes one
8888 // tier up and "frontierAfter" failures in a row to frontier; "learning"
8989 // steps work down or up by the repository's own recent runs of the kind.
90− "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 4.5\",\"model\":\"claude-haiku-4-5-20251001\",\"price\":{\"input\":1,\"output\":5,\"cacheRead\":0.1,\"cacheWrite\":1.25}},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\",\"price\":{\"input\":2,\"output\":10,\"cacheRead\":0.2,\"cacheWrite\":2.5}},\"frontier\":{\"modelName\":\"Claude Opus 5.5\",\"model\":\"claude-opus-5-5\",\"price\":{\"input\":4,\"output\":20,\"cacheRead\":0.2,\"cacheWrite\":5}}},\"tasks\":{\"implement\":\"large\",\"revise\":\"large\",\"answer\":\"small\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"large\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeChange\":{\"files\":60,\"lines\":3000},\"largeLabels\":[\"security\"],\"frontierLabels\":[\"architecture\"],\"smallLabels\":[\"documentation\",\"docs\",\"typo\"],\"frontierAfter\":2,\"learning\":{\"window\":20,\"minRuns\":5,\"stepDownAt\":0.9,\"stepUpAt\":0.5}}",
90+ "AGENT_ROUTING": "{\"tiers\":{\"small\":{\"modelName\":\"Claude Haiku 5.5\",\"model\":\"claude-haiku-5-5\",\"price\":{\"input\":0.1,\"output\":0.5,\"cacheRead\":0.01,\"cacheWrite\":0.125}},\"large\":{\"modelName\":\"Claude Sonnet 5.5\",\"model\":\"claude-sonnet-5-5\",\"price\":{\"input\":2,\"output\":10,\"cacheRead\":0.1,\"cacheWrite\":2.5}},\"frontier\":{\"modelName\":\"Claude Opus 5.5\",\"model\":\"claude-opus-5-5\",\"price\":{\"input\":4,\"output\":20,\"cacheRead\":0.2,\"cacheWrite\":5}}},\"tasks\":{\"implement\":\"large\",\"revise\":\"large\",\"answer\":\"small\",\"review\":\"change\",\"update\":\"small\",\"plan\":\"small\"},\"effort\":{\"plan\":\"high\",\"answer\":\"medium\",\"update\":\"low\"},\"smallChange\":{\"files\":10,\"lines\":200},\"largeChange\":{\"files\":60,\"lines\":3000},\"largeLabels\":[\"security\"],\"frontierLabels\":[\"architecture\"],\"smallLabels\":[\"documentation\",\"docs\",\"typo\"],\"frontierAfter\":2,\"learning\":{\"window\":20,\"minRuns\":5,\"stepDownAt\":0.9,\"stepUpAt\":0.5}}",
9191 // Where sandboxes send model requests, with a token for their run.
9292 // The proxy holds the keys: g1t's gateway's, or the workspace's own.
9393 "MODELS_URL": "https://models.g1t.sh",