Skip to content

Compare changes

Choose two branches to see what one has that the other does not, then open a pull request for it.

Open a pull request

1 commit

35 files+3654−610/35 viewed
+5−1
253253 your key.
254254
255255 `GET /openai/v1/models` lists what the workspace can use: its own
256−providers' models first, then g1t's, cheapest Claude first. Each has
256+providers' models first, then g1t's, starting with the Claude g1t suggests
257+starting with (Claude Haiku 5.5 today). g1t adds models as providers
258+release them, once their prices are confirmed, so the list and the tables
259+below grow over time; a model a provider stops offering is listed until it
260+is retired. Each has
257261 `billed_to` (`workspace` or `g1t`), `connection` (your provider's name) and,
258262 on g1t's models, `pricing` in dollars per million tokens:
259263
+7−0
4747 of the same kind, Auto moves that work down a tier there; when a model
4848 keeps failing, up.
4949
50+The models behind the tiers are today's. g1t keeps up with new models
51+as providers release them: it checks for new ones every day, and when g1t
52+moves a tier to a new model, your runs use it within a minute, with
53+nothing for you to change. Nobody picks a model; Auto keeps choosing by the
54+work. A model a provider retires is never used again: the next model for
55+that tier runs instead, and the run says so.
56+
5057 Every run says which model it used and why, in one line on its run and in
5158 its pull request's session, such as *Used a fast model (Claude Haiku 5.5):
5259 small change, 3 files and 80 lines.* The full rules are in
+4−0
452452 | Standard | Claude Sonnet 5.5 | Most changes and reviews |
453453 | Most capable | Claude Opus 5.5 | Hard work, and work that failed on the standard model |
454454
455+The models are today's: when g1t moves a tier to a newer model, runs use
456+it within a minute. A model its provider retires is never used; the next
457+model for the tier runs, and the run's line says so.
458+
455459 The job starts on its tier:
456460
457461 | Work | Starts on |
+23−1
102102 to Stripe's hosted invoice page. An invoice also goes out on its own as
103103 each month closes: one Stripe invoice, a line per workspace for what it
104104 owes, net 30, emailed by Stripe.
105+- **Agents & models** (`/agents`, under Platform): billing's model
106+ catalogue, every model g1t can use (`admin_models`). **Check for new
107+ models** lists each provider's models now, through the model proxy's
108+ `Discovery` entrypoint (the `MODELS` binding), as the daily check does;
109+ the result says what each provider listed, what is new or gone, or why a
110+ provider could not be listed. **Defaults**: the model behind each of
111+ Auto's tiers, the harness's background model and the AI Gateway's first
112+ Claude (an available, priced Claude each, shown with what a typical run
113+ costs on it), and each kind of job's starting tier and effort. A change
114+ needs a reason and shows a review first, the current and new value side
115+ by side with what a typical run would cost on each, before **Save**
116+ (`admin_set_model_default`); runs pick it up within a minute. A default
117+ that has fallen back (its model retired or no longer listed) says so.
118+ **New models**: each model a check found, with its prices filled in
119+ where known; confirm its name, tier and prices per million tokens (and
120+ long-prompt prices) and **Approve**, or **Retire** it
121+ (`admin_decide_model`). **Catalogue**: every other model with its status,
122+ context, prices, typical run and when its provider last listed it, each
123+ with **Retire** or **Restore**. **Checks**: the latest checks of each
124+ provider. Every change names the staff member and why in the audit log
125+ (account `models`). See docs/BILLING_OPERATIONS.md, "The model
126+ catalogue".
105127 - **Stripe**: whether billing's key is in test or live mode (or off), the
106128 webhook Stripe calls (URL, endpoint id, events, who registered it and
107129 when), and the events Stripe sent lately with what billing did with
193215
194216 ```sh
195217 npm run typecheck -w @g1t/sudo
196−npm test -w @g1t/sudo # JWT verification, forms, money, the workspace join, paging, nav, charts, signals
218+npm test -w @g1t/sudo # JWT verification, forms, money, the workspace join, paging, nav, charts, signals, models
197219 npm run build -w @g1t/sudo
198220 ```
199221
+9−0
6767 */
6868 export function accountPath(account: string | null | undefined): string | null {
6969 const id = (account ?? "").trim().toLowerCase();
70+ // The model catalogue's changes (billing's catalogue.rs).
71+ if (id === "models") return "/agents";
7072 const status = /^(incident|maintenance):([a-z0-9-]{1,64})$/.exec(id);
7173 if (status) return status[1] === "incident" ? `/incidents/${status[2]}` : `/incidents/maintenance/${status[2]}`;
7274 if (ENTERPRISE.test(id)) return `/enterprises/${encodeURIComponent(id)}`;
7779 /** What to call an account: a workspace by its slug, an enterprise by its name when known. */
7880 export function accountName(account: string, names: Map<string, string> = new Map()): string {
7981 if (names.has(account)) return names.get(account) as string;
82+ if (account === "models") return "Agents & models";
8083 if (account.startsWith("incident:")) return "Incident";
8184 if (account.startsWith("maintenance:")) return "Maintenance";
8285 if (account.startsWith("ws_")) return account.slice(3);
133136 maintenance_cancelled: "Maintenance cancelled",
134137 credit_revoked: "Credit revoked",
135138 credit_expired: "Credit expired",
139+ // The model catalogue (Agents & models).
140+ models_discovered: "Models found or gone",
141+ model_approved: "Model approved",
142+ model_retired: "Model retired",
143+ model_restored: "Model restored",
144+ model_default: "Model default changed",
136145 };
137146
138147 /**
+134−0
1+import assert from "node:assert/strict";
2+import { test } from "node:test";
3+
4+import type { CatalogueModel } from "@g1t/contracts";
5+
6+import { choicesFor, describeDefault, impact, parseApproval, parseDefault, parsePerMillion, perMillion, tokens } from "./models.ts";
7+
8+function model(id: string, name: string, typical: number, extra: Partial<CatalogueModel> = {}): CatalogueModel {
9+ return {
10+ model: id,
11+ name,
12+ provider: "anthropic",
13+ kind: "chat",
14+ inputMicros: 1_000_000,
15+ outputMicros: 5_000_000,
16+ cacheReadMicros: 100_000,
17+ cacheWriteMicros: 1_250_000,
18+ aliases: [],
19+ family: "haiku",
20+ tierHint: "small",
21+ contextWindow: 0,
22+ maxOutput: 0,
23+ capabilities: [],
24+ dimensions: 0,
25+ status: "available",
26+ priced: true,
27+ source: "staff",
28+ firstSeenAt: null,
29+ lastSeenAt: null,
30+ missingSince: null,
31+ approvedBy: null,
32+ approvedAt: null,
33+ note: "",
34+ typicalRunMicros: typical,
35+ ...extra,
36+ };
37+}
38+
39+const catalogue = [
40+ model("claude-haiku-5-5", "Claude Haiku 5.5", 76_000),
41+ model("claude-sonnet-5-5", "Claude Sonnet 5.5", 1_340_000, { tierHint: "large" }),
42+ model("claude-haiku-6", "Claude Haiku 6", 0, { status: "new", priced: false }),
43+ model("@cf/openai/gpt-oss-120b", "gpt-oss-120b", 30_000, { provider: "workers-ai" }),
44+ model("claude-haiku-4-5", "Claude Haiku 4.5", 600_000, { status: "retired" }),
45+];
46+
47+function form(fields: Record<string, string>): FormData {
48+ const data = new FormData();
49+ for (const [name, value] of Object.entries(fields)) data.set(name, value);
50+ return data;
51+}
52+
53+test("prices per million are typed in dollars to a millionth", () => {
54+ assert.equal(parsePerMillion("0.125"), 125_000);
55+ assert.equal(parsePerMillion("$2"), 2_000_000);
56+ assert.equal(parsePerMillion("12.50"), 12_500_000);
57+ assert.equal(parsePerMillion("0.000001"), 1);
58+ assert.equal(parsePerMillion(""), 0);
59+ assert.equal(parsePerMillion("-1"), null);
60+ assert.equal(parsePerMillion("0.0000001"), null);
61+ assert.equal(parsePerMillion("abc"), null);
62+ assert.equal(perMillion(125_000), "0.125");
63+ assert.equal(perMillion(100_000), "0.10");
64+ assert.equal(perMillion(2_000_000), "2");
65+});
66+
67+test("approving a model takes its name, tier, prices and why", () => {
68+ const approved = parseApproval(
69+ form({ name: "Claude Haiku 6", tier_hint: "small", input: "0.08", output: "0.40", cache_read: "0.008", cache_write: "0.10", reason: "Anthropic's price page" }),
70+ "chat",
71+ );
72+ assert.ok(approved.ok);
73+ assert.equal(approved.value.prices.inputMicros, 80_000);
74+ assert.equal(approved.value.prices.outputMicros, 400_000);
75+ // No hour-long price: what a five-minute write costs.
76+ assert.equal(approved.value.prices.cacheWrite1hMicros, 100_000);
77+ assert.equal(approved.value.prices.threshold, 0);
78+ assert.equal(parseApproval(form({ name: "X", input: "1", output: "", reason: "r" }), "chat").ok, false);
79+ assert.equal(parseApproval(form({ name: "X", input: "0.012", reason: "r" }), "embeddings").ok, true);
80+ assert.deepEqual(parseApproval(form({ name: "X", input: "1", output: "5", reason: "" }), "chat"), { ok: false, error: "Say why, for whoever looks next." });
81+ assert.equal(parseApproval(form({ name: "X", input: "1", output: "5", over_input: "2", reason: "r" }), "chat").ok, false);
82+ const long = parseApproval(form({ name: "X", input: "1", output: "5", threshold: "200,000", over_input: "2", over_output: "10", reason: "r" }), "chat");
83+ assert.ok(long.ok);
84+ assert.equal(long.value.prices.threshold, 200_000);
85+ assert.equal(parseApproval(form({ name: "X", tier_hint: "huge", input: "1", output: "5", reason: "r" }), "chat").ok, false);
86+ assert.equal(parseApproval(form({ name: "", input: "1", output: "5", reason: "r" }), "chat").ok, false);
87+});
88+
89+test("a model purpose takes only an available, priced Claude", () => {
90+ const choices = choicesFor(catalogue);
91+ assert.deepEqual(choices.map((m) => m.model), ["claude-haiku-5-5", "claude-sonnet-5-5"]);
92+ const set = parseDefault(form({ purpose: "tier_small", model: "claude-sonnet-5-5", reason: "testing" }), choices);
93+ assert.deepEqual(set, { ok: true, value: { purpose: "tier_small", model: "claude-sonnet-5-5", tier: null, effort: null, reason: "testing" } });
94+ assert.equal(parseDefault(form({ purpose: "tier_small", model: "claude-haiku-6", reason: "x" }), choices).ok, false);
95+ assert.equal(parseDefault(form({ purpose: "background", model: "@cf/openai/gpt-oss-120b", reason: "x" }), choices).ok, false);
96+ assert.equal(parseDefault(form({ purpose: "tier_small", model: "claude-haiku-5-5", reason: "" }), choices).ok, false);
97+ assert.equal(parseDefault(form({ purpose: "nonsense", model: "claude-haiku-5-5", reason: "x" }), choices).ok, false);
98+});
99+
100+test("a job takes a tier and an effort, or the harness's own", () => {
101+ const choices = choicesFor(catalogue);
102+ assert.deepEqual(parseDefault(form({ purpose: "job_plan", tier: "small", effort: "xhigh", reason: "x" }), choices), {
103+ ok: true,
104+ value: { purpose: "job_plan", model: null, tier: "small", effort: "xhigh", reason: "x" },
105+ });
106+ const own = parseDefault(form({ purpose: "job_update", tier: "small", effort: "", reason: "x" }), choices);
107+ assert.ok(own.ok);
108+ assert.equal(own.value.effort, null);
109+ // Only a review is sized by the change.
110+ assert.equal(parseDefault(form({ purpose: "job_review", tier: "change", reason: "x" }), choices).ok, true);
111+ assert.equal(parseDefault(form({ purpose: "job_plan", tier: "change", reason: "x" }), choices).ok, false);
112+ assert.equal(parseDefault(form({ purpose: "job_plan", tier: "small", effort: "huge", reason: "x" }), choices).ok, false);
113+});
114+
115+test("a change shows what a typical run would cost before it is saved", () => {
116+ const before = { purpose: "tier_small", chosen: "claude-haiku-5-5", model: catalogue[0], capabilities: [], note: null };
117+ assert.equal(
118+ impact(before, catalogue[1], catalogue),
119+ "A typical run: about $1.34 on Claude Sonnet 5.5, 17.6 times $0.076 on Claude Haiku 5.5.",
120+ );
121+ assert.equal(impact({ ...before, model: catalogue[1] }, catalogue[0], catalogue), "A typical run: about $0.076 on Claude Haiku 5.5, 94% less than $1.34 on Claude Sonnet 5.5.");
122+ assert.equal(impact(before, catalogue[0], catalogue), "No change: Claude Haiku 5.5 already runs it, at about $0.076 a typical run.");
123+ assert.equal(impact(undefined, catalogue[1], catalogue), "A typical run would cost about $1.34 on Claude Sonnet 5.5.");
124+ assert.equal(impact(before, undefined, catalogue), null);
125+});
126+
127+test("defaults and sizes read plainly", () => {
128+ assert.equal(describeDefault({ model: "claude-haiku-5-5", tier: null, effort: null }, catalogue), "Claude Haiku 5.5");
129+ assert.equal(describeDefault({ model: null, tier: "small", effort: "high" }, catalogue), "Fast, high effort");
130+ assert.equal(describeDefault({ model: null, tier: "change", effort: null }, catalogue), "By the change's size, the harness's own effort");
131+ assert.equal(tokens(1_000_000), "1M");
132+ assert.equal(tokens(200_000), "200k");
133+ assert.equal(tokens(0), "—");
134+});
+235−0
1+/**
2+ * Agents & models, as the page reads its forms and states its figures:
3+ * prices per million tokens, each purpose's default, and what a change
4+ * does to the cost of a typical run. Billing checks every value again. No
5+ * Workers imports, so it can be tested under Node.
6+ */
7+import type { CatalogueModel, ModelDefault, ModelPrices, ResolvedModel } from "@g1t/contracts";
8+
9+import type { Parsed } from "./forms.ts";
10+
11+/** The longest reason billing keeps. */
12+export const MAX_REASON = 500;
13+
14+/** A model purpose: what it is called and what it is for. */
15+export type PurposeInfo = { purpose: string; label: string; about: string };
16+
17+/** The model purposes, in the order the page shows them. */
18+export const MODEL_PURPOSES: PurposeInfo[] = [
19+ { purpose: "tier_small", label: "Fast", about: "Auto's fast tier: catching up, answers, plans, small reviews." },
20+ { purpose: "tier_large", label: "Standard", about: "Auto's standard tier: making and revising changes, most reviews." },
21+ { purpose: "tier_frontier", label: "Most capable", about: "Auto's top tier: large reviews, architecture, work that failed twice." },
22+ { purpose: "background", label: "Background", about: "The harness's own small tasks in every run, such as titles and summaries." },
23+ {
24+ purpose: "gateway_first",
25+ label: "AI Gateway's first Claude",
26+ about: "Listed first by GET /openai/v1/models, the one people start with.",
27+ },
28+];
29+
30+/** Each kind of agent job, as the page names it. */
31+export const JOBS: { kind: string; label: string }[] = [
32+ { kind: "implement", label: "Making a change" },
33+ { kind: "revise", label: "Revising a change" },
34+ { kind: "answer", label: "Answering a question" },
35+ { kind: "review", label: "Reviewing" },
36+ { kind: "update", label: "Catching up" },
37+ { kind: "plan", label: "Planning" },
38+];
39+
40+export const TIER_LABELS: Record<string, string> = {
41+ small: "Fast",
42+ large: "Standard",
43+ frontier: "Most capable",
44+ change: "By the change's size",
45+};
46+
47+export const EFFORTS = ["low", "medium", "high", "xhigh", "max"] as const;
48+
49+/** What a purpose is called anywhere on the page or in a notice. */
50+export function purposeLabel(purpose: string): string {
51+ const model = MODEL_PURPOSES.find((p) => p.purpose === purpose);
52+ if (model) return model.label;
53+ const job = JOBS.find((j) => `job_${j.kind}` === purpose);
54+ return job ? job.label : purpose;
55+}
56+
57+/** The catalogue models a purpose may be set to: available, priced Claude chat models. */
58+export function choicesFor(catalogue: CatalogueModel[]): CatalogueModel[] {
59+ return catalogue.filter((m) => m.status === "available" && m.priced && m.provider === "anthropic" && (m.kind ?? "chat") === "chat");
60+}
61+
62+/** A price per million tokens, in dollars, to show or to fill a field: `0.125`, `2`, `12.50`. */
63+export function perMillion(micros: number | null | undefined): string {
64+ if (micros == null) return "";
65+ const text = (Math.max(0, micros) / 1_000_000).toFixed(6).replace(/0+$/, "").replace(/\.$/, "");
66+ const [whole, cents] = text.split(".");
67+ return cents && cents.length === 1 ? `${whole}.${cents}0` : text;
68+}
69+
70+/**
71+ * Micros from a price per million tokens as typed: `0.125`, `$2`, `12.50`.
72+ * Up to six decimal places (a millionth of a dollar per million tokens);
73+ * an empty field is 0.
74+ */
75+export function parsePerMillion(input: string): number | null {
76+ const text = input.trim().replace(/^\$/, "").replace(/,/g, "");
77+ if (text === "") return 0;
78+ const match = /^(\d{1,4})(?:\.(\d{1,6}))?$/.exec(text);
79+ if (!match) return null;
80+ return Number(match[1]) * 1_000_000 + Number((match[2] ?? "").padEnd(6, "0"));
81+}
82+
83+/** The price fields of the approval form, by the name each is posted as. */
84+export const PRICE_FIELDS: { name: string; key: keyof ModelPrices; label: string }[] = [
85+ { name: "input", key: "inputMicros", label: "Input" },
86+ { name: "output", key: "outputMicros", label: "Output" },
87+ { name: "cache_read", key: "cacheReadMicros", label: "Cache read" },
88+ { name: "cache_write", key: "cacheWriteMicros", label: "Cache write, 5 min" },
89+ { name: "cache_write_1h", key: "cacheWrite1hMicros", label: "Cache write, 1 h" },
90+];
91+
92+export const OVER_FIELDS: { name: string; key: keyof ModelPrices; label: string }[] = [
93+ { name: "over_input", key: "overInputMicros", label: "Input" },
94+ { name: "over_output", key: "overOutputMicros", label: "Output" },
95+ { name: "over_cache_read", key: "overCacheReadMicros", label: "Cache read" },
96+ { name: "over_cache_write", key: "overCacheWriteMicros", label: "Cache write, 5 min" },
97+ { name: "over_cache_write_1h", key: "overCacheWrite1hMicros", label: "Cache write, 1 h" },
98+];
99+
100+/** A model's prices as the catalogue has them, for the approval form. */
101+export function pricesOf(model: CatalogueModel): ModelPrices {
102+ return {
103+ inputMicros: model.inputMicros,
104+ outputMicros: model.outputMicros,
105+ cacheReadMicros: model.cacheReadMicros,
106+ cacheWriteMicros: model.cacheWriteMicros,
107+ cacheWrite1hMicros: model.cacheWrite1hMicros ?? 0,
108+ threshold: model.threshold ?? 0,
109+ overInputMicros: model.overInputMicros ?? 0,
110+ overOutputMicros: model.overOutputMicros ?? 0,
111+ overCacheReadMicros: model.overCacheReadMicros ?? 0,
112+ overCacheWriteMicros: model.overCacheWriteMicros ?? 0,
113+ overCacheWrite1hMicros: model.overCacheWrite1hMicros ?? 0,
114+ };
115+}
116+
117+/** A reason as the forms take it: required, kept short. */
118+export function parseReason(raw: string): Parsed<string> {
119+ const reason = raw.trim();
120+ if (!reason) return { ok: false, error: "Say why, for whoever looks next." };
121+ if ([...reason].length > MAX_REASON) return { ok: false, error: `Keep the reason to ${MAX_REASON} characters.` };
122+ return { ok: true, value: reason };
123+}
124+
125+export type Approval = { name: string; tierHint: string; prices: ModelPrices; reason: string };
126+
127+/**
128+ * The approval form: the name people see, the tier it suits, its prices
129+ * per million tokens (and above a long-prompt threshold, if it has one),
130+ * and why. Input is required; a chat model needs an output price too.
131+ */
132+export function parseApproval(form: FormData, kind: string): Parsed<Approval> {
133+ const get = (name: string) => {
134+ const value = form.get(name);
135+ return typeof value === "string" ? value : "";
136+ };
137+ const name = get("name").trim();
138+ if (!name) return { ok: false, error: "Give the name people see, such as Claude Haiku 6." };
139+ if (name.length > 120) return { ok: false, error: "Keep the name to 120 characters." };
140+ const tierHint = get("tier_hint").trim();
141+ if (tierHint && !["small", "large", "frontier"].includes(tierHint)) return { ok: false, error: "Choose a tier it suits, or none." };
142+ const prices = {} as ModelPrices;
143+ for (const field of [...PRICE_FIELDS, ...OVER_FIELDS]) {
144+ const micros = parsePerMillion(get(field.name));
145+ if (micros == null) return { ok: false, error: `${field.label}: a price in dollars per million tokens, such as 0.125.` };
146+ prices[field.key] = micros;
147+ }
148+ const threshold = get("threshold").trim().replace(/,/g, "");
149+ if (threshold && !/^\d{1,9}$/.test(threshold)) return { ok: false, error: "The long-prompt threshold is a number of tokens, such as 200000." };
150+ prices.threshold = threshold ? Number(threshold) : 0;
151+ if (prices.inputMicros === 0) return { ok: false, error: "Give the input price per million tokens." };
152+ if (kind === "chat" && prices.outputMicros === 0) return { ok: false, error: "Give the output price per million tokens." };
153+ // An hour-long write with no price of its own costs what a five-minute one does.
154+ if (prices.cacheWrite1hMicros === 0) prices.cacheWrite1hMicros = prices.cacheWriteMicros;
155+ const over = prices.overInputMicros + prices.overOutputMicros + prices.overCacheReadMicros + prices.overCacheWriteMicros;
156+ if (prices.threshold === 0 && over > 0) return { ok: false, error: "Long-prompt prices need the prompt length they start above." };
157+ if (prices.threshold > 0 && prices.overInputMicros === 0) return { ok: false, error: "With a long-prompt threshold, give the prices above it." };
158+ if (prices.threshold > 0 && prices.overCacheWrite1hMicros === 0) prices.overCacheWrite1hMicros = prices.overCacheWriteMicros;
159+ const reason = parseReason(get("reason"));
160+ if (!reason.ok) return reason;
161+ return { ok: true, value: { name, tierHint, prices, reason: reason.value } };
162+}
163+
164+/** A default as the forms post it: a model, or a job's tier and effort. */
165+export type DefaultChange = { purpose: string; model: string | null; tier: string | null; effort: string | null; reason: string };
166+
167+/**
168+ * One default's form: a model purpose takes one of `choices`; a job takes
169+ * a tier (`change` only for reviews) and an effort, or none for the
170+ * harness's own.
171+ */
172+export function parseDefault(form: FormData, choices: CatalogueModel[]): Parsed<DefaultChange> {
173+ const get = (name: string) => {
174+ const value = form.get(name);
175+ return typeof value === "string" ? value.trim() : "";
176+ };
177+ const purpose = get("purpose");
178+ const reason = parseReason(get("reason"));
179+ if (MODEL_PURPOSES.some((p) => p.purpose === purpose)) {
180+ const model = get("model");
181+ if (!choices.some((m) => m.model === model)) return { ok: false, error: "Choose an available Claude model." };
182+ if (!reason.ok) return reason;
183+ return { ok: true, value: { purpose, model, tier: null, effort: null, reason: reason.value } };
184+ }
185+ const job = JOBS.find((j) => `job_${j.kind}` === purpose);
186+ if (!job) return { ok: false, error: "That is not something a default is chosen for." };
187+ const tier = get("tier");
188+ const tiers = job.kind === "review" ? ["small", "large", "frontier", "change"] : ["small", "large", "frontier"];
189+ if (!tiers.includes(tier)) return { ok: false, error: `Choose ${job.kind === "review" ? "a tier, or by the change's size" : "a tier"}.` };
190+ const effort = get("effort");
191+ if (effort && !(EFFORTS as readonly string[]).includes(effort)) return { ok: false, error: "Effort is low, medium, high, xhigh or max, or the harness's own." };
192+ if (!reason.ok) return reason;
193+ return { ok: true, value: { purpose, model: null, tier, effort: effort || null, reason: reason.value } };
194+}
195+
196+/** How a default reads: a model's name, or `Fast, high effort`. */
197+export function describeDefault(value: { model: string | null; tier: string | null; effort: string | null }, catalogue: CatalogueModel[]): string {
198+ if (value.model) return catalogue.find((m) => m.model === value.model)?.name ?? value.model;
199+ const tier = TIER_LABELS[value.tier ?? ""] ?? value.tier ?? "none";
200+ return value.effort ? `${tier}, ${value.effort} effort` : `${tier}, the harness's own effort`;
201+}
202+
203+/** What a change to a default does to a typical run's cost, in a sentence; null for a job. */
204+export function impact(before: ResolvedModel | undefined, after: CatalogueModel | undefined, catalogue: CatalogueModel[]): string | null {
205+ if (!after) return null;
206+ const was = before?.model ? catalogue.find((m) => m.model === before.model!.model) : undefined;
207+ const now = after.typicalRunMicros;
208+ const money = (micros: number) => (micros < 1_000_000 ? `$${(micros / 1_000_000).toFixed(3)}` : `$${(micros / 1_000_000).toFixed(2)}`);
209+ if (!was || was.typicalRunMicros <= 0) return `A typical run would cost about ${money(now)} on ${after.name}.`;
210+ if (was.model === after.model) return `No change: ${after.name} already runs it, at about ${money(now)} a typical run.`;
211+ const ratio = now / was.typicalRunMicros;
212+ const change = Math.round((ratio - 1) * 100);
213+ const direction =
214+ ratio >= 2
215+ ? `${Number(ratio.toFixed(1))} times`
216+ : change === 0
217+ ? "the same as"
218+ : change > 0
219+ ? `${change}% more than`
220+ : `${-change}% less than`;
221+ return `A typical run: about ${money(now)} on ${after.name}, ${direction} ${money(was.typicalRunMicros)} on ${was.name}.`;
222+}
223+
224+/** The current default for each purpose, by purpose. */
225+export function defaultsByPurpose(defaults: ModelDefault[]): Map<string, ModelDefault> {
226+ return new Map(defaults.map((d) => [d.purpose, d]));
227+}
228+
229+/** A context window or output limit: `1M`, `200k`, or a dash when not known. */
230+export function tokens(n: number): string {
231+ if (!n) return "—";
232+ if (n >= 1_000_000) return `${Number((n / 1_000_000).toFixed(1))}M`;
233+ if (n >= 1_000) return `${Math.round(n / 1_000)}k`;
234+ return String(n);
235+}
+2−2
4040
4141 test("the built pages are not marked soon", () => {
4242 const built = navItems().filter((item) => !item.soon).map((item) => item.to);
43− assert.deepEqual(built, ["/", "/reach-out", "/workspaces", "/enterprises", "/invites", "/aliases", "/requests", "/overages", "/velocity", "/invoices", "/credits", "/stripe", "/costs", "/abuse", "/incidents", "/audit"]);
43+ assert.deepEqual(built, ["/", "/reach-out", "/workspaces", "/enterprises", "/invites", "/aliases", "/requests", "/overages", "/velocity", "/invoices", "/credits", "/stripe", "/costs", "/agents", "/abuse", "/incidents", "/audit"]);
4444 });
4545
4646 test("every soon page says what it will do, why, and what it will have", () => {
4747 const soon = soonItems();
48− assert.ok(soon.length >= 8);
48+ assert.ok(soon.length >= 7);
4949 for (const item of soon) {
5050 assert.ok(item.soon.summary.length >= 1 && item.soon.summary.length <= 4, item.label);
5151 assert.ok(item.soon.plans.length >= 3 && item.soon.plans.length <= 6, item.label);
+1−13
223223 label: "Agents & models",
224224 to: "/agents",
225225 icon: "agents",
226− about: "The models agents run on, what each costs, and how runs are going.",
227− soon: {
228− summary: [
229− "The models g1t's agents run on, and how they are doing: runs, failures, tokens and cost per model, and which workspaces bring their own provider. It is where staff decide which models to offer and see what a change in a provider's price means.",
230− "Hosted models are open to some workspaces and not others; that list belongs here, edited and recorded, not in configuration.",
231− ],
232− plans: [
233− "Runs, failures and cost per model, per day",
234− "Who may use g1t's hosted models, and the free allowance's pool",
235− "Workspaces on their own provider, and the sandbox time their runs use",
236− "Stuck or long-running agents, with a way to stop one",
237− ],
238− },
226+ about: "Every model g1t can use, new ones to approve, and which model each tier and job uses by default.",
239227 },
240228 {
241229 label: "Abuse & fraud",
+4−0
11 import { env } from "cloudflare:workers";
22
33 import {
4+ type ModelDiscoveryApi,
45 type StatusAdminApi,
56 accountsAdminClient,
67 billingAdminClient,
3031 /** The status page's incidents (apps/status's `StatusAdmin` entrypoint). */
3132 export const statusAdmin = env.STATUS as unknown as StatusAdminApi;
3233
34+/** "Check for new models": the model proxy's `Discovery` entrypoint (services/models). */
35+export const modelDiscovery = env.MODELS as unknown as ModelDiscoveryApi;
36+
3337 /** STAFF_EMAILS, for suggesting staff in the incident roles. */
3438 export const staffEmails = (): string => env.STAFF_EMAILS ?? "";
+1−0
1919 route("velocity", "routes/velocity.tsx"),
2020 route("costs", "routes/costs.tsx"),
2121 route("costs/bill", "routes/costs-bill.tsx"),
22+ route("agents", "routes/agents.tsx"),
2223 route("abuse", "routes/abuse.tsx"),
2324 route("invoices", "routes/invoices.tsx"),
2425 route("credits", "routes/credits.tsx"),
+575−0
1+import { Check, RefreshCw } from "lucide-react";
2+import { data, redirect } from "react-router";
3+
4+import type { AdminModels, CatalogueModel, DiscoveryResult, ModelStatus } from "@g1t/contracts";
5+
6+import type { Route } from "./+types/agents";
7+import { Badge, Button, ButtonLink, EmptyState, Field, Input, Notice, PageHeader, Section, Select, When } from "~/components/ui";
8+import { text } from "~/lib/forms";
9+import {
10+ EFFORTS,
11+ JOBS,
12+ MAX_REASON,
13+ MODEL_PURPOSES,
14+ OVER_FIELDS,
15+ PRICE_FIELDS,
16+ TIER_LABELS,
17+ type DefaultChange,
18+ choicesFor,
19+ defaultsByPurpose,
20+ describeDefault,
21+ impact,
22+ parseApproval,
23+ parseDefault,
24+ parseReason,
25+ perMillion,
26+ pricesOf,
27+ purposeLabel,
28+ tokens,
29+} from "~/lib/models";
30+import { usd } from "~/lib/money";
31+import { admin, modelDiscovery } from "~/lib/services.server";
32+import { settle } from "~/lib/settle";
33+import { requireStaff } from "~/lib/staff";
34+
35+export const meta: Route.MetaFunction = () => [{ title: "Agents & models · sudo" }, { name: "robots", content: "noindex, nofollow" }];
36+
37+export async function loader({ request, context }: Route.LoaderArgs) {
38+ requireStaff(context);
39+ const url = new URL(request.url);
40+ const models = await settle(admin.models());
41+ const done = url.searchParams.get("done");
42+ const subject = url.searchParams.get("subject") ?? "";
43+ const message: Record<string, string> = {
44+ approved: `${subject} is available now: defaults can use it, and the AI Gateway offers it.`,
45+ retired: `${subject} is retired. Any default that chose it uses the next suitable model until you choose another.`,
46+ restored: `${subject} is back where it stood.`,
47+ default: `The default for ${subject} is saved. Runs pick it up within a minute.`,
48+ };
49+ return {
50+ models: models.ok ? models.value : null,
51+ error: models.ok ? null : models.error,
52+ done: done && message[done] ? message[done] : null,
53+ };
54+}
55+
56+type Review = { change: DefaultChange; before: string; after: string; impact: string | null };
57+
58+type ActionData =
59+ | { kind: "error"; error: string; target: string; values?: Record<string, string> }
60+ | { kind: "review"; review: Review }
61+ | { kind: "checked"; results: DiscoveryResult[] | null; error: string | null };
62+
63+const back = (done: string, subject: string, anchor: string) =>
64+ redirect(`/agents?done=${done}&subject=${encodeURIComponent(subject)}#${anchor}`);
65+
66+/**
67+ * Checking for new models, approving, retiring and restoring them, and
68+ * changing a default. Billing checks each again and records it in the
69+ * audit log with the staff member and why; a default shows what it does
70+ * to a typical run's cost before it is saved.
71+ */
72+export async function action({ request, context }: Route.ActionArgs) {
73+ const staff = requireStaff(context);
74+ const form = await request.formData();
75+ const failed = (error: string, target: string, values?: Record<string, string>) =>
76+ data<ActionData>({ kind: "error", error, target, values }, { status: 422 });
77+ const intent = text(form, "intent");
78+
79+ if (intent === "check") {
80+ const results = await settle(modelDiscovery.check(staff.email));
81+ return { kind: "checked", results: results.ok ? results.value : null, error: results.ok ? null : results.error } satisfies ActionData;
82+ }
83+
84+ const catalogue = await admin.models().then((m) => m.catalogue);
85+ if (intent === "approve" || intent === "retire" || intent === "restore") {
86+ const id = text(form, "model");
87+ const model = catalogue.find((m) => m.model === id);
88+ if (!model) return failed(`${id} is not in the catalogue.`, `model-${id}`);
89+ if (intent === "approve") {
90+ const approval = parseApproval(form, model.kind ?? "chat");
91+ if (!approval.ok) return failed(approval.error, `model-${id}`, Object.fromEntries([...form.entries()].map(([k, v]) => [k, String(v)])));
92+ const { name, tierHint, prices, reason } = approval.value;
93+ const result = await admin.decideModel(id, "approve", { name, tierHint: tierHint || "", prices }, reason, staff.email);
94+ if (!result.ok) return failed(result.error.message, `model-${id}`);
95+ throw back("approved", result.value.name, "catalogue");
96+ }
97+ const reason = parseReason(text(form, "reason"));
98+ if (!reason.ok) return failed(reason.error, `model-${id}`);
99+ const result = await admin.decideModel(id, intent, {}, reason.value, staff.email);
100+ if (!result.ok) return failed(result.error.message, `model-${id}`);
101+ throw back(intent === "retire" ? "retired" : "restored", result.value.name, "catalogue");
102+ }
103+
104+ if (intent === "default") {
105+ const parsed = parseDefault(form, choicesFor(catalogue));
106+ const purpose = text(form, "purpose");
107+ if (!parsed.ok) return failed(parsed.error, `default-${purpose}`, { reason: text(form, "reason") });
108+ const change = parsed.value;
109+ if (text(form, "confirm") !== "yes") {
110+ const models = await admin.models();
111+ const current = defaultsByPurpose(models.defaults).get(change.purpose);
112+ const resolved = models.resolved.models.find((m) => m.purpose === change.purpose);
113+ return {
114+ kind: "review",
115+ review: {
116+ change,
117+ before: current ? describeDefault(current, catalogue) : "nothing",
118+ after: describeDefault(change, catalogue),
119+ impact: change.model ? impact(resolved, catalogue.find((m) => m.model === change.model), catalogue) : null,
120+ },
121+ } satisfies ActionData;
122+ }
123+ const result = await admin.setModelDefault(change.purpose, change, change.reason, staff.email);
124+ if (!result.ok) return failed(result.error.message, `default-${change.purpose}`);
125+ throw back("default", purposeLabel(change.purpose), "defaults");
126+ }
127+ return data<ActionData>({ kind: "error", error: "Unknown action.", target: "" }, { status: 400 });
128+}
129+
130+const STATUS: Record<ModelStatus, { label: string; tone: "mint" | "lavender" | "warn" | "danger" }> = {
131+ available: { label: "Available", tone: "mint" },
132+ new: { label: "New", tone: "lavender" },
133+ deprecated: { label: "Deprecated", tone: "warn" },
134+ retired: { label: "Retired", tone: "danger" },
135+};
136+
137+const PROVIDER: Record<string, string> = { anthropic: "Anthropic", "workers-ai": "Workers AI" };
138+
139+/** A typical run's cost, with the cents a fast model's run is measured in. */
140+function runCost(micros: number): string {
141+ return micros > 0 ? usd(micros) : "—";
142+}
143+
144+export default function Agents({ loaderData, actionData }: Route.ComponentProps) {
145+ const { models, error, done } = loaderData;
146+ const result = actionData as ActionData | undefined;
147+ const failure = result?.kind === "error" ? result : null;
148+ return (
149+ <main id="top" className="mx-auto max-w-6xl scroll-mt-20 px-4 py-8 sm:py-10">
150+ <PageHeader
151+ title="Agents & models"
152+ description="Every model g1t can use, found by listing each provider daily; new ones wait here for their prices to be confirmed. Staff choose which model each tier and job uses. Customers never pick a model: Auto routes their work with these."
153+ actions={
154+ <form method="post" action="/agents#top">
155+ <input type="hidden" name="intent" value="check" />
156+ <Button type="submit" variant="lavender">
157+ <RefreshCw size={14} />
158+ Check for new models
159+ </Button>
160+ </form>
161+ }
162+ />
163+ <div className="mt-6 space-y-3">
164+ {done && <Notice tone="ok">{done}</Notice>}
165+ {error && <Notice tone="error">Billing did not answer: {error}</Notice>}
166+ {failure && !failure.target && <Notice tone="error">{failure.error}</Notice>}
167+ {result?.kind === "checked" && <Checked results={result.results} error={result.error} />}
168+ </div>
169+ {result?.kind === "review" && <ReviewPanel review={result.review} />}
170+ {models && <Page models={models} failure={failure} />}
171+ </main>
172+ );
173+}
174+
175+function Checked({ results, error }: { results: DiscoveryResult[] | null; error: string | null }) {
176+ if (!results) return <Notice tone="error">The model proxy did not answer: {error}</Notice>;
177+ return (
178+ <Notice tone={results.some((r) => r.error) ? "warn" : "ok"}>
179+ <ul className="space-y-1">
180+ {results.map((r) => (
181+ <li key={r.provider}>
182+ <span className="font-medium text-fg">{PROVIDER[r.provider] ?? r.provider}</span>:{" "}
183+ {r.error
184+ ? `could not be listed (${r.error}). Nothing changed.`
185+ : `${r.listed} listed; ${r.added.length ? `new: ${r.added.join(", ")}` : "nothing new"}${r.deprecated.length ? `; no longer listed: ${r.deprecated.join(", ")}` : ""}${r.restored.length ? `; listed again: ${r.restored.join(", ")}` : ""}.`}
186+ </li>
187+ ))}
188+ </ul>
189+ </Notice>
190+ );
191+}
192+
193+function ReviewPanel({ review }: { review: Review }) {
194+ const { change } = review;
195+ return (
196+ <Section id="review" className="mt-6 border-accent/40" title={`Change ${purposeLabel(change.purpose)}?`} description="Nothing is saved until you confirm.">
197+ <dl className="grid gap-x-6 gap-y-2 text-sm sm:grid-cols-[max-content_minmax(0,1fr)]">
198+ <dt className="text-muted">Now</dt>
199+ <dd>{review.before}</dd>
200+ <dt className="text-muted">After</dt>
201+ <dd className="font-medium">{review.after}</dd>
202+ {review.impact && (
203+ <>
204+ <dt className="text-muted">Cost</dt>
205+ <dd>{review.impact}</dd>
206+ </>
207+ )}
208+ <dt className="text-muted">Why</dt>
209+ <dd className="text-fg-soft">{change.reason}</dd>
210+ </dl>
211+ <form method="post" action="/agents#defaults" className="mt-4 flex flex-wrap gap-2">
212+ <input type="hidden" name="intent" value="default" />
213+ <input type="hidden" name="confirm" value="yes" />
214+ <input type="hidden" name="purpose" value={change.purpose} />
215+ {change.model && <input type="hidden" name="model" value={change.model} />}
216+ {change.tier && <input type="hidden" name="tier" value={change.tier} />}
217+ <input type="hidden" name="effort" value={change.effort ?? ""} />
218+ <input type="hidden" name="reason" value={change.reason} />
219+ <Button type="submit" variant="lavender">
220+ <Check size={14} />
221+ Save
222+ </Button>
223+ <ButtonLink to="/agents#defaults" variant="quiet">
224+ Cancel
225+ </ButtonLink>
226+ </form>
227+ </Section>
228+ );
229+}
230+
231+type Failure = Extract<ActionData, { kind: "error" }> | null;
232+
233+function Page({ models, failure }: { models: AdminModels; failure: Failure }) {
234+ const waiting = models.catalogue.filter((m) => m.status === "new");
235+ const t = models.typical;
236+ return (
237+ <>
238+ <Defaults models={models} failure={failure} />
239+ <Section
240+ id="waiting"
241+ className="mt-6"
242+ title="New models"
243+ description="Found by a check and not used for anything yet. Confirm the name and prices (from the provider's price page) to make one available; until then nothing routes to it, offers it or charges for it."
244+ >
245+ {waiting.length === 0 ? (
246+ <EmptyState title="Nothing waiting">A model a provider starts listing appears here after the next check.</EmptyState>
247+ ) : (
248+ <ul className="space-y-4">
249+ {waiting.map((model) => (
250+ <Waiting key={model.model} model={model} failure={failure?.target === `model-${model.model}` ? failure : null} />
251+ ))}
252+ </ul>
253+ )}
254+ </Section>
255+ <Section
256+ id="catalogue"
257+ className="mt-6"
258+ title="Catalogue"
259+ description={`Prices per million tokens. A typical run is ${t.requests} requests of ${t.input.toLocaleString("en-US")} input, ${t.output.toLocaleString("en-US")} output, ${t.cacheRead.toLocaleString("en-US")} cache-read and ${t.cacheWrite.toLocaleString("en-US")} cache-write tokens: an estimate for comparing models, never a charge.`}
260+ >
261+ <Catalogue catalogue={models.catalogue} failure={failure} />
262+ </Section>
263+ <Section id="checks" className="mt-6" title="Checks" description="Each provider is listed daily at 05:29 UTC, and whenever someone checks from here. Listing models is free; nothing calls a model.">
264+ {models.checks.length === 0 ? (
265+ <EmptyState title="No checks yet">Check for new models above, or wait for the daily check.</EmptyState>
266+ ) : (
267+ <ul className="divide-y divide-line text-sm">
268+ {models.checks.map((check) => (
269+ <li key={check.id} className="flex flex-col gap-1 py-2.5 first:pt-0 last:pb-0 sm:flex-row sm:items-baseline sm:gap-4">
270+ <span className="w-28 shrink-0 font-medium">{PROVIDER[check.provider] ?? check.provider}</span>
271+ <span className="w-44 shrink-0 text-muted">
272+ <When at={check.checkedAt} time />
273+ </span>
274+ <span className="min-w-0 flex-1 text-fg-soft">
275+ {check.error ? (
276+ <span className="text-danger">Failed: {check.error}</span>
277+ ) : (
278+ <>
279+ {check.listed} listed
280+ {check.added.length > 0 && <>; new: {check.added.join(", ")}</>}
281+ {check.deprecated.length > 0 && <>; gone: {check.deprecated.join(", ")}</>}
282+ </>
283+ )}
284+ </span>
285+ <span className="text-xs text-faint">{check.by === "schedule" ? "Daily check" : check.by}</span>
286+ </li>
287+ ))}
288+ </ul>
289+ )}
290+ </Section>
291+ </>
292+ );
293+}
294+
295+function Defaults({ models, failure }: { models: AdminModels; failure: Failure }) {
296+ const choices = choicesFor(models.catalogue);
297+ const current = defaultsByPurpose(models.defaults);
298+ return (
299+ <Section
300+ id="defaults"
301+ className="mt-6"
302+ title="Defaults"
303+ description="What Auto runs each tier on, the harness's background model, the AI Gateway's first Claude, and where each kind of job starts and how hard it thinks. Runs read these within a minute; a model that is retired or no longer listed is never used, and the next suitable one runs instead."
304+ >
305+ <ul className="divide-y divide-line">
306+ {MODEL_PURPOSES.map((info) => {
307+ const row = current.get(info.purpose);
308+ const resolved = models.resolved.models.find((m) => m.purpose === info.purpose);
309+ const running = resolved?.model ? models.catalogue.find((m) => m.model === resolved.model!.model) : undefined;
310+ const error = failure?.target === `default-${info.purpose}` ? failure : null;
311+ return (
312+ <li key={info.purpose} id={`default-${info.purpose}`} className="scroll-mt-20 py-4 first:pt-0 last:pb-0">
313+ <div className="flex flex-wrap items-baseline justify-between gap-x-4 gap-y-1">
314+ <div>
315+ <h3 className="text-sm font-semibold">{info.label}</h3>
316+ <p className="text-xs text-muted">{info.about}</p>
317+ </div>
318+ <div className="text-right text-sm">
319+ <span className="font-medium">{running?.name ?? resolved?.chosen ?? "Not set"}</span>
320+ {running && running.typicalRunMicros > 0 && <span className="ml-2 text-xs text-muted">{runCost(running.typicalRunMicros)} a typical run</span>}
321+ </div>
322+ </div>
323+ {resolved?.note && (
324+ <div className="mt-2">
325+ <Notice tone="warn">{resolved.note}</Notice>
326+ </div>
327+ )}
328+ {row && (
329+ <p className="mt-1 text-xs text-faint">
330+ Set by {row.updatedBy} <When at={row.updatedAt} />
331+ {row.reason ? `: ${row.reason}` : ""}
332+ </p>
333+ )}
334+ <form method="post" action={`/agents#review`} className="mt-3 grid gap-2 sm:grid-cols-[minmax(0,14rem)_minmax(0,1fr)_auto] sm:items-end">
335+ <input type="hidden" name="intent" value="default" />
336+ <input type="hidden" name="purpose" value={info.purpose} />
337+ <Select name="model" defaultValue={row?.model ?? ""} aria-label={`Model for ${info.label}`}>
338+ {choices.map((m) => (
339+ <option key={m.model} value={m.model}>
340+ {m.name} · {runCost(m.typicalRunMicros)}
341+ </option>
342+ ))}
343+ </Select>
344+ <Input name="reason" required maxLength={MAX_REASON} placeholder="Why" aria-label={`Why change ${info.label}`} defaultValue={error?.values?.reason} />
345+ <Button type="submit" variant="quiet">
346+ Review
347+ </Button>
348+ </form>
349+ {error && (
350+ <div className="mt-2">
351+ <Notice tone="error">{error.error}</Notice>
352+ </div>
353+ )}
354+ </li>
355+ );
356+ })}
357+ </ul>
358+ <h3 className="mt-6 border-t border-line pt-4 text-sm font-semibold">Jobs</h3>
359+ <p className="text-xs text-muted">
360+ Where each kind of job starts. Failures, labels and the repository's own history still move it up or down. Effort applies on models that take
361+ it; on one that does not, the harness's own.
362+ </p>
363+ <ul className="mt-3 divide-y divide-line">
364+ {JOBS.map((job) => {
365+ const purpose = `job_${job.kind}`;
366+ const row = current.get(purpose);
367+ const error = failure?.target === `default-${purpose}` ? failure : null;
368+ const tiers = job.kind === "review" ? ["small", "large", "frontier", "change"] : ["small", "large", "frontier"];
369+ return (
370+ <li key={job.kind} id={`default-${purpose}`} className="scroll-mt-20 py-3 first:pt-0 last:pb-0">
371+ <form method="post" action="/agents#review" className="grid gap-2 sm:grid-cols-[10rem_minmax(0,11rem)_minmax(0,9rem)_minmax(0,1fr)_auto] sm:items-center">
372+ <input type="hidden" name="intent" value="default" />
373+ <input type="hidden" name="purpose" value={purpose} />
374+ <div>
375+ <span className="text-sm font-medium">{job.label}</span>
376+ {row && (
377+ <span className="block text-xs text-faint">
378+ {row.updatedBy} <When at={row.updatedAt} />
379+ </span>
380+ )}
381+ </div>
382+ <Select name="tier" defaultValue={row?.tier ?? "large"} aria-label={`Starting tier for ${job.label}`}>
383+ {tiers.map((tier) => (
384+ <option key={tier} value={tier}>
385+ {TIER_LABELS[tier]}
386+ </option>
387+ ))}
388+ </Select>
389+ <Select name="effort" defaultValue={row?.effort ?? ""} aria-label={`Effort for ${job.label}`}>
390+ <option value="">Harness's own</option>
391+ {EFFORTS.map((effort) => (
392+ <option key={effort} value={effort}>
393+ {effort} effort
394+ </option>
395+ ))}
396+ </Select>
397+ <Input name="reason" required maxLength={MAX_REASON} placeholder="Why" aria-label={`Why change ${job.label}`} />
398+ <Button type="submit" variant="quiet">
399+ Review
400+ </Button>
401+ </form>
402+ {error && (
403+ <div className="mt-2">
404+ <Notice tone="error">{error.error}</Notice>
405+ </div>
406+ )}
407+ </li>
408+ );
409+ })}
410+ </ul>
411+ </Section>
412+ );
413+}
414+
415+function Waiting({ model, failure }: { model: CatalogueModel; failure: Failure }) {
416+ const prices = pricesOf(model);
417+ const value = (name: string, fallback: string) => failure?.values?.[name] ?? fallback;
418+ return (
419+ <li id={`model-${model.model}`} className="scroll-mt-20 rounded-lg border border-line bg-bg p-4 sm:p-5">
420+ <div className="flex flex-wrap items-center gap-2">
421+ <span className="font-medium">{model.name}</span>
422+ <code className="text-xs text-muted">{model.model}</code>
423+ <Badge>{PROVIDER[model.provider] ?? model.provider}</Badge>
424+ {!model.priced && <Badge tone="warn">No price known</Badge>}
425+ </div>
426+ <p className="mt-1 text-xs text-muted">
427+ Found <When at={model.firstSeenAt} /> · {model.kind === "embeddings" ? "Embeddings" : "Chat"} · context {tokens(model.contextWindow)}
428+ {model.capabilities.length > 0 && <> · {model.capabilities.join(", ")}</>}
429+ {model.priced && " · prices filled from the provider; check them"}
430+ </p>
431+ <form method="post" action={`/agents#model-${model.model}`} className="mt-4 space-y-3">
432+ <input type="hidden" name="intent" value="approve" />
433+ <input type="hidden" name="model" value={model.model} />
434+ <div className="grid gap-3 sm:grid-cols-[minmax(0,1fr)_12rem]">
435+ <Field label="Name people see">
436+ <Input name="name" required maxLength={120} defaultValue={value("name", model.name)} />
437+ </Field>
438+ <Field label="Suits the tier">
439+ <Select name="tier_hint" defaultValue={value("tier_hint", model.tierHint)}>
440+ <option value="">None</option>
441+ <option value="small">Fast</option>
442+ <option value="large">Standard</option>
443+ <option value="frontier">Most capable</option>
444+ </Select>
445+ </Field>
446+ </div>
447+ <fieldset>
448+ <legend className="mb-1.5 text-sm font-medium text-muted">Dollars per million tokens</legend>
449+ <div className="grid grid-cols-2 gap-3 sm:grid-cols-5">
450+ {PRICE_FIELDS.map((field) => (
451+ <Field key={field.name} label={field.label}>
452+ <Input name={field.name} inputMode="decimal" className="font-mono" defaultValue={value(field.name, model.priced ? perMillion(prices[field.key]) : "")} />
453+ </Field>
454+ ))}
455+ </div>
456+ </fieldset>
457+ <details className="rounded-md border border-line px-3 py-2" open={prices.threshold > 0}>
458+ <summary className="cursor-pointer text-sm text-muted">Priced by prompt length</summary>
459+ <div className="mt-3 space-y-3">
460+ <Field label="Above this many prompt tokens" hint="The whole request is charged at the prices below once its prompt (input and cache tokens) is longer.">
461+ <Input name="threshold" inputMode="numeric" className="font-mono" defaultValue={value("threshold", prices.threshold ? String(prices.threshold) : "")} />
462+ </Field>
463+ <div className="grid grid-cols-2 gap-3 sm:grid-cols-5">
464+ {OVER_FIELDS.map((field) => (
465+ <Field key={field.name} label={field.label}>
466+ <Input name={field.name} inputMode="decimal" className="font-mono" defaultValue={value(field.name, prices[field.key] ? perMillion(prices[field.key]) : "")} />
467+ </Field>
468+ ))}
469+ </div>
470+ </div>
471+ </details>
472+ <Field label="Why" hint="Kept with the model, and in the audit log.">
473+ <Input name="reason" required maxLength={MAX_REASON} placeholder="e.g. Prices from the provider's price page, 2026-10-08." defaultValue={value("reason", "")} />
474+ </Field>
475+ {failure && <Notice tone="error">{failure.error}</Notice>}
476+ <div className="flex justify-end">
477+ <Button type="submit" variant="lavender">
478+ <Check size={14} />
479+ Approve
480+ </Button>
481+ </div>
482+ </form>
483+ <Decide model={model} intent="retire" label="Retire instead" />
484+ </li>
485+ );
486+}
487+
488+/** A retire or restore form, folded away until opened. */
489+function Decide({ model, intent, label }: { model: CatalogueModel; intent: "retire" | "restore"; label: string }) {
490+ return (
491+ <details className="mt-3">
492+ <summary className="cursor-pointer text-xs text-muted hover:text-fg">{label}</summary>
493+ <form method="post" action={`/agents#model-${model.model}`} className="mt-2 flex flex-col gap-2 sm:flex-row sm:items-center">
494+ <input type="hidden" name="intent" value={intent} />
495+ <input type="hidden" name="model" value={model.model} />
496+ <Input name="reason" required maxLength={MAX_REASON} placeholder="Why" aria-label={`Why ${intent} ${model.name}`} />
497+ <Button type="submit" variant={intent === "retire" ? "danger" : "quiet"}>
498+ {intent === "retire" ? "Retire" : "Restore"}
499+ </Button>
500+ </form>
501+ </details>
502+ );
503+}
504+
505+function Catalogue({ catalogue, failure }: { catalogue: CatalogueModel[]; failure: Failure }) {
506+ const listed = catalogue.filter((m) => m.status !== "new");
507+ return (
508+ <div className="-mx-4 overflow-x-auto sm:-mx-5">
509+ <table className="w-full min-w-[56rem] text-sm">
510+ <thead>
511+ <tr className="border-b border-line text-left text-xs text-muted">
512+ <th className="px-4 py-2 font-medium sm:px-5">Model</th>
513+ <th className="px-4 py-2 font-medium">Status</th>
514+ <th className="px-4 py-2 font-medium">Suits</th>
515+ <th className="px-4 py-2 text-right font-medium">Context</th>
516+ <th className="px-4 py-2 text-right font-medium">In / out</th>
517+ <th className="px-4 py-2 text-right font-medium">Cache read</th>
518+ <th className="px-4 py-2 text-right font-medium">Typical run</th>
519+ <th className="px-4 py-2 font-medium sm:pr-5">Listed</th>
520+ </tr>
521+ </thead>
522+ <tbody>
523+ {listed.map((model) => {
524+ const status = STATUS[model.status] ?? STATUS.available;
525+ const error = failure?.target === `model-${model.model}` ? failure : null;
526+ return (
527+ <tr key={model.model} id={`model-${model.model}`} className="scroll-mt-20 border-b border-line align-top last:border-0">
528+ <td className="px-4 py-2.5 sm:px-5">
529+ <span className="font-medium">{model.name}</span>
530+ <code className="block text-xs text-muted">{model.model}</code>
531+ {model.aliases.length > 0 && <span className="block text-xs text-faint">also {model.aliases.join(", ")}</span>}
532+ <span className="block text-xs text-faint">{PROVIDER[model.provider] ?? model.provider}</span>
533+ {error && <span className="mt-1 block text-xs text-danger">{error.error}</span>}
534+ </td>
535+ <td className="px-4 py-2.5">
536+ <Badge tone={status.tone}>{status.label}</Badge>
537+ {!model.priced && (
538+ <span className="mt-1 block">
539+ <Badge tone="warn">Unpriced</Badge>
540+ </span>
541+ )}
542+ {model.status === "available" ? (
543+ <Decide model={model} intent="retire" label="Retire" />
544+ ) : (
545+ <Decide model={model} intent="restore" label="Restore" />
546+ )}
547+ </td>
548+ <td className="px-4 py-2.5 text-muted">{model.kind === "embeddings" ? `Embeddings${model.dimensions ? `, ${model.dimensions}` : ""}` : TIER_LABELS[model.tierHint] ?? "—"}</td>
549+ <td className="px-4 py-2.5 text-right font-mono tabular-nums">{tokens(model.contextWindow)}</td>
550+ <td className="px-4 py-2.5 text-right font-mono tabular-nums whitespace-nowrap">
551+ ${perMillion(model.inputMicros)}
552+ {model.kind !== "embeddings" && <> / ${perMillion(model.outputMicros)}</>}
553+ {model.threshold ? <span className="block text-xs text-faint">over {tokens(model.threshold)}: ${perMillion(model.overInputMicros)} / ${perMillion(model.overOutputMicros)}</span> : null}
554+ </td>
555+ <td className="px-4 py-2.5 text-right font-mono tabular-nums">{model.kind === "embeddings" ? "—" : `$${perMillion(model.cacheReadMicros)}`}</td>
556+ <td className="px-4 py-2.5 text-right font-mono tabular-nums">{runCost(model.typicalRunMicros)}</td>
557+ <td className="px-4 py-2.5 text-xs text-muted sm:pr-5">
558+ {model.missingSince ? (
559+ <span className="text-warn">
560+ Not since <When at={model.missingSince} />
561+ </span>
562+ ) : model.lastSeenAt ? (
563+ <When at={model.lastSeenAt} />
564+ ) : (
565+ "Not checked yet"
566+ )}
567+ </td>
568+ </tr>
569+ );
570+ })}
571+ </tbody>
572+ </table>
573+ </div>
574+ );
575+}
+2−0
88 EVENTS: ServiceBinding;
99 /** apps/status's `StatusAdmin` entrypoint: see `StatusAdminApi`. */
1010 STATUS: Fetcher;
11+ /** services/models's `Discovery` entrypoint: see `ModelDiscoveryApi`. */
12+ MODELS: Fetcher;
1113 ASSETS: Fetcher;
1214 ACCESS_TEAM_DOMAIN: string;
1315 ACCESS_AUD: string;
+5−1
2323 { "binding": "EVENTS", "service": "g1t-events" },
2424 // The status page's staff-only entrypoint: posting incidents to
2525 // status.g1t.sh (apps/status). Only bindings reach it.
26− { "binding": "STATUS", "service": "g1t-status", "entrypoint": "StatusAdmin" }
26+ { "binding": "STATUS", "service": "g1t-status", "entrypoint": "StatusAdmin" },
27+ // "Check for new models" on Agents & models: the model proxy's
28+ // `Discovery` entrypoint lists each provider's models now and records
29+ // them with billing's catalogue. Only bindings reach it.
30+ { "binding": "MODELS", "service": "g1t-models", "entrypoint": "Discovery" }
2731 ],
2832 "vars": {
2933 // The Zero Trust team domain, such as `g1t.cloudflareaccess.com`.
+6−4
237237 failing holds it for a person. There is no model or
238238 agent count to choose: to put more agents to work, assign more issues.
239239 On g1t's hosted models, Auto routes each job to the cheapest of three
240− tiers that can do it, fast, standard and most capable: catching up,
241− answering and reviews of small changes that touch no sensitive path
242− start fast; making changes, revising, planning and other reviews start
243− standard; reviews of very large changes and issues labelled
240+ tiers that can do it, fast, standard and most capable (the model behind
241+ each is today's, and moves to newer models as g1t adopts them; a
242+ retired model is never used): catching up, answering, planning and
243+ reviews of small changes that touch no sensitive path start fast;
244+ making changes, revising and other reviews start standard; reviews of
245+ very large changes and issues labelled
244246 `architecture` start most capable. A failed attempt moves the next one
245247 up a tier (two in a row: most capable), and a repository's own recent
246248 runs move work down or up. Each run states its model and why in one
+318−0
539539 "anthropic".to_owned()
540540 }
541541
542+// --- The model catalogue ----------------------------------------------------
543+//
544+// Every model g1t can use, in one table (billing's `gateway_models`): the
545+// models agents run on, the AI Gateway's, and the embeddings model. New
546+// models are found by the models service listing each provider daily
547+// (`record_discovery`) and wait as `new` until staff approve them in sudo.
548+// Which model each purpose uses by default is staff's choice
549+// (`model_defaults`), read by the runner and the AI Gateway.
550+
551+/// Where a model stands. Only `available` models are routed to; the AI
552+/// Gateway offers `available` and `deprecated` ones that have a price.
553+pub mod model_status {
554+ /// Approved and priced: routed to and offered.
555+ pub const AVAILABLE: &str = "available";
556+ /// Found by discovery and not approved yet: never routed to, offered or charged.
557+ pub const NEW: &str = "new";
558+ /// The provider stopped listing it: still offered to anyone who names
559+ /// it, but no default routes to it.
560+ pub const DEPRECATED: &str = "deprecated";
561+ /// Staff retired it: neither routed to nor offered.
562+ pub const RETIRED: &str = "retired";
563+}
564+
565+/// One model in the catalogue: its prices (as the AI Gateway reads them)
566+/// and what g1t knows about it. `admin_models` returns these.
567+#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
568+#[serde(rename_all = "camelCase")]
569+pub struct CatalogueModel {
570+ #[serde(flatten)]
571+ pub prices: GatewayModel,
572+ /// Other ids the provider lists it by, such as a dated one.
573+ #[serde(default)]
574+ pub aliases: Vec<String>,
575+ /// `haiku`, `sonnet`, `opus`, `fable`, or a Workers AI author.
576+ #[serde(default)]
577+ pub family: String,
578+ /// The agent tier it suits: `small`, `large`, `frontier`, or empty.
579+ #[serde(default)]
580+ pub tier_hint: String,
581+ /// Tokens it reads at most; 0 when not known.
582+ #[serde(default)]
583+ pub context_window: u64,
584+ /// Tokens it writes at most; 0 when not known.
585+ #[serde(default)]
586+ pub max_output: u64,
587+ /// Any of `effort`, `thinking`, `tools`, `vision`, `embeddings`.
588+ #[serde(default)]
589+ pub capabilities: Vec<String>,
590+ /// An embeddings model's vector length; 0 otherwise or when not known.
591+ #[serde(default)]
592+ pub dimensions: u32,
593+ /// `available`, `new`, `deprecated` or `retired` (`model_status`).
594+ pub status: String,
595+ /// Whether its prices are known. An unpriced model is never routed to,
596+ /// offered or charged for.
597+ pub priced: bool,
598+ /// `discovered` (found by listing its provider) or `staff`.
599+ pub source: String,
600+ #[serde(default)]
601+ pub first_seen_at: Option<String>,
602+ /// When its provider last listed it.
603+ #[serde(default)]
604+ pub last_seen_at: Option<String>,
605+ /// Since when its provider has not listed it.
606+ #[serde(default)]
607+ pub missing_since: Option<String>,
608+ #[serde(default)]
609+ pub approved_by: Option<String>,
610+ #[serde(default)]
611+ pub approved_at: Option<String>,
612+ #[serde(default)]
613+ pub note: String,
614+ /// What a typical agent run would cost on it, in millionths of a
615+ /// dollar, from its prices (`typical_run` in billing's catalogue.rs);
616+ /// 0 for an embeddings or unpriced model.
617+ #[serde(default)]
618+ pub typical_run_micros: i64,
619+}
620+
621+/// A provider's list price per million tokens, as its listing gives it, in
622+/// millionths of a dollar.
623+#[derive(Clone, Debug, Default, PartialEq, Serialize, Deserialize)]
624+#[serde(rename_all = "camelCase")]
625+pub struct ListedPrice {
626+ pub input_micros: i64,
627+ #[serde(default)]
628+ pub output_micros: i64,
629+}
630+
631+/// One model as its provider lists it, from the models service's
632+/// discovery.
633+#[derive(Clone, Debug, Default, PartialEq, Serialize, Deserialize)]
634+#[serde(rename_all = "camelCase")]
635+pub struct ProviderModel {
636+ /// The provider's id: `claude-haiku-5-5`, `@cf/openai/gpt-oss-120b`.
637+ pub id: String,
638+ /// For people, as the provider names it.
639+ #[serde(default)]
640+ pub name: String,
641+ /// `chat`, `embeddings`, or anything else (not added to the catalogue,
642+ /// but still counted as listed).
643+ #[serde(default)]
644+ pub kind: String,
645+ #[serde(default)]
646+ pub context_window: u64,
647+ #[serde(default)]
648+ pub max_output: u64,
649+ #[serde(default)]
650+ pub capabilities: Vec<String>,
651+ /// Workers AI lists a price with each model; Anthropic does not.
652+ #[serde(default)]
653+ pub price: Option<ListedPrice>,
654+}
655+
656+/// `record_discovery`: what one provider lists now, from the models
657+/// service (daily, or when staff press "Check for new models"). Billing
658+/// adds new ids as `new`, marks ones no longer listed `deprecated`, records
659+/// the check and emails staff about anything new. With `error` (the listing
660+/// failed) only the check is recorded. Returns `DiscoveryResult`.
661+#[derive(Clone, Debug, Default, Serialize, Deserialize)]
662+#[serde(rename_all = "camelCase")]
663+pub struct RecordDiscoveryArgs {
664+ /// `anthropic` or `workers-ai`.
665+ pub provider: String,
666+ #[serde(default)]
667+ pub models: Vec<ProviderModel>,
668+ /// The staff member who asked, or `schedule`.
669+ pub by: String,
670+ #[serde(default)]
671+ pub error: Option<String>,
672+}
673+
674+/// What one check found.
675+#[derive(Clone, Debug, Default, PartialEq, Serialize, Deserialize)]
676+#[serde(rename_all = "camelCase")]
677+pub struct DiscoveryResult {
678+ pub provider: String,
679+ pub checked_at: String,
680+ pub by: String,
681+ /// Ids the provider listed.
682+ pub listed: u32,
683+ /// Ids added to the catalogue as `new`.
684+ pub added: Vec<String>,
685+ /// Catalogue models the provider no longer lists, now `deprecated`.
686+ pub deprecated: Vec<String>,
687+ /// Deprecated models listed again.
688+ pub restored: Vec<String>,
689+ /// The listing failed: nothing changed.
690+ #[serde(default)]
691+ pub error: Option<String>,
692+}
693+
694+/// Which model, tier or effort one purpose uses by default, as staff last
695+/// set it.
696+#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
697+#[serde(rename_all = "camelCase")]
698+pub struct ModelDefault {
699+ /// `tier_small`, `tier_large`, `tier_frontier`, `background`,
700+ /// `gateway_first`, or `job_<kind>` for `implement`, `revise`,
701+ /// `answer`, `review`, `update` and `plan`.
702+ pub purpose: String,
703+ /// The model, for a model purpose.
704+ #[serde(default)]
705+ pub model: Option<String>,
706+ /// For a job: `small`, `large`, `frontier`, or `change` (sized by the change).
707+ #[serde(default)]
708+ pub tier: Option<String>,
709+ /// For a job: `low`, `medium`, `high`, `xhigh` or `max`; none for the harness's own.
710+ #[serde(default)]
711+ pub effort: Option<String>,
712+ pub updated_at: String,
713+ pub updated_by: String,
714+ #[serde(default)]
715+ pub reason: String,
716+}
717+
718+/// A model purpose's default as it applies now: the chosen model, or the
719+/// one routing falls back to when the chosen one cannot be used.
720+#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
721+#[serde(rename_all = "camelCase")]
722+pub struct ResolvedModel {
723+ pub purpose: String,
724+ /// The model staff chose.
725+ pub chosen: String,
726+ /// The model to use, with its prices; none when neither the chosen
727+ /// model nor any other suits (callers keep their own fallback).
728+ #[serde(default)]
729+ pub model: Option<GatewayModel>,
730+ #[serde(default)]
731+ pub capabilities: Vec<String>,
732+ /// Why it is not the chosen one, in a sentence: `Claude Haiku 5.5 is
733+ /// retired; using Claude Haiku 4.5.`
734+ #[serde(default)]
735+ pub note: Option<String>,
736+}
737+
738+/// One kind of agent job's starting tier and effort.
739+#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
740+#[serde(rename_all = "camelCase")]
741+pub struct JobDefault {
742+ /// `implement`, `revise`, `answer`, `review`, `update` or `plan`.
743+ pub kind: String,
744+ /// `small`, `large`, `frontier` or `change`.
745+ pub tier: String,
746+ #[serde(default)]
747+ pub effort: Option<String>,
748+}
749+
750+/// `model_defaults` takes nothing and returns this: every purpose's model
751+/// as it applies now, and each job's tier and effort. Read by the runner
752+/// (cached a minute) and the AI Gateway.
753+#[derive(Clone, Debug, Default, PartialEq, Serialize, Deserialize)]
754+#[serde(rename_all = "camelCase")]
755+pub struct ModelDefaults {
756+ pub models: Vec<ResolvedModel>,
757+ pub jobs: Vec<JobDefault>,
758+}
759+
760+/// A check of one provider, as `model_checks` keeps it.
761+#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
762+#[serde(rename_all = "camelCase")]
763+pub struct ModelCheck {
764+ pub id: String,
765+ pub provider: String,
766+ pub checked_at: String,
767+ pub by: String,
768+ pub listed: u32,
769+ pub added: Vec<String>,
770+ pub deprecated: Vec<String>,
771+ #[serde(default)]
772+ pub error: Option<String>,
773+}
774+
775+/// `admin_models` takes nothing and returns this: sudo's Agents & models.
776+#[derive(Clone, Debug, Default, Serialize, Deserialize)]
777+#[serde(rename_all = "camelCase")]
778+pub struct AdminModels {
779+ /// Every model, `new` ones first, then by provider and position.
780+ pub catalogue: Vec<CatalogueModel>,
781+ pub defaults: Vec<ModelDefault>,
782+ pub resolved: ModelDefaults,
783+ /// The latest checks, newest first.
784+ pub checks: Vec<ModelCheck>,
785+ /// The token mix `typical_run_micros` prices, for the page to state.
786+ pub typical: TypicalRun,
787+}
788+
789+/// The tokens of the typical agent run estimates are priced from.
790+#[derive(Clone, Copy, Debug, Default, PartialEq, Serialize, Deserialize)]
791+#[serde(rename_all = "camelCase")]
792+pub struct TypicalRun {
793+ pub requests: u64,
794+ /// Per request.
795+ pub input: u64,
796+ pub output: u64,
797+ pub cache_read: u64,
798+ pub cache_write: u64,
799+}
800+
801+/// A model's prices as staff confirm them, per million tokens in
802+/// millionths of a dollar.
803+#[derive(Clone, Debug, Default, PartialEq, Serialize, Deserialize)]
804+pub struct ModelPrices {
805+ pub input_micros: i64,
806+ pub output_micros: i64,
807+ pub cache_read_micros: i64,
808+ pub cache_write_micros: i64,
809+ #[serde(default)]
810+ pub cache_write_1h_micros: i64,
811+ #[serde(default)]
812+ pub threshold: u64,
813+ #[serde(default)]
814+ pub over_input_micros: i64,
815+ #[serde(default)]
816+ pub over_output_micros: i64,
817+ #[serde(default)]
818+ pub over_cache_read_micros: i64,
819+ #[serde(default)]
820+ pub over_cache_write_micros: i64,
821+ #[serde(default)]
822+ pub over_cache_write_1h_micros: i64,
823+}
824+
825+/// `admin_decide_model`: `approve` a model (its prices confirmed, made
826+/// available), `retire` one, or `restore` a retired or deprecated one.
827+/// Audited. Returns `Outcome<CatalogueModel>`.
828+#[derive(Clone, Debug, Default, Serialize, Deserialize)]
829+pub struct AdminDecideModelArgs {
830+ pub model: String,
831+ pub decision: String,
832+ /// On approval: the name people see, and the prices.
833+ #[serde(default)]
834+ pub name: Option<String>,
835+ #[serde(default)]
836+ pub tier_hint: Option<String>,
837+ #[serde(default)]
838+ pub prices: Option<ModelPrices>,
839+ pub reason: String,
840+ pub by: String,
841+}
842+
843+/// `admin_set_model_default`: one purpose's default. A model purpose takes
844+/// `model` (available, priced, and suited to the purpose); a job takes
845+/// `tier` and `effort`. Audited with the old and new values and why.
846+/// Returns `Outcome<ModelDefault>`.
847+#[derive(Clone, Debug, Default, Serialize, Deserialize)]
848+pub struct AdminSetModelDefaultArgs {
849+ pub purpose: String,
850+ #[serde(default)]
851+ pub model: Option<String>,
852+ #[serde(default)]
853+ pub tier: Option<String>,
854+ #[serde(default)]
855+ pub effort: Option<String>,
856+ pub reason: String,
857+ pub by: String,
858+}
859+
542860 /// `gateway_admit`: whether a workspace's next AI Gateway request may go to
543861 /// g1t's models. Fails with `payment_required` and what to do when it may
544862 /// not: over its spend limit, out of AI credit, or not on the plan. Returns
+152−5
303303 keeps `runs.gateway_note`, its correction says why, and it raises the
304304 **Unpriced** drift. For agent runs g1t keeps no token rates of its own:
305305 the first figure is Claude Code's, the final one the gateway's. AI Gateway
306−requests are the exception: they are charged from `gateway_models` (one
307−row per model offered), which has to follow the provider's price list by
306+requests are the exception: they are charged from `gateway_models` (the
307+model catalogue, below), which has to follow the provider's price list by
308308 hand until they are settled like runs. It is not part of the price book's
309−`price_versions`: a price change is a migration that updates the row and
310−its `updated_at` (as 0047 did for Sonnet 5.5's cache reads), and the
311−`gateway_models` meter's markup is the only price-book number on it. Open
309+`price_versions`: a new model's prices are confirmed when staff approve it
310+in sudo, and a changed price on a model already offered is a migration that
311+updates the row and its `updated_at` (as 0047 did for Sonnet 5.5's cache
312+reads). The `gateway_models` meter's markup is the only price-book number
313+on it. Open
312314 models need `WORKERS_AI_TOKEN` (a Cloudflare API token with Workers AI on
313315 g1t's account) on the model proxy; without it, and without
314316 `AI_GATEWAY_TOKEN` holding that permission, they are refused with `503`.
344346 Every usage path goes through this: `finish_run`, settling, sandbox time,
345347 features and builds (`charge_feature`), and the month-end meters.
346348
349+## The model catalogue
350+
351+Every model g1t can use is one row of `gateway_models` (migrations 0045,
352+0047 and `0048_model_catalogue.sql`; code in `catalogue.rs`): the agents'
353+tiers, the AI Gateway's Claude and open models, and the embeddings model.
354+Billing keeps it because billing owns prices: the gateway charges from the
355+same row, so there is one price per model, not two that drift. Page: sudo
356+**Agents & models** (`/agents`).
357+
358+| Column | What |
359+| --- | --- |
360+| `model`, `name`, `provider`, `kind` | The provider's id, the name people see, `anthropic` or `workers-ai`, `chat` or `embeddings` |
361+| `aliases` | Other ids the provider lists it by, comma separated (a dated id such as `claude-haiku-4-5-20251001`) |
362+| `family`, `tier_hint` | `haiku`, `sonnet`, `opus`, `fable` (or a Workers AI author); the agent tier it suits: `small`, `large`, `frontier` |
363+| `context_window`, `max_output`, `capabilities`, `dimensions` | From the provider's list where it gives them: `effort`, `thinking`, `tools`, `vision`, `embeddings`; an embeddings model's vector length |
364+| Prices | Per million tokens in millionths: input, output, cache read, five-minute and hour-long cache writes, and for a model priced by prompt length the `threshold` and `over_` prices |
365+| `status` | `available` (routed to and offered), `new` (found, not approved), `deprecated` (the provider stopped listing it), `retired` (staff took it out) |
366+| `priced` | 1 when its prices are known |
367+| `source`, `first_seen_at`, `last_seen_at`, `missing_since`, `approved_by`, `approved_at`, `note` | Where it came from (`discovered` or `staff`) and its history |
368+
369+What each status allows:
370+
371+| Status | Default for a purpose | Offered on the AI Gateway | Charged |
372+| --- | --- | --- | --- |
373+| `available`, priced | Yes | Yes | Yes |
374+| `new`, or unpriced | No | No: refused before it reaches a provider | No |
375+| `deprecated` | No: a default that chose it falls back | Yes, to whoever names it | Yes |
376+| `retired` | No: a default that chose it falls back | No | No |
377+
378+### Discovery
379+
380+The models service lists each provider once a day (`29 5 * * *`, after
381+billing's daily run) and whenever staff press **Check for new models**
382+(`services/models/src/discover.ts`, through its `Discovery` entrypoint,
383+which only sudo binds). Listing models is free; nothing calls a model.
384+
385+| Provider | How it is listed | Credentials (on g1t-models) |
386+| --- | --- | --- |
387+| Anthropic | `GET /v1/models` through g1t's AI Gateway (`…/anthropic/v1/models`, tagged `task: discovery`), or straight to Anthropic with g1t's key when there is no gateway | `AI_GATEWAY_TOKEN` (the gateway holds Anthropic's key), or `ANTHROPIC_API_KEY` |
388+| Workers AI | `GET /accounts/{account}/ai/models/search` | `WORKERS_AI_TOKEN` (Workers AI Read), else `AI_GATEWAY_TOKEN` |
389+
390+Each provider's list goes to billing's `record_discovery`, which compares
391+it with the catalogue:
392+
393+- **An id it has never seen** is added as `new`. Chat and embeddings
394+ models only. Anthropic models are priced from the price table in
395+ `catalogue.rs` (`ANTHROPIC_PRICES`: Anthropic's list prices by model, a
396+ dated id priced as its model) when it has them; Workers AI models from
397+ the price their listing gives, with cached tokens at the input price.
398+ Anything else is added unpriced. Its family and tier hint come from its
399+ id (`claude-haiku-*` is fast, `claude-sonnet-*` standard,
400+ `claude-opus-*` and `claude-fable-*` most capable).
401+- **A dated id of a model it has** (`claude-sonnet-5-5-20261001`) is that
402+ model: the id is added to its aliases.
403+- **A model it has that is listed** gets `last_seen_at`, and its context
404+ window, output limit and capabilities from the list.
405+- **A model it has that is not listed** (available or new) becomes
406+ `deprecated`, with `missing_since`. Listed again, it goes back to
407+ `available` if it was ever approved, else to `new`.
408+- **An empty list or a failed one** changes nothing: it is recorded as a
409+ failed check with the provider's answer (never a key).
410+
411+Every check is kept 90 days in `model_checks` and shown on the page. When a
412+check adds, deprecates or restores anything, it is in the audit log
413+(`models_discovered`, as `schedule` or the staff member) and staff are
414+emailed at `COSTS_ALERT_EMAIL` with a link to Agents & models.
415+
416+When Anthropic publishes a new model's price, add a row to
417+`ANTHROPIC_PRICES`, so the next one of its kind arrives priced. Until then
418+staff enter the prices when they approve it.
419+
420+### Approving, retiring and restoring
421+
422+On Agents & models, **New models** lists every `new` model with its prices
423+filled in where they are known. To approve one:
424+
425+1. Check the name people see and the tier it suits.
426+2. Check or enter its prices per million tokens against the provider's
427+ price page: input, output, cache read, cache writes (five-minute and
428+ hour-long; an hour-long price left empty is the five-minute one), and
429+ under **Priced by prompt length** the threshold and the prices above it.
430+3. Say why (for example where the prices came from), and **Approve**
431+ (`admin_decide_model`, `approve`).
432+
433+It is `available` at once: the AI Gateway offers it within five minutes
434+(the proxy keeps the catalogue that long) and it can be chosen as a
435+default. **Retire** takes a model out of routing and the gateway; any
436+default that chose it falls back to the next suitable model until staff
437+choose another, and the audit line names those defaults. **Restore** puts
438+a retired or deprecated model back: `available` if it was ever approved,
439+else `new`. Each needs a reason and is in the audit log (`model_approved`,
440+`model_retired`, `model_restored`).
441+
442+### Defaults
443+
444+`model_defaults` holds staff's choice per purpose, each with when, who and
445+why (`admin_set_model_default`; the audit log, `model_default`, has the old
446+value, the new one and why):
447+
448+| Purpose | What it chooses | Read by |
449+| --- | --- | --- |
450+| `tier_small`, `tier_large`, `tier_frontier` | The model behind Auto's fast, standard and most capable tiers | The runner |
451+| `background` | The harness's own small tasks in every run on g1t's tiers (`ANTHROPIC_SMALL_FAST_MODEL`, `ANTHROPIC_DEFAULT_HAIKU_MODEL`) | The runner |
452+| `gateway_first` | The Claude listed first by `GET /openai/v1/models` | Billing's `gateway_models` |
453+| `job_implement`, `job_revise`, `job_answer`, `job_review`, `job_update`, `job_plan` | Each kind of job's starting tier (`small`, `large`, `frontier`; a review also `change`, sized by its change) and effort (`low` to `max`, or none for the harness's own) | The runner |
454+
455+A model purpose takes only an available, priced Claude chat model (the
456+harness speaks Anthropic's API). The migration seeded each with what
457+`AGENT_ROUTING` had.
458+
459+Before a default is saved, sudo shows it beside the current one with what
460+a **typical run** would cost on each: 40 requests of 2,000 input, 1,500
461+output, 45,000 cache-read and 4,000 cache-write tokens, each request priced
462+on its own at the catalogue's prices (`catalogue::TYPICAL`; on 2026-10-08,
463+$0.076 on Claude Haiku 5.5, $1.34 on Sonnet 5.5, $2.68 on Opus 5.5). It is
464+an estimate for comparing models, never a charge.
465+
466+**How the runner reads them.** `model_defaults` (the RPC) returns each
467+purpose's model as it applies now, with its prices and capabilities, and
468+each job's tier and effort. The runner reads it at most once a minute per
469+isolate and puts it over `AGENT_ROUTING` (`withDefaults` in
470+`services/runner/src/model-env.ts`), so a change reaches runs within a
471+minute. When billing cannot be read, the runner uses `AGENT_ROUTING` (and
472+`DEFAULT_ROUTING` under it) alone and asks again ten seconds later. The
473+labels, change sizes, `frontierAfter` and learning always come from
474+`AGENT_ROUTING`. Effort is not sent on a model whose catalogue entry lacks
475+`effort` (Claude Haiku 4.5).
476+
477+**Never a retired model.** A chosen model that is deprecated, retired or
478+unpriced is not handed out: the purpose falls back to the first available,
479+priced Claude with the same tier hint (any Claude for `gateway_first`), in
480+the catalogue's order, with a sentence such as *Claude Haiku 5.5 is
481+retired; using Claude Haiku 4.5 instead.* The runner adds that sentence to
482+the run's reason line, and sudo shows it beside the default. With no model
483+left for a purpose, the runner keeps `AGENT_ROUTING`'s.
484+
485+**Not a default:** the context hub's embeddings model
486+(`@cf/baai/bge-base-en-v1.5`, in `services/context`). Its vectors are
487+only comparable with others from the same model, so changing it means a
488+new index, rebuilt; it is pinned in code, and listed in the catalogue with
489+its price. No other g1t service calls a model.
490+
491+Customers never choose among these: they keep **Auto**, or pin a tier per
492+kind of work, and nothing customer-facing names the catalogue.
493+
347494 ## Discounts
348495
349496 An account's terms are standard, or custom: a **discount** from 1 to 100%,
+21−4
11421142 - **Tiers and catalogue.** `small` (Claude Haiku 5.5 since 2026-10-08,
11431143 $0.10/$0.50 per million input/output up to 100k-token prompts, five
11441144 times that above; it was Haiku 4.5 at $1/$5), `large` (Claude Sonnet 5.5, $2/$10) and `frontier`
1145− (Claude Opus 5.5, $4/$20). Models, names and list prices are
1146− configuration (`AGENT_ROUTING`), never code; prices there are for
1147− estimates only, runs are charged what AI Gateway priced them at.
1145+ (Claude Opus 5.5, $4/$20). Models, names and list prices are data,
1146+ never code: since 2026-10-08 billing's model catalogue
1147+ (`gateway_models`, one row per model g1t can use) and staff's defaults
1148+ in sudo, **Agents & models** (`model_defaults`: each tier's model, the
1149+ harness's background model, the AI Gateway's first Claude, each job's
1150+ starting tier and effort), read by the runner once a minute over
1151+ `AGENT_ROUTING`, which is only the fallback when billing cannot be read.
1152+ Catalogue prices are for estimates only; runs are charged what AI
1153+ Gateway priced them at.
1154+- **Keeping up with new models (built 2026-10-08).** The models service
1155+ lists Anthropic's models (through the AI Gateway) and Workers AI's daily
1156+ and on demand; a new id lands in the catalogue as `new`, priced from a
1157+ maintained table of Anthropic's list prices or Workers AI's listing, or
1158+ unpriced, and staff are emailed. Nothing routes to it, offers it or
1159+ charges for it until staff approve it with its prices. A model a provider
1160+ stops listing is `deprecated`; routing never sends work to a deprecated
1161+ or retired model, falling back to the next model of the tier and saying
1162+ so on the run. Customers keep Auto: the catalogue is staff's. See
1163+ [BILLING_OPERATIONS.md](BILLING_OPERATIONS.md#the-model-catalogue).
11481164 - **Starting tier by job.** Catch-up, answering a question, and reviews of
11491165 at most 10 files and 200 lines touching no sensitive path: small.
11501166 Plans: small at high effort (Haiku 5.5 takes an effort level).
11631179 *Used a fast model (Claude Haiku 5.5): small change, 3 files and 80
11641180 lines.* Effort per kind of job (`effort` in `AGENT_ROUTING`: plan high,
11651181 answer medium, update low) is sent as `CLAUDE_CODE_EFFORT_LEVEL` on
1166− g1t's tiers and named in that line.
1182+ g1t's tiers (never on a model the catalogue says takes none) and named
1183+ in that line, with the catalogue's name for the model.
11671184 - **Chosen instead.** `model_routes` rows to g1t's models name `small`,
11681185 `large` or `frontier`, or nothing for Auto (Integrations → Models).
11691186 A workspace's own Anthropic key with no model named is routed by Auto
+155−0
681681 setCostMapping(mapping: CostMappingInput, by: string): Promise<Result<CostMapping>>;
682682 /** Reads Cloudflare's bill and reconciles now, as the daily run does. */
683683 runCosts(by: string): Promise<Result<CostsRun>>;
684+ /** Agents & models: the catalogue, each purpose's default, and the latest checks. */
685+ models(): Promise<AdminModels>;
686+ /** Approve a model (its prices confirmed), retire it, or restore it. Needs a reason. */
687+ decideModel(
688+ model: string,
689+ decision: "approve" | "retire" | "restore",
690+ details: { name?: string | null; tierHint?: string | null; prices?: ModelPrices | null },
691+ reason: string,
692+ by: string,
693+ ): Promise<Result<CatalogueModel>>;
694+ /** One purpose's default: a model, or for a job its tier and effort. Needs a reason. */
695+ setModelDefault(
696+ purpose: string,
697+ value: { model?: string | null; tier?: string | null; effort?: string | null },
698+ reason: string,
699+ by: string,
700+ ): Promise<Result<ModelDefault>>;
684701 }
685702
686703 /** How much a workspace has earned g1t's trust with money. */
923940 */
924941 /** What the AI Gateway offers on g1t's key, with prices per million tokens. */
925942 gatewayModels(): Promise<GatewayModel[]>;
943+ /** Every purpose's default model as it applies now, and each job's tier and effort. */
944+ modelDefaults(): Promise<ModelDefaults>;
945+ /** What one provider lists now, from the models service's discovery. */
946+ recordDiscovery(provider: string, models: ProviderModel[], by: string, error?: string | null): Promise<DiscoveryResult>;
926947 /**
927948 * Whether a workspace's next AI Gateway request may go to g1t's models:
928949 * fails with `payment_required` and what to do when it is over its spend
12081229 overCacheWrite1hMicros?: number;
12091230 };
12101231
1232+// --- The model catalogue ------------------------------------------------------
1233+
1234+/** Where a model stands: only `available` ones are routed to. */
1235+export type ModelStatus = "available" | "new" | "deprecated" | "retired";
1236+
1237+/** One model in g1t's catalogue: its prices and what g1t knows about it. */
1238+export type CatalogueModel = GatewayModel & {
1239+ /** Other ids the provider lists it by, such as a dated one. */
1240+ aliases: string[];
1241+ family: string;
1242+ /** The agent tier it suits: `small`, `large`, `frontier`, or empty. */
1243+ tierHint: string;
1244+ contextWindow: number;
1245+ maxOutput: number;
1246+ /** Any of `effort`, `thinking`, `tools`, `vision`, `embeddings`. */
1247+ capabilities: string[];
1248+ dimensions: number;
1249+ status: ModelStatus;
1250+ /** Its prices are known. An unpriced model is never routed to, offered or charged for. */
1251+ priced: boolean;
1252+ source: "discovered" | "staff";
1253+ firstSeenAt: string | null;
1254+ lastSeenAt: string | null;
1255+ missingSince: string | null;
1256+ approvedBy: string | null;
1257+ approvedAt: string | null;
1258+ note: string;
1259+ /** A typical agent run on it, in millionths of a dollar; 0 when unpriced or embeddings. */
1260+ typicalRunMicros: number;
1261+};
1262+
1263+/** One model as its provider lists it, from the models service's discovery. */
1264+export type ProviderModel = {
1265+ id: string;
1266+ name: string;
1267+ /** `chat`, `embeddings`, or anything else (counted as listed, never added). */
1268+ kind: string;
1269+ contextWindow: number;
1270+ maxOutput: number;
1271+ capabilities: string[];
1272+ /** Workers AI lists a price per million tokens; Anthropic does not. */
1273+ price: { inputMicros: number; outputMicros: number } | null;
1274+};
1275+
1276+/** What one check of a provider found. */
1277+export type DiscoveryResult = {
1278+ provider: string;
1279+ checkedAt: string;
1280+ by: string;
1281+ listed: number;
1282+ added: string[];
1283+ deprecated: string[];
1284+ restored: string[];
1285+ error: string | null;
1286+};
1287+
1288+/** The purposes a default model is chosen for. */
1289+export type ModelPurpose = "tier_small" | "tier_large" | "tier_frontier" | "background" | "gateway_first";
1290+
1291+/** One purpose's default, as staff last set it. */
1292+export type ModelDefault = {
1293+ /** A `ModelPurpose`, or `job_<kind>`. */
1294+ purpose: string;
1295+ model: string | null;
1296+ /** For a job: `small`, `large`, `frontier` or `change`. */
1297+ tier: string | null;
1298+ effort: string | null;
1299+ updatedAt: string;
1300+ updatedBy: string;
1301+ reason: string;
1302+};
1303+
1304+/** A model purpose's default as it applies now. */
1305+export type ResolvedModel = {
1306+ purpose: string;
1307+ chosen: string;
1308+ /** The model to use; null when nothing suits (callers keep their own fallback). */
1309+ model: GatewayModel | null;
1310+ capabilities: string[];
1311+ /** Why it is not the chosen model, in a sentence. */
1312+ note: string | null;
1313+};
1314+
1315+/** One kind of agent job's starting tier and effort. */
1316+export type JobDefault = { kind: string; tier: string; effort: string | null };
1317+
1318+/** Every purpose's model as it applies now, and each job's tier and effort. */
1319+export type ModelDefaults = { models: ResolvedModel[]; jobs: JobDefault[] };
1320+
1321+/** One check of one provider. */
1322+export type ModelCheck = {
1323+ id: string;
1324+ provider: string;
1325+ checkedAt: string;
1326+ by: string;
1327+ listed: number;
1328+ added: string[];
1329+ deprecated: string[];
1330+ error: string | null;
1331+};
1332+
1333+/** The tokens of the typical agent run estimates are priced from. */
1334+export type TypicalRun = { requests: number; input: number; output: number; cacheRead: number; cacheWrite: number };
1335+
1336+/** sudo's Agents & models. */
1337+export type AdminModels = {
1338+ catalogue: CatalogueModel[];
1339+ defaults: ModelDefault[];
1340+ resolved: ModelDefaults;
1341+ checks: ModelCheck[];
1342+ typical: TypicalRun;
1343+};
1344+
1345+/** A model's prices as staff confirm them, per million tokens in millionths of a dollar. */
1346+export type ModelPrices = {
1347+ inputMicros: number;
1348+ outputMicros: number;
1349+ cacheReadMicros: number;
1350+ cacheWriteMicros: number;
1351+ cacheWrite1hMicros: number;
1352+ threshold: number;
1353+ overInputMicros: number;
1354+ overOutputMicros: number;
1355+ overCacheReadMicros: number;
1356+ overCacheWriteMicros: number;
1357+ overCacheWrite1hMicros: number;
1358+};
1359+
1360+/** The models service's `Discovery` entrypoint, for sudo's "Check for new models". */
1361+export interface ModelDiscoveryApi {
1362+ /** Lists every provider's models now and records what changed: one result per provider. */
1363+ check(by: string): Promise<DiscoveryResult[]>;
1364+}
1365+
12111366 /** The format a gateway request was sent in. */
12121367 export type GatewayFormat = "anthropic" | "openai";
12131368
+36−0
487487 call("token_usage", { workspace, viewer, person: options.person ?? null, days: options.days ?? null }),
488488 recordTokens: (usage) => call("record_tokens", usage),
489489 gatewayModels: () => call("gateway_models", {}),
490+ modelDefaults: () => call("model_defaults", {}),
491+ recordDiscovery: (provider, models, by, error = null) => call("record_discovery", { provider, models, by, error }),
490492 gatewayAdmit: (workspace) => call("gateway_admit", { workspace }),
491493 recordGateway: (record) => call("record_gateway", record),
492494 gatewayRequests: (workspace, viewer, options = {}) =>
619621 by,
620622 }),
621623 runCosts: (by) => call("admin_run_costs", { by }),
624+ models: () => call("admin_models", {}),
625+ decideModel: (model, decision, details, reason, by) =>
626+ call("admin_decide_model", {
627+ model,
628+ decision,
629+ name: details.name ?? null,
630+ tier_hint: details.tierHint ?? null,
631+ prices: details.prices
632+ ? {
633+ input_micros: details.prices.inputMicros,
634+ output_micros: details.prices.outputMicros,
635+ cache_read_micros: details.prices.cacheReadMicros,
636+ cache_write_micros: details.prices.cacheWriteMicros,
637+ cache_write_1h_micros: details.prices.cacheWrite1hMicros,
638+ threshold: details.prices.threshold,
639+ over_input_micros: details.prices.overInputMicros,
640+ over_output_micros: details.prices.overOutputMicros,
641+ over_cache_read_micros: details.prices.overCacheReadMicros,
642+ over_cache_write_micros: details.prices.overCacheWriteMicros,
643+ over_cache_write_1h_micros: details.prices.overCacheWrite1hMicros,
644+ }
645+ : null,
646+ reason,
647+ by,
648+ }),
649+ setModelDefault: (purpose, value, reason, by) =>
650+ call("admin_set_model_default", {
651+ purpose,
652+ model: value.model ?? null,
653+ tier: value.tier ?? null,
654+ effort: value.effort ?? null,
655+ reason,
656+ by,
657+ }),
622658 };
623659 }
624660
+96−0
1+-- One model catalogue: `gateway_models` becomes every model g1t can use,
2+-- the agents' tiers, the AI Gateway's and the embeddings model, with what
3+-- g1t knows about each and where it stands. See src/catalogue.rs.
4+--
5+-- New models are found by the models service listing each provider daily
6+-- (`record_discovery`) and wait as `new` until staff approve them in sudo,
7+-- Agents & models. Only `available` models are routed to; the AI Gateway
8+-- offers `available` and `deprecated` ones that are priced.
9+
10+-- Other ids the provider lists it by, comma separated (a dated id).
11+ALTER TABLE gateway_models ADD COLUMN aliases TEXT NOT NULL DEFAULT '';
12+-- `haiku`, `sonnet`, `opus`, `fable`, or a Workers AI author.
13+ALTER TABLE gateway_models ADD COLUMN family TEXT NOT NULL DEFAULT '';
14+-- The agent tier it suits: `small`, `large`, `frontier`, or empty.
15+ALTER TABLE gateway_models ADD COLUMN tier_hint TEXT NOT NULL DEFAULT '';
16+ALTER TABLE gateway_models ADD COLUMN context_window INTEGER NOT NULL DEFAULT 0;
17+ALTER TABLE gateway_models ADD COLUMN max_output INTEGER NOT NULL DEFAULT 0;
18+-- Comma separated: effort, thinking, tools, vision, embeddings.
19+ALTER TABLE gateway_models ADD COLUMN capabilities TEXT NOT NULL DEFAULT '';
20+-- An embeddings model's vector length.
21+ALTER TABLE gateway_models ADD COLUMN dimensions INTEGER NOT NULL DEFAULT 0;
22+-- `available`, `new` (found, not approved), `deprecated` (the provider
23+-- stopped listing it) or `retired` (staff took it out).
24+ALTER TABLE gateway_models ADD COLUMN status TEXT NOT NULL DEFAULT 'available';
25+-- 1 when its prices are known. An unpriced model is never routed to,
26+-- offered or charged for.
27+ALTER TABLE gateway_models ADD COLUMN priced INTEGER NOT NULL DEFAULT 1;
28+-- `discovered` or `staff`.
29+ALTER TABLE gateway_models ADD COLUMN source TEXT NOT NULL DEFAULT 'staff';
30+ALTER TABLE gateway_models ADD COLUMN first_seen_at TEXT;
31+ALTER TABLE gateway_models ADD COLUMN last_seen_at TEXT;
32+ALTER TABLE gateway_models ADD COLUMN missing_since TEXT;
33+ALTER TABLE gateway_models ADD COLUMN approved_by TEXT;
34+ALTER TABLE gateway_models ADD COLUMN approved_at TEXT;
35+ALTER TABLE gateway_models ADD COLUMN note TEXT NOT NULL DEFAULT '';
36+
37+UPDATE gateway_models SET first_seen_at = updated_at, approved_at = updated_at, approved_by = 'migration';
38+
39+-- What g1t knows about the models it had.
40+UPDATE gateway_models SET family = 'opus', tier_hint = 'frontier', context_window = 1000000, max_output = 128000,
41+ capabilities = 'effort,thinking,tools,vision' WHERE model = 'claude-opus-5-5';
42+UPDATE gateway_models SET family = 'sonnet', tier_hint = 'large', context_window = 1000000, max_output = 128000,
43+ capabilities = 'effort,thinking,tools,vision' WHERE model = 'claude-sonnet-5-5';
44+UPDATE gateway_models SET family = 'haiku', tier_hint = 'small',
45+ capabilities = 'effort,thinking,tools,vision' WHERE model = 'claude-haiku-5-5';
46+-- Haiku 4.5 takes no effort level.
47+UPDATE gateway_models SET family = 'haiku', tier_hint = 'small', context_window = 200000, max_output = 64000,
48+ capabilities = 'thinking,tools,vision' WHERE model IN ('claude-haiku-4-5', 'claude-haiku-4-5-20251001');
49+UPDATE gateway_models SET aliases = 'claude-haiku-4-5-20251001' WHERE model = 'claude-haiku-4-5';
50+UPDATE gateway_models SET family = substr(model, 5, instr(substr(model, 5), '/') - 1) WHERE provider = 'workers-ai';
51+UPDATE gateway_models SET capabilities = 'embeddings', dimensions = 768 WHERE model = '@cf/baai/bge-base-en-v1.5';
52+UPDATE gateway_models SET capabilities = 'embeddings', dimensions = 1024 WHERE model = '@cf/baai/bge-m3';
53+
54+-- Which model each purpose uses by default, chosen by staff in sudo
55+-- (`admin_set_model_default`; every change in the audit log with the old
56+-- and new value and why). A model purpose names a model; a job
57+-- (`job_<kind>`) its starting tier (`small`, `large`, `frontier`, or
58+-- `change` to size the change it reads) and effort. Read by the runner and
59+-- the AI Gateway; the runner's AGENT_ROUTING is only the fallback when
60+-- billing cannot be reached. Seeded with what AGENT_ROUTING said.
61+CREATE TABLE IF NOT EXISTS model_defaults (
62+ purpose TEXT PRIMARY KEY,
63+ model TEXT,
64+ tier TEXT,
65+ effort TEXT,
66+ updated_at TEXT NOT NULL,
67+ updated_by TEXT NOT NULL,
68+ reason TEXT NOT NULL DEFAULT ''
69+);
70+
71+INSERT OR IGNORE INTO model_defaults (purpose, model, tier, effort, updated_at, updated_by, reason) VALUES
72+ ('tier_small', 'claude-haiku-5-5', NULL, NULL, '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
73+ ('tier_large', 'claude-sonnet-5-5', NULL, NULL, '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
74+ ('tier_frontier', 'claude-opus-5-5', NULL, NULL, '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
75+ ('background', 'claude-haiku-5-5', NULL, NULL, '2026-10-08T00:00:00Z', 'migration', 'The fast tier''s model, as before'),
76+ ('gateway_first', 'claude-haiku-5-5', NULL, NULL, '2026-10-08T00:00:00Z', 'migration', 'The cheapest Claude, listed first since 0047'),
77+ ('job_implement', NULL, 'large', NULL, '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
78+ ('job_revise', NULL, 'large', NULL, '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
79+ ('job_answer', NULL, 'small', 'medium', '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
80+ ('job_review', NULL, 'change', NULL, '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
81+ ('job_update', NULL, 'small', 'low', '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it'),
82+ ('job_plan', NULL, 'small', 'high', '2026-10-08T00:00:00Z', 'migration', 'As AGENT_ROUTING had it');
83+
84+-- Every check of a provider's list, daily or from sudo: what it listed,
85+-- added and found gone, or why it failed. Kept 90 days.
86+CREATE TABLE IF NOT EXISTS model_checks (
87+ id TEXT PRIMARY KEY,
88+ provider TEXT NOT NULL,
89+ checked_at TEXT NOT NULL,
90+ by TEXT NOT NULL,
91+ listed INTEGER NOT NULL DEFAULT 0,
92+ added TEXT NOT NULL DEFAULT '',
93+ deprecated TEXT NOT NULL DEFAULT '',
94+ error TEXT
95+);
96+CREATE INDEX IF NOT EXISTS model_checks_by_time ON model_checks (checked_at);
+1051−0
1+//! The model catalogue: every model g1t can use, where each stands, and
2+//! which one each purpose uses by default.
3+//!
4+//! - **One table.** `gateway_models` (migrations 0045, 0047, 0048) holds
5+//! every model: the agents' tiers, the AI Gateway's models and the
6+//! embeddings model, with prices per million tokens by kind (long-prompt
7+//! and cache prices included), and its aliases, family, tier hint,
8+//! context window, capabilities, status and where it came from. Billing
9+//! owns it because billing owns prices: the gateway charges from it, and
10+//! a second table of models would drift from the first.
11+//! - **Discovery.** The models service lists each provider daily (and when
12+//! staff press "Check for new models" in sudo): Anthropic's
13+//! `GET /v1/models` through g1t's AI Gateway, and Workers AI's model
14+//! search. Listing is free; nothing here calls a paid model.
15+//! `record_discovery` compares the list with the catalogue (`diff`):
16+//! an id it has never seen is added as `new` (priced from `known` or the
17+//! listing's own price when either knows it, otherwise unpriced); a dated
18+//! id of a model it has (`claude-haiku-4-5-20251001`) is that model; a
19+//! model the provider stopped listing becomes `deprecated`, and comes
20+//! back when it is listed again. Staff are emailed about anything new or
21+//! gone, with a link to sudo. A `new` model is never routed to, offered
22+//! or charged for until staff approve it with its prices.
23+//! - **Defaults.** `model_defaults` holds staff's choice per purpose: the
24+//! model behind each agent tier, the harness's background model, the AI
25+//! Gateway's first Claude, and each job's starting tier and effort. Every
26+//! change is audited with the old value, the new one and why.
27+//! `resolve` never hands out a model that cannot be used: a chosen model
28+//! that is deprecated, retired or unpriced falls back to the next
29+//! available model suited to the purpose, with a sentence saying so.
30+
31+use g1t_contracts::billing::{
32+ model_status, AdminDecideModelArgs, AdminModels, AdminSetModelDefaultArgs, CatalogueModel, DiscoveryResult, GatewayModel,
33+ JobDefault, ModelCheck, ModelDefault, ModelDefaults, ModelPrices, ProviderModel, RecordDiscoveryArgs, ResolvedModel,
34+ TypicalRun,
35+};
36+use g1t_contracts::time::rfc3339;
37+use g1t_contracts::{FailureCode, Outcome, new_id};
38+use g1t_kit::now_ms;
39+use serde::Deserialize;
40+use worker::Result;
41+use worker::wasm_bindgen::JsValue;
42+
43+use crate::gateway::{ModelRow, Used, cost_micros};
44+use crate::Billing;
45+
46+/// The tokens of a typical agent run, which sudo prices each model at: 40
47+/// requests, each reading most of its context from the prompt cache, as a
48+/// Sonnet implement run does (about 90% of its tokens cache reads). An
49+/// estimate for comparing models, never a charge.
50+pub(crate) const TYPICAL: TypicalRun = TypicalRun { requests: 40, input: 2_000, output: 1_500, cache_read: 45_000, cache_write: 4_000 };
51+
52+/// What `TYPICAL` costs on a model, in millionths of a dollar: 0 for an
53+/// embeddings model. Each request is priced on its own, so a model priced
54+/// by prompt length is at its lower prices unless one request's prompt is
55+/// over the threshold.
56+pub(crate) fn typical_run(model: &GatewayModel) -> i64 {
57+ if model.kind != "chat" {
58+ return 0;
59+ }
60+ let used = Used { input: TYPICAL.input, output: TYPICAL.output, cache_read: TYPICAL.cache_read, cache_write: TYPICAL.cache_write, cache_write_1h: 0 };
61+ cost_micros(model, &used).saturating_mul(TYPICAL.requests as i64)
62+}
63+
64+/// The purposes a default model is chosen for, in the order sudo shows them.
65+pub(crate) const MODEL_PURPOSES: [&str; 5] = ["tier_small", "tier_large", "tier_frontier", "background", "gateway_first"];
66+/// The kinds of agent job, each with a starting tier and effort (`job_<kind>`).
67+pub(crate) const JOB_KINDS: [&str; 6] = ["implement", "revise", "answer", "review", "update", "plan"];
68+pub(crate) const TIERS: [&str; 3] = ["small", "large", "frontier"];
69+pub(crate) const EFFORTS: [&str; 5] = ["low", "medium", "high", "xhigh", "max"];
70+
71+/// What sudo and the audit log call a purpose.
72+pub(crate) fn purpose_label(purpose: &str) -> String {
73+ match purpose {
74+ "tier_small" => "Fast tier".to_owned(),
75+ "tier_large" => "Standard tier".to_owned(),
76+ "tier_frontier" => "Most capable tier".to_owned(),
77+ "background" => "Background model".to_owned(),
78+ "gateway_first" => "AI Gateway's first Claude".to_owned(),
79+ other => match other.strip_prefix("job_") {
80+ Some(kind) => format!("Job: {kind}"),
81+ None => other.to_owned(),
82+ },
83+ }
84+}
85+
86+/// The tier a model purpose is for, if it is one of the agent tiers or
87+/// the background model (which is a fast model's job).
88+fn tier_of_purpose(purpose: &str) -> Option<&'static str> {
89+ match purpose {
90+ "tier_small" | "background" => Some("small"),
91+ "tier_large" => Some("large"),
92+ "tier_frontier" => Some("frontier"),
93+ _ => None,
94+ }
95+}
96+
97+// ---------------------------------------------------------------------------
98+// What g1t knows about models before anyone tells it.
99+// ---------------------------------------------------------------------------
100+
101+/// Prices per million tokens in millionths of a dollar: input, output,
102+/// cache read, five-minute and hour-long cache write.
103+type Five = [i64; 5];
104+
105+const fn five(input: i64, output: i64, read: i64, write: i64, write_1h: i64) -> Five {
106+ [input, output, read, write, write_1h]
107+}
108+
109+/// A model id, its prices, and its long-prompt threshold and prices above it.
110+type Listed = (&'static str, Five, Option<(u64, Five)>);
111+
112+/// Anthropic's list prices by model, as published (checked 2026-10-08).
113+/// The maintained price table: when Anthropic lists a model here, it is
114+/// added priced and only needs staff to confirm it. Add a row when
115+/// Anthropic publishes a new model's price; until then a new model is
116+/// added unpriced and staff enter its prices when they approve it.
117+const ANTHROPIC_PRICES: &[Listed] = &[
118+ ("claude-fable-5-1", five(10_000_000, 50_000_000, 250_000, 12_500_000, 20_000_000), None),
119+ ("claude-fable-5", five(10_000_000, 50_000_000, 1_000_000, 12_500_000, 20_000_000), None),
120+ ("claude-opus-5-5", five(4_000_000, 20_000_000, 200_000, 5_000_000, 8_000_000), None),
121+ ("claude-opus-5", five(5_000_000, 25_000_000, 500_000, 6_250_000, 10_000_000), None),
122+ ("claude-opus-4-8", five(5_000_000, 25_000_000, 500_000, 6_250_000, 10_000_000), None),
123+ ("claude-opus-4-7", five(5_000_000, 25_000_000, 500_000, 6_250_000, 10_000_000), None),
124+ ("claude-opus-4-6", five(5_000_000, 25_000_000, 500_000, 6_250_000, 10_000_000), None),
125+ ("claude-sonnet-5-5", five(2_000_000, 10_000_000, 100_000, 2_500_000, 4_000_000), None),
126+ ("claude-sonnet-5", five(2_000_000, 10_000_000, 200_000, 2_500_000, 4_000_000), None),
127+ ("claude-sonnet-4-6", five(3_000_000, 15_000_000, 300_000, 3_750_000, 6_000_000), None),
128+ (
129+ "claude-haiku-5-5",
130+ five(100_000, 500_000, 10_000, 125_000, 200_000),
131+ Some((100_000, five(500_000, 2_500_000, 50_000, 625_000, 1_000_000))),
132+ ),
133+ ("claude-haiku-4-5", five(1_000_000, 5_000_000, 100_000, 1_250_000, 2_000_000), None),
134+];
135+
136+/// Whether `listed` is `model` with a date after it (`-YYYYMMDD`), as
137+/// Anthropic lists a model it also names without one.
138+pub(crate) fn is_dated(listed: &str, model: &str) -> bool {
139+ listed
140+ .strip_prefix(model)
141+ .and_then(|rest| rest.strip_prefix('-'))
142+ .is_some_and(|date| date.len() == 8 && date.bytes().all(|b| b.is_ascii_digit()))
143+}
144+
145+/// The list price of an Anthropic model (by its id, or its dated id), if
146+/// the table has it.
147+pub(crate) fn known_prices(id: &str) -> Option<ModelPrices> {
148+ let (_, base, over) = ANTHROPIC_PRICES.iter().find(|(model, _, _)| id == *model || is_dated(id, model))?;
149+ let (threshold, above) = over.unwrap_or((0, [0; 5]));
150+ Some(ModelPrices {
151+ input_micros: base[0],
152+ output_micros: base[1],
153+ cache_read_micros: base[2],
154+ cache_write_micros: base[3],
155+ cache_write_1h_micros: base[4],
156+ threshold,
157+ over_input_micros: above[0],
158+ over_output_micros: above[1],
159+ over_cache_read_micros: above[2],
160+ over_cache_write_micros: above[3],
161+ over_cache_write_1h_micros: above[4],
162+ })
163+}
164+
165+/// A model's family, and the agent tier it suits, from its id:
166+/// `claude-haiku-*` is fast, `claude-sonnet-*` standard, `claude-opus-*`
167+/// and `claude-fable-*` the most capable. A Workers AI model's family is
168+/// its author (`@cf/<author>/…`) and it suits no tier.
169+pub(crate) fn family_of(provider: &str, id: &str) -> (String, String) {
170+ if provider == "anthropic" {
171+ for (family, tier) in [("haiku", "small"), ("sonnet", "large"), ("opus", "frontier"), ("fable", "frontier"), ("mythos", "frontier")] {
172+ if id.starts_with(&format!("claude-{family}")) {
173+ return (family.to_owned(), tier.to_owned());
174+ }
175+ }
176+ return (String::new(), String::new());
177+ }
178+ let author = id.trim_start_matches('@').split('/').nth(1).unwrap_or_default();
179+ (author.to_owned(), String::new())
180+}
181+
182+/// A name for people when the provider gives none: a Workers AI id's last
183+/// part.
184+fn name_of(listed: &ProviderModel) -> String {
185+ let name = listed.name.trim();
186+ if !name.is_empty() && !name.starts_with('@') {
187+ return name.chars().take(120).collect();
188+ }
189+ listed.id.rsplit('/').next().unwrap_or(&listed.id).chars().take(120).collect()
190+}
191+
192+// ---------------------------------------------------------------------------
193+// Discovery: what a provider lists, against the catalogue.
194+// ---------------------------------------------------------------------------
195+
196+/// What a check found, before it is written.
197+#[derive(Debug, Default, PartialEq)]
198+pub(crate) struct Diff {
199+ /// Ids never seen, to add as `new` (chat and embeddings models only).
200+ pub added: Vec<ProviderModel>,
201+ /// Catalogue models listed, with an id to add to their aliases when the
202+ /// provider listed them by one they did not have.
203+ pub seen: Vec<(String, Option<String>)>,
204+ /// Catalogue models (available or new) the provider no longer lists.
205+ pub gone: Vec<String>,
206+ /// Deprecated models listed again.
207+ pub restored: Vec<String>,
208+}
209+
210+/// Whether the provider listing `listed` lists catalogue model `row`, and
211+/// by an id it did not know (to keep as an alias).
212+fn lists(row: &CatalogueModel, listed: &str) -> Option<Option<String>> {
213+ if listed == row.prices.model || row.aliases.iter().any(|alias| alias == listed) {
214+ return Some(None);
215+ }
216+ is_dated(listed, &row.prices.model).then(|| Some(listed.to_owned()))
217+}
218+
219+/// Compares what `provider` lists with the catalogue. An empty list says
220+/// nothing (a failed or empty answer is not every model gone), so nothing
221+/// is found gone then. Retired models are left as they are.
222+pub(crate) fn diff(provider: &str, catalogue: &[CatalogueModel], listed: &[ProviderModel]) -> Diff {
223+ let mine: Vec<&CatalogueModel> = catalogue.iter().filter(|row| row.prices.provider == provider).collect();
224+ let mut out = Diff::default();
225+ let mut matched: Vec<&str> = vec![];
226+ for model in listed {
227+ let id = model.id.trim();
228+ if id.is_empty() {
229+ continue;
230+ }
231+ let hits: Vec<(&CatalogueModel, Option<String>)> = mine.iter().filter_map(|row| lists(row, id).map(|alias| (*row, alias))).collect();
232+ if hits.is_empty() {
233+ let kind = model.kind.as_str();
234+ if (kind == "chat" || kind == "embeddings") && !out.added.iter().any(|added| added.id == id) {
235+ out.added.push(ProviderModel { id: id.to_owned(), ..model.clone() });
236+ }
237+ continue;
238+ }
239+ for (row, alias) in hits {
240+ let name = row.prices.model.as_str();
241+ if matched.contains(&name) {
242+ continue;
243+ }
244+ matched.push(name);
245+ if row.status == model_status::DEPRECATED && row.missing_since.is_some() {
246+ out.restored.push(name.to_owned());
247+ }
248+ out.seen.push((name.to_owned(), alias));
249+ }
250+ }
251+ if listed.iter().any(|model| !model.id.trim().is_empty()) {
252+ for row in mine {
253+ let gone = !matched.contains(&row.prices.model.as_str());
254+ if gone && (row.status == model_status::AVAILABLE || row.status == model_status::NEW) {
255+ out.gone.push(row.prices.model.clone());
256+ }
257+ }
258+ }
259+ out
260+}
261+
262+/// The row a newly found model is added as: `new`, from discovery, priced
263+/// from `known_prices` (Anthropic), else the listing's own price (Workers
264+/// AI, which has no cache prices: cached tokens cost what input does),
265+/// else unpriced.
266+pub(crate) fn new_row(provider: &str, listed: &ProviderModel, position: i64) -> (GatewayModel, CatalogueFacts) {
267+ let (family, tier) = family_of(provider, &listed.id);
268+ let kind = if listed.kind == "embeddings" { "embeddings" } else { "chat" };
269+ let prices = known_prices(&listed.id).or_else(|| {
270+ listed.price.as_ref().filter(|price| price.input_micros > 0).map(|price| ModelPrices {
271+ input_micros: price.input_micros,
272+ output_micros: price.output_micros.max(0),
273+ cache_read_micros: price.input_micros,
274+ cache_write_micros: price.input_micros,
275+ cache_write_1h_micros: price.input_micros,
276+ ..ModelPrices::default()
277+ })
278+ });
279+ let priced = prices.is_some();
280+ let prices = prices.unwrap_or_default();
281+ let mut capabilities = listed.capabilities.clone();
282+ if kind == "embeddings" && !capabilities.iter().any(|c| c == "embeddings") {
283+ capabilities.push("embeddings".to_owned());
284+ }
285+ (
286+ with_prices(
287+ GatewayModel {
288+ model: listed.id.clone(),
289+ name: name_of(listed),
290+ provider: provider.to_owned(),
291+ kind: kind.to_owned(),
292+ input_micros: 0,
293+ output_micros: 0,
294+ cache_read_micros: 0,
295+ cache_write_micros: 0,
296+ cache_write_1h_micros: 0,
297+ threshold: 0,
298+ over_input_micros: 0,
299+ over_output_micros: 0,
300+ over_cache_read_micros: 0,
301+ over_cache_write_micros: 0,
302+ over_cache_write_1h_micros: 0,
303+ },
304+ &prices,
305+ ),
306+ CatalogueFacts { family, tier_hint: tier, capabilities, priced, position },
307+ )
308+}
309+
310+/// What a new row carries besides its prices.
311+#[derive(Debug, PartialEq)]
312+pub(crate) struct CatalogueFacts {
313+ pub family: String,
314+ pub tier_hint: String,
315+ pub capabilities: Vec<String>,
316+ pub priced: bool,
317+ pub position: i64,
318+}
319+
320+/// `model` with `prices`.
321+pub(crate) fn with_prices(model: GatewayModel, prices: &ModelPrices) -> GatewayModel {
322+ GatewayModel {
323+ input_micros: prices.input_micros,
324+ output_micros: prices.output_micros,
325+ cache_read_micros: prices.cache_read_micros,
326+ cache_write_micros: prices.cache_write_micros,
327+ cache_write_1h_micros: prices.cache_write_1h_micros,
328+ threshold: prices.threshold,
329+ over_input_micros: prices.over_input_micros,
330+ over_output_micros: prices.over_output_micros,
331+ over_cache_read_micros: prices.over_cache_read_micros,
332+ over_cache_write_micros: prices.over_cache_write_micros,
333+ over_cache_write_1h_micros: prices.over_cache_write_1h_micros,
334+ ..model
335+ }
336+}
337+
338+/// Whether prices staff confirm make sense: none below nothing, a chat
339+/// model with input and output prices, an embeddings model with an input
340+/// price, and long-prompt prices only with a threshold.
341+pub(crate) fn check_prices(kind: &str, prices: &ModelPrices) -> std::result::Result<(), String> {
342+ let all = [
343+ prices.input_micros,
344+ prices.output_micros,
345+ prices.cache_read_micros,
346+ prices.cache_write_micros,
347+ prices.cache_write_1h_micros,
348+ prices.over_input_micros,
349+ prices.over_output_micros,
350+ prices.over_cache_read_micros,
351+ prices.over_cache_write_micros,
352+ prices.over_cache_write_1h_micros,
353+ ];
354+ if all.iter().any(|price| *price < 0) {
355+ return Err("A price cannot be less than nothing.".into());
356+ }
357+ // $1,000 per million tokens: no model costs that; a slipped finger does.
358+ if all.iter().any(|price| *price > 1_000_000_000) {
359+ return Err("No price is over $1,000 per million tokens; check the decimal point.".into());
360+ }
361+ if prices.input_micros == 0 {
362+ return Err("Give the input price per million tokens.".into());
363+ }
364+ if kind == "chat" && prices.output_micros == 0 {
365+ return Err("Give the output price per million tokens.".into());
366+ }
367+ let over = prices.over_input_micros + prices.over_output_micros + prices.over_cache_read_micros + prices.over_cache_write_micros;
368+ if prices.threshold == 0 && over > 0 {
369+ return Err("Long-prompt prices need the prompt length they start above.".into());
370+ }
371+ if prices.threshold > 0 && (prices.over_input_micros == 0 || (kind == "chat" && prices.over_output_micros == 0)) {
372+ return Err("With a long-prompt threshold, give the input and output prices above it.".into());
373+ }
374+ Ok(())
375+}
376+
377+// ---------------------------------------------------------------------------
378+// Defaults: what each purpose uses, and what it falls back to.
379+// ---------------------------------------------------------------------------
380+
381+/// Why `model` cannot serve `purpose`, or None when it can. The agents'
382+/// purposes need Claude (the harness speaks Anthropic's API), and so does
383+/// the AI Gateway's first model.
384+pub(crate) fn unsuited(purpose: &str, model: &CatalogueModel) -> Option<String> {
385+ let name = &model.prices.name;
386+ if !model.priced {
387+ return Some(format!("{name} has no price yet: approve it with its prices first."));
388+ }
389+ match model.status.as_str() {
390+ model_status::AVAILABLE => {}
391+ model_status::NEW => return Some(format!("{name} is new: approve it first.")),
392+ status => return Some(format!("{name} is {status}.")),
393+ }
394+ if MODEL_PURPOSES.contains(&purpose) && (model.prices.provider != "anthropic" || model.prices.kind != "chat") {
395+ return Some(format!("{name} is not a Claude chat model, which {} needs.", purpose_label(purpose).to_lowercase()));
396+ }
397+ None
398+}
399+
400+/// What one model purpose uses now: the chosen model while it can be used;
401+/// otherwise the first available Claude suited to the purpose's tier (or,
402+/// for the gateway, any available Claude), in the catalogue's order, with a
403+/// sentence saying so. A model is never handed out that cannot be used.
404+pub(crate) fn resolve(purpose: &str, chosen: &str, catalogue: &[CatalogueModel]) -> ResolvedModel {
405+ let resolved = |model: &CatalogueModel, note: Option<String>| ResolvedModel {
406+ purpose: purpose.to_owned(),
407+ chosen: chosen.to_owned(),
408+ model: Some(model.prices.clone()),
409+ capabilities: model.capabilities.clone(),
410+ note,
411+ };
412+ let picked = catalogue.iter().find(|model| model.prices.model == chosen);
413+ if let Some(model) = picked
414+ && unsuited(purpose, model).is_none()
415+ {
416+ return resolved(model, None);
417+ }
418+ let why = match picked {
419+ Some(model) => match model.status.as_str() {
420+ model_status::AVAILABLE if !model.priced => format!("{} has no price", model.prices.name),
421+ model_status::AVAILABLE => format!("{} does not suit it", model.prices.name),
422+ status => format!("{} is {status}", model.prices.name),
423+ },
424+ None => format!("{chosen} is not in the catalogue"),
425+ };
426+ let tier = tier_of_purpose(purpose);
427+ let fallback = catalogue.iter().filter(|model| model.prices.model != chosen && unsuited(purpose, model).is_none()).find(|model| {
428+ // Same tier first; a purpose with no tier takes any Claude.
429+ tier.is_none_or(|tier| model.tier_hint == tier)
430+ });
431+ match fallback {
432+ Some(model) => resolved(model, Some(format!("{why}; using {} instead.", model.prices.name))),
433+ None => ResolvedModel {
434+ purpose: purpose.to_owned(),
435+ chosen: chosen.to_owned(),
436+ model: None,
437+ capabilities: vec![],
438+ note: Some(format!("{why}, and no other model suits it.")),
439+ },
440+ }
441+}
442+
443+/// Every purpose's model as it applies now, and each job's tier and
444+/// effort; a purpose or job with no row is left out (callers keep theirs).
445+pub(crate) fn defaults_view(defaults: &[ModelDefault], catalogue: &[CatalogueModel]) -> ModelDefaults {
446+ let models = MODEL_PURPOSES
447+ .iter()
448+ .filter_map(|purpose| {
449+ let row = defaults.iter().find(|row| row.purpose == *purpose)?;
450+ Some(resolve(purpose, row.model.as_deref()?, catalogue))
451+ })
452+ .collect();
453+ let jobs = JOB_KINDS
454+ .iter()
455+ .filter_map(|kind| {
456+ let row = defaults.iter().find(|row| row.purpose == format!("job_{kind}"))?;
457+ let tier = row.tier.as_deref().filter(|tier| TIERS.contains(tier) || *tier == "change")?;
458+ Some(JobDefault {
459+ kind: (*kind).to_owned(),
460+ tier: tier.to_owned(),
461+ effort: row.effort.clone().filter(|effort| EFFORTS.contains(&effort.as_str())),
462+ })
463+ })
464+ .collect();
465+ ModelDefaults { models, jobs }
466+}
467+
468+/// Puts `first` at the top of the gateway's list, the rest in order.
469+pub(crate) fn put_first(models: &mut Vec<GatewayModel>, first: Option<&str>) {
470+ if let Some(at) = first.and_then(|first| models.iter().position(|model| model.model == first)) {
471+ let model = models.remove(at);
472+ models.insert(0, model);
473+ }
474+}
475+
476+/// A default as staff set it, checked: the purpose exists, a model purpose
477+/// names a model that suits it, a job names a tier (and an effort or none).
478+/// The reason is required.
479+/// A checked default: its model, or its tier and effort.
480+pub(crate) type Checked = (Option<String>, Option<String>, Option<String>);
481+
482+pub(crate) fn check_default(a: &AdminSetModelDefaultArgs, catalogue: &[CatalogueModel]) -> std::result::Result<Checked, String> {
483+ if a.reason.trim().is_empty() {
484+ return Err("Say why, for whoever looks next.".into());
485+ }
486+ let purpose = a.purpose.as_str();
487+ if MODEL_PURPOSES.contains(&purpose) {
488+ let wanted = a.model.as_deref().map(str::trim).filter(|m| !m.is_empty()).ok_or("Choose a model.")?;
489+ let model = catalogue.iter().find(|model| model.prices.model == wanted).ok_or_else(|| format!("{wanted} is not in the catalogue."))?;
490+ if let Some(why) = unsuited(purpose, model) {
491+ return Err(why);
492+ }
493+ return Ok((Some(wanted.to_owned()), None, None));
494+ }
495+ let Some(kind) = purpose.strip_prefix("job_").filter(|kind| JOB_KINDS.contains(kind)) else {
496+ return Err(format!("{purpose} is not something a default is chosen for."));
497+ };
498+ let tier = a.tier.as_deref().map(str::trim).unwrap_or_default();
499+ // Only a review is sized by the change it reads.
500+ if !(TIERS.contains(&tier) || (tier == "change" && kind == "review")) {
501+ return Err(if kind == "review" { "Choose fast, standard, most capable, or by the change's size." } else { "Choose fast, standard or most capable." }.into());
502+ }
503+ let effort = a.effort.as_deref().map(str::trim).filter(|effort| !effort.is_empty());
504+ if let Some(effort) = effort
505+ && !EFFORTS.contains(&effort)
506+ {
507+ return Err("Effort is low, medium, high, xhigh or max, or the harness's own.".into());
508+ }
509+ Ok((None, Some(tier.to_owned()), effort.map(str::to_owned)))
510+}
511+
512+/// How a default reads in the audit log: `claude-haiku-5-5`, or
513+/// `small at high effort`.
514+fn describe_default(model: Option<&str>, tier: Option<&str>, effort: Option<&str>) -> String {
515+ match (model, tier) {
516+ (Some(model), _) => model.to_owned(),
517+ (None, Some(tier)) => match effort {
518+ Some(effort) => format!("{tier} at {effort} effort"),
519+ None => tier.to_owned(),
520+ },
521+ (None, None) => "nothing".to_owned(),
522+ }
523+}
524+
525+fn split(list: Option<&str>) -> Vec<String> {
526+ list.unwrap_or_default().split(',').map(str::trim).filter(|s| !s.is_empty()).map(str::to_owned).collect()
527+}
528+
529+impl From<ModelRow> for CatalogueModel {
530+ fn from(row: ModelRow) -> Self {
531+ let n = |value: Option<f64>| value.unwrap_or(0.0).max(0.0) as u64;
532+ let aliases = split(row.aliases.as_deref());
533+ let capabilities = split(row.capabilities.as_deref());
534+ let (family, tier_hint) = (row.family.clone().unwrap_or_default(), row.tier_hint.clone().unwrap_or_default());
535+ let status = row.status.clone().unwrap_or_else(|| model_status::AVAILABLE.to_owned());
536+ let priced = row.priced.is_none_or(|priced| priced != 0.0);
537+ let source = row.source.clone().unwrap_or_else(|| "staff".to_owned());
538+ let (context_window, max_output, dimensions) = (n(row.context_window), n(row.max_output), n(row.dimensions) as u32);
539+ let (first_seen_at, last_seen_at, missing_since) = (row.first_seen_at.clone(), row.last_seen_at.clone(), row.missing_since.clone());
540+ let (approved_by, approved_at, note) = (row.approved_by.clone(), row.approved_at.clone(), row.note.clone().unwrap_or_default());
541+ let prices = GatewayModel::from(row);
542+ let typical_run_micros = if priced { typical_run(&prices) } else { 0 };
543+ CatalogueModel {
544+ prices,
545+ aliases,
546+ family,
547+ tier_hint,
548+ context_window,
549+ max_output,
550+ capabilities,
551+ dimensions,
552+ status,
553+ priced,
554+ source,
555+ first_seen_at,
556+ last_seen_at,
557+ missing_since,
558+ approved_by,
559+ approved_at,
560+ note,
561+ typical_run_micros,
562+ }
563+ }
564+}
565+
566+#[derive(Deserialize)]
567+struct DefaultRow {
568+ purpose: String,
569+ model: Option<String>,
570+ tier: Option<String>,
571+ effort: Option<String>,
572+ updated_at: String,
573+ updated_by: String,
574+ reason: Option<String>,
575+}
576+
577+impl From<DefaultRow> for ModelDefault {
578+ fn from(row: DefaultRow) -> Self {
579+ ModelDefault {
580+ purpose: row.purpose,
581+ model: row.model,
582+ tier: row.tier,
583+ effort: row.effort,
584+ updated_at: row.updated_at,
585+ updated_by: row.updated_by,
586+ reason: row.reason.unwrap_or_default(),
587+ }
588+ }
589+}
590+
591+#[derive(Deserialize)]
592+struct CheckRow {
593+ id: String,
594+ provider: String,
595+ checked_at: String,
596+ by: String,
597+ listed: f64,
598+ added: Option<String>,
599+ deprecated: Option<String>,
600+ error: Option<String>,
601+}
602+
603+impl From<CheckRow> for ModelCheck {
604+ fn from(row: CheckRow) -> Self {
605+ ModelCheck {
606+ id: row.id,
607+ provider: row.provider,
608+ checked_at: row.checked_at,
609+ by: row.by,
610+ listed: row.listed.max(0.0) as u32,
611+ added: split(row.added.as_deref()),
612+ deprecated: split(row.deprecated.as_deref()),
613+ error: row.error,
614+ }
615+ }
616+}
617+
618+/// How long checks are kept.
619+const CHECK_DAYS: u64 = 90;
620+const DAY_MS: u64 = 86_400_000;
621+
622+fn number(n: u64) -> JsValue {
623+ JsValue::from_f64(n as f64)
624+}
625+
626+impl Billing {
627+ /// Every model in the catalogue, whatever its status: `new` ones first,
628+ /// then by provider and position.
629+ pub(crate) async fn catalogue(&self) -> Result<Vec<CatalogueModel>> {
630+ Ok(self
631+ .db
632+ .prepare("SELECT * FROM gateway_models ORDER BY CASE status WHEN 'new' THEN 0 ELSE 1 END, provider, position, model")
633+ .all()
634+ .await?
635+ .results::<ModelRow>()?
636+ .into_iter()
637+ .map(CatalogueModel::from)
638+ .collect())
639+ }
640+
641+ pub(crate) async fn model_default(&self, purpose: &str) -> Result<Option<ModelDefault>> {
642+ Ok(self
643+ .db
644+ .prepare("SELECT * FROM model_defaults WHERE purpose = ?")
645+ .bind(&[purpose.into()])?
646+ .first::<DefaultRow>(None)
647+ .await?
648+ .map(ModelDefault::from))
649+ }
650+
651+ async fn model_default_rows(&self) -> Result<Vec<ModelDefault>> {
652+ Ok(self
653+ .db
654+ .prepare("SELECT * FROM model_defaults ORDER BY purpose")
655+ .all()
656+ .await?
657+ .results::<DefaultRow>()?
658+ .into_iter()
659+ .map(ModelDefault::from)
660+ .collect())
661+ }
662+
663+ /// `model_defaults`: what the runner and the gateway use now.
664+ pub(crate) async fn model_defaults(&self) -> Result<ModelDefaults> {
665+ let (defaults, catalogue) = futures_util::future::try_join(self.model_default_rows(), self.catalogue()).await?;
666+ Ok(defaults_view(&defaults, &catalogue))
667+ }
668+
669+ /// `admin_models`.
670+ pub(crate) async fn admin_models(&self) -> Result<AdminModels> {
671+ let (defaults, catalogue) = futures_util::future::try_join(self.model_default_rows(), self.catalogue()).await?;
672+ let checks = self
673+ .db
674+ .prepare("SELECT * FROM model_checks ORDER BY checked_at DESC, id DESC LIMIT 20")
675+ .all()
676+ .await?
677+ .results::<CheckRow>()?
678+ .into_iter()
679+ .map(ModelCheck::from)
680+ .collect();
681+ let resolved = defaults_view(&defaults, &catalogue);
682+ Ok(AdminModels { catalogue, defaults, resolved, checks, typical: TYPICAL })
683+ }
684+
685+ /// `record_discovery`: what one provider lists, against the catalogue.
686+ pub(crate) async fn record_discovery(&self, a: RecordDiscoveryArgs) -> Result<DiscoveryResult> {
687+ let provider: String = a.provider.trim().chars().take(40).collect();
688+ let by: String = a.by.trim().chars().take(200).collect();
689+ let now = now_ms();
690+ let checked_at = rfc3339(now);
691+ let mut result = DiscoveryResult {
692+ provider: provider.clone(),
693+ checked_at: checked_at.clone(),
694+ by: by.clone(),
695+ listed: a.models.len().min(u32::MAX as usize) as u32,
696+ error: a.error.as_deref().map(|e| e.chars().take(500).collect()),
697+ ..DiscoveryResult::default()
698+ };
699+ if result.error.is_none() {
700+ let catalogue = self.catalogue().await?;
701+ let found = diff(&provider, &catalogue, &a.models);
702+ let mut writes = vec![];
703+ let position = catalogue.iter().filter(|row| row.prices.provider == provider).count() as i64 + 100;
704+ for (at, listed) in found.added.iter().enumerate() {
705+ let (model, facts) = new_row(&provider, listed, position + at as i64);
706+ writes.push(
707+ self.db
708+ .prepare(
709+ "INSERT OR IGNORE INTO gateway_models
710+ (model, name, provider, kind, input_micros, output_micros, cache_read_micros, cache_write_micros,
711+ cache_write_1h_micros, threshold, over_input_micros, over_output_micros, over_cache_read_micros,
712+ over_cache_write_micros, over_cache_write_1h_micros, position, updated_at, family, tier_hint,
713+ context_window, max_output, capabilities, status, priced, source, first_seen_at, last_seen_at)
714+ VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12, ?13, ?14, ?15, ?16, ?17, ?18, ?19, ?20, ?21, ?22,
715+ 'new', ?23, 'discovered', ?17, ?17)",
716+ )
717+ .bind(&[
718+ model.model.as_str().into(),
719+ model.name.as_str().into(),
720+ provider.as_str().into(),
721+ model.kind.as_str().into(),
722+ (model.input_micros as f64).into(),
723+ (model.output_micros as f64).into(),
724+ (model.cache_read_micros as f64).into(),
725+ (model.cache_write_micros as f64).into(),
726+ (model.cache_write_1h_micros as f64).into(),
727+ number(model.threshold),
728+ (model.over_input_micros as f64).into(),
729+ (model.over_output_micros as f64).into(),
730+ (model.over_cache_read_micros as f64).into(),
731+ (model.over_cache_write_micros as f64).into(),
732+ (model.over_cache_write_1h_micros as f64).into(),
733+ (facts.position as f64).into(),
734+ checked_at.as_str().into(),
735+ facts.family.as_str().into(),
736+ facts.tier_hint.as_str().into(),
737+ number(listed.context_window),
738+ number(listed.max_output),
739+ facts.capabilities.join(",").into(),
740+ f64::from(u8::from(facts.priced)).into(),
741+ ])?,
742+ );
743+ result.added.push(model.model);
744+ }
745+ for (model, alias) in &found.seen {
746+ let listed = a.models.iter().find(|listed| &listed.id == model || alias.as_ref() == Some(&listed.id));
747+ let (context, output) = listed.map_or((0, 0), |l| (l.context_window, l.max_output));
748+ let capabilities = listed.map(|l| l.capabilities.join(",")).unwrap_or_default();
749+ // What the provider says now fills in what g1t did not know;
750+ // a restored model goes back to where it stood.
751+ writes.push(
752+ self.db
753+ .prepare(
754+ "UPDATE gateway_models SET last_seen_at = ?2,
755+ context_window = CASE WHEN ?3 > 0 THEN ?3 ELSE context_window END,
756+ max_output = CASE WHEN ?4 > 0 THEN ?4 ELSE max_output END,
757+ capabilities = CASE WHEN ?5 <> '' AND kind = 'chat' THEN ?5 ELSE capabilities END,
758+ aliases = CASE WHEN ?6 = '' THEN aliases WHEN aliases = '' THEN ?6 ELSE aliases || ',' || ?6 END,
759+ status = CASE WHEN status = 'deprecated' AND missing_since IS NOT NULL
760+ THEN CASE WHEN approved_at IS NOT NULL THEN 'available' ELSE 'new' END
761+ ELSE status END,
762+ missing_since = NULL
763+ WHERE model = ?1",
764+ )
765+ .bind(&[
766+ model.as_str().into(),
767+ checked_at.as_str().into(),
768+ number(context),
769+ number(output),
770+ capabilities.into(),
771+ alias.as_deref().unwrap_or_default().into(),
772+ ])?,
773+ );
774+ }
775+ for model in &found.gone {
776+ writes.push(
777+ self.db
778+ .prepare(
779+ "UPDATE gateway_models SET status = 'deprecated', missing_since = ?2
780+ WHERE model = ?1 AND status IN ('available', 'new')",
781+ )
782+ .bind(&[model.as_str().into(), checked_at.as_str().into()])?,
783+ );
784+ }
785+ result.deprecated = found.gone.clone();
786+ result.restored = found.restored.clone();
787+ if !writes.is_empty() {
788+ self.db.batch(writes).await?;
789+ }
790+ }
791+ self.db
792+ .prepare("INSERT INTO model_checks (id, provider, checked_at, by, listed, added, deprecated, error) VALUES (?, ?, ?, ?, ?, ?, ?, ?)")
793+ .bind(&[
794+ new_id("mck", now).into(),
795+ provider.as_str().into(),
796+ checked_at.as_str().into(),
797+ by.as_str().into(),
798+ f64::from(result.listed).into(),
799+ result.added.join(",").into(),
800+ result.deprecated.join(",").into(),
801+ crate::optional(result.error.as_deref()),
802+ ])?
803+ .run()
804+ .await?;
805+ if !result.added.is_empty() || !result.deprecated.is_empty() || !result.restored.is_empty() {
806+ let detail = discovery_detail(&result);
807+ self.audit("models", "models_discovered", &detail, &by).await?;
808+ self.tell_staff(&result).await;
809+ }
810+ Ok(result)
811+ }
812+
813+ /// Emails staff what a check found; a failure is logged, never raised.
814+ async fn tell_staff(&self, result: &DiscoveryResult) {
815+ if self.caps.alert_to.is_empty() {
816+ return;
817+ }
818+ let (subject, lines) = discovery_email(result);
819+ let sent = crate::margin::email_staff_page(
820+ &self.env,
821+ &self.caps.alert_to,
822+ &subject,
823+ &lines,
824+ ("Agents & models", "https://sudo.g1t.sh/agents"),
825+ "g1t-billing's model catalogue",
826+ )
827+ .await;
828+ if let Err(error) = sent {
829+ worker::console_error!("emailing staff about models failed: {error}");
830+ }
831+ }
832+
833+ /// `admin_decide_model`.
834+ pub(crate) async fn admin_decide_model(&self, a: AdminDecideModelArgs) -> Result<Outcome<CatalogueModel>> {
835+ let reason = a.reason.trim();
836+ if reason.is_empty() {
837+ return Ok(Outcome::fail(FailureCode::Invalid, "Say why, for whoever looks next."));
838+ }
839+ let catalogue = self.catalogue().await?;
840+ let Some(model) = catalogue.iter().find(|model| model.prices.model == a.model.trim()) else {
841+ return Ok(Outcome::fail(FailureCode::NotFound, format!("{} is not in the catalogue.", a.model.trim())));
842+ };
843+ let now = rfc3339(now_ms());
844+ let id = model.prices.model.as_str();
845+ let detail;
846+ match a.decision.as_str() {
847+ "approve" => {
848+ let prices = a.prices.clone().unwrap_or(ModelPrices {
849+ input_micros: model.prices.input_micros,
850+ output_micros: model.prices.output_micros,
851+ cache_read_micros: model.prices.cache_read_micros,
852+ cache_write_micros: model.prices.cache_write_micros,
853+ cache_write_1h_micros: model.prices.cache_write_1h_micros,
854+ threshold: model.prices.threshold,
855+ over_input_micros: model.prices.over_input_micros,
856+ over_output_micros: model.prices.over_output_micros,
857+ over_cache_read_micros: model.prices.over_cache_read_micros,
858+ over_cache_write_micros: model.prices.over_cache_write_micros,
859+ over_cache_write_1h_micros: model.prices.over_cache_write_1h_micros,
860+ });
861+ if let Err(why) = check_prices(&model.prices.kind, &prices) {
862+ return Ok(Outcome::fail(FailureCode::Invalid, why));
863+ }
864+ let name: String = a.name.as_deref().map(str::trim).filter(|n| !n.is_empty()).unwrap_or(&model.prices.name).chars().take(120).collect();
865+ let tier = a.tier_hint.as_deref().map(str::trim).unwrap_or(&model.tier_hint).to_owned();
866+ if !tier.is_empty() && !TIERS.contains(&tier.as_str()) {
867+ return Ok(Outcome::fail(FailureCode::Invalid, "A tier hint is small, large, frontier, or none."));
868+ }
869+ self.db
870+ .prepare(
871+ "UPDATE gateway_models SET status = 'available', priced = 1, name = ?2, tier_hint = ?3,
872+ input_micros = ?4, output_micros = ?5, cache_read_micros = ?6, cache_write_micros = ?7,
873+ cache_write_1h_micros = ?8, threshold = ?9, over_input_micros = ?10, over_output_micros = ?11,
874+ over_cache_read_micros = ?12, over_cache_write_micros = ?13, over_cache_write_1h_micros = ?14,
875+ approved_by = ?15, approved_at = ?16, updated_at = ?16, missing_since = NULL, note = ?17
876+ WHERE model = ?1",
877+ )
878+ .bind(&[
879+ id.into(),
880+ name.as_str().into(),
881+ tier.as_str().into(),
882+ (prices.input_micros as f64).into(),
883+ (prices.output_micros as f64).into(),
884+ (prices.cache_read_micros as f64).into(),
885+ (prices.cache_write_micros as f64).into(),
886+ (prices.cache_write_1h_micros as f64).into(),
887+ number(prices.threshold),
888+ (prices.over_input_micros as f64).into(),
889+ (prices.over_output_micros as f64).into(),
890+ (prices.over_cache_read_micros as f64).into(),
891+ (prices.over_cache_write_micros as f64).into(),
892+ (prices.over_cache_write_1h_micros as f64).into(),
893+ a.by.as_str().into(),
894+ now.as_str().into(),
895+ reason.into(),
896+ ])?
897+ .run()
898+ .await?;
899+ detail = format!(
900+ "{id} approved as {name}: ${} in, ${} out per million. {reason}",
901+ dollars(prices.input_micros),
902+ dollars(prices.output_micros)
903+ );
904+ self.audit("models", "model_approved", &detail, &a.by).await?;
905+ }
906+ "retire" => {
907+ // A default that names it falls back on its own (`resolve`);
908+ // say which, so staff can choose another.
909+ let using: Vec<String> = self
910+ .model_default_rows()
911+ .await?
912+ .into_iter()
913+ .filter(|row| row.model.as_deref() == Some(id))
914+ .map(|row| purpose_label(&row.purpose))
915+ .collect();
916+ self.db
917+ .prepare("UPDATE gateway_models SET status = 'retired', updated_at = ?2, note = ?3 WHERE model = ?1")
918+ .bind(&[id.into(), now.as_str().into(), reason.into()])?
919+ .run()
920+ .await?;
921+ let defaults = if using.is_empty() { String::new() } else { format!(" Defaults that fall back now: {}.", using.join(", ")) };
922+ detail = format!("{id} retired. {reason}{defaults}");
923+ self.audit("models", "model_retired", &detail, &a.by).await?;
924+ }
925+ "restore" => {
926+ // Back to available if it was ever approved; else waiting again.
927+ self.db
928+ .prepare(
929+ "UPDATE gateway_models SET status = CASE WHEN approved_at IS NOT NULL THEN 'available' ELSE 'new' END,
930+ missing_since = NULL, updated_at = ?2, note = ?3 WHERE model = ?1",
931+ )
932+ .bind(&[id.into(), now.as_str().into(), reason.into()])?
933+ .run()
934+ .await?;
935+ detail = format!("{id} restored. {reason}");
936+ self.audit("models", "model_restored", &detail, &a.by).await?;
937+ }
938+ _ => return Ok(Outcome::fail(FailureCode::Invalid, "Approve, retire or restore.")),
939+ }
940+ let updated = self.catalogue().await?.into_iter().find(|model| model.prices.model == id);
941+ Ok(match updated {
942+ Some(model) => Outcome::Ok(model),
943+ None => Outcome::fail(FailureCode::NotFound, format!("{id} is not in the catalogue.")),
944+ })
945+ }
946+
947+ /// `admin_set_model_default`.
948+ pub(crate) async fn admin_set_model_default(&self, a: AdminSetModelDefaultArgs) -> Result<Outcome<ModelDefault>> {
949+ let catalogue = self.catalogue().await?;
950+ let (model, tier, effort) = match check_default(&a, &catalogue) {
951+ Ok(value) => value,
952+ Err(why) => return Ok(Outcome::fail(FailureCode::Invalid, why)),
953+ };
954+ let before = self.model_default(&a.purpose).await?;
955+ let reason: String = a.reason.trim().chars().take(500).collect();
956+ let now = rfc3339(now_ms());
957+ self.db
958+ .prepare(
959+ "INSERT INTO model_defaults (purpose, model, tier, effort, updated_at, updated_by, reason)
960+ VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7)
961+ ON CONFLICT (purpose) DO UPDATE SET model = ?2, tier = ?3, effort = ?4, updated_at = ?5, updated_by = ?6, reason = ?7",
962+ )
963+ .bind(&[
964+ a.purpose.as_str().into(),
965+ crate::optional(model.as_deref()),
966+ crate::optional(tier.as_deref()),
967+ crate::optional(effort.as_deref()),
968+ now.as_str().into(),
969+ a.by.as_str().into(),
970+ reason.as_str().into(),
971+ ])?
972+ .run()
973+ .await?;
974+ let old = before
975+ .as_ref()
976+ .map_or("nothing".to_owned(), |row| describe_default(row.model.as_deref(), row.tier.as_deref(), row.effort.as_deref()));
977+ let new = describe_default(model.as_deref(), tier.as_deref(), effort.as_deref());
978+ self.audit("models", "model_default", &format!("{}: {old} → {new}. {reason}", purpose_label(&a.purpose)), &a.by).await?;
979+ Ok(Outcome::Ok(ModelDefault { purpose: a.purpose, model, tier, effort, updated_at: now, updated_by: a.by, reason }))
980+ }
981+
982+ /// Daily: checks older than `CHECK_DAYS` go.
983+ pub(crate) async fn forget_model_checks(&self) -> Result<()> {
984+ let cutoff = rfc3339(now_ms().saturating_sub(CHECK_DAYS * DAY_MS));
985+ self.db.prepare("DELETE FROM model_checks WHERE checked_at < ?").bind(&[cutoff.into()])?.run().await?;
986+ Ok(())
987+ }
988+}
989+
990+/// `$0.10`, `$2`, `$12.50`: dollars per million from millionths.
991+pub(crate) fn dollars(micros: i64) -> String {
992+ let text = format!("{:.4}", micros as f64 / 1_000_000.0);
993+ let text = text.trim_end_matches('0').trim_end_matches('.');
994+ match text.split_once('.') {
995+ Some((whole, cents)) if cents.len() == 1 => format!("{whole}.{cents}0"),
996+ _ => text.to_owned(),
997+ }
998+}
999+
1000+/// What a check found, for the audit log.
1001+pub(crate) fn discovery_detail(result: &DiscoveryResult) -> String {
1002+ let mut parts = vec![];
1003+ if !result.added.is_empty() {
1004+ parts.push(format!("new: {}", result.added.join(", ")));
1005+ }
1006+ if !result.deprecated.is_empty() {
1007+ parts.push(format!("no longer listed: {}", result.deprecated.join(", ")));
1008+ }
1009+ if !result.restored.is_empty() {
1010+ parts.push(format!("listed again: {}", result.restored.join(", ")));
1011+ }
1012+ format!("{}: {}", result.provider, parts.join("; "))
1013+}
1014+
1015+/// The email to staff about a check: its subject and paragraphs.
1016+pub(crate) fn discovery_email(result: &DiscoveryResult) -> (String, Vec<String>) {
1017+ let provider = match result.provider.as_str() {
1018+ "anthropic" => "Anthropic",
1019+ "workers-ai" => "Workers AI",
1020+ other => other,
1021+ };
1022+ let subject = if result.added.is_empty() {
1023+ format!("g1t: {provider} no longer lists {}", result.deprecated.join(", "))
1024+ } else {
1025+ let more = if result.added.len() > 3 { format!(" and {} more", result.added.len() - 3) } else { String::new() };
1026+ format!("g1t: new {provider} models: {}{more}", result.added.iter().take(3).cloned().collect::<Vec<_>>().join(", "))
1027+ };
1028+ let mut lines = vec![];
1029+ if !result.added.is_empty() {
1030+ lines.push(format!(
1031+ "{provider} lists {} model{} g1t has not used before: {}. Each is in the catalogue as new: nothing routes to it, offers it or charges for it until you approve it with its prices in sudo.",
1032+ result.added.len(),
1033+ if result.added.len() == 1 { "" } else { "s" },
1034+ result.added.join(", ")
1035+ ));
1036+ }
1037+ if !result.deprecated.is_empty() {
1038+ lines.push(format!(
1039+ "{provider} no longer lists {}. Each is deprecated now: no default routes to it, and any default that chose it uses the next suitable model. Choose another default, or retire it.",
1040+ result.deprecated.join(", ")
1041+ ));
1042+ }
1043+ if !result.restored.is_empty() {
1044+ lines.push(format!("{provider} lists {} again; each is back where it stood.", result.restored.join(", ")));
1045+ }
1046+ (subject, lines)
1047+}
1048+
1049+#[cfg(test)]
1050+#[path = "catalogue_tests.rs"]
1051+mod tests;
+333−0
1+//! The model catalogue: discovery's diff, prices, typical runs, defaults
2+//! and their fallbacks (`catalogue.rs`).
3+
4+use super::*;
5+
6+fn model(id: &str, name: &str, tier: &str, status: &str) -> CatalogueModel {
7+ let prices = known_prices(id).unwrap_or(ModelPrices { input_micros: 1_000_000, output_micros: 5_000_000, ..ModelPrices::default() });
8+ let base = GatewayModel {
9+ model: id.into(),
10+ name: name.into(),
11+ provider: "anthropic".into(),
12+ kind: "chat".into(),
13+ input_micros: 0,
14+ output_micros: 0,
15+ cache_read_micros: 0,
16+ cache_write_micros: 0,
17+ cache_write_1h_micros: 0,
18+ threshold: 0,
19+ over_input_micros: 0,
20+ over_output_micros: 0,
21+ over_cache_read_micros: 0,
22+ over_cache_write_micros: 0,
23+ over_cache_write_1h_micros: 0,
24+ };
25+ CatalogueModel {
26+ prices: with_prices(base, &prices),
27+ aliases: vec![],
28+ family: String::new(),
29+ tier_hint: tier.into(),
30+ context_window: 0,
31+ max_output: 0,
32+ capabilities: vec!["effort".into()],
33+ dimensions: 0,
34+ status: status.into(),
35+ priced: true,
36+ source: "staff".into(),
37+ first_seen_at: None,
38+ last_seen_at: None,
39+ missing_since: None,
40+ approved_by: None,
41+ approved_at: None,
42+ note: String::new(),
43+ typical_run_micros: 0,
44+ }
45+}
46+
47+fn listed(id: &str) -> ProviderModel {
48+ ProviderModel { id: id.into(), name: String::new(), kind: "chat".into(), ..ProviderModel::default() }
49+}
50+
51+fn catalogue() -> Vec<CatalogueModel> {
52+ let mut haiku_4_5 = model("claude-haiku-4-5", "Claude Haiku 4.5", "small", "available");
53+ haiku_4_5.aliases = vec!["claude-haiku-4-5-20251001".into()];
54+ vec![
55+ model("claude-haiku-5-5", "Claude Haiku 5.5", "small", "available"),
56+ haiku_4_5,
57+ model("claude-haiku-4-5-20251001", "Claude Haiku 4.5", "small", "available"),
58+ model("claude-sonnet-5-5", "Claude Sonnet 5.5", "large", "available"),
59+ model("claude-opus-5-5", "Claude Opus 5.5", "frontier", "available"),
60+ ]
61+}
62+
63+#[test]
64+fn a_model_never_seen_is_added_and_one_no_longer_listed_is_gone() {
65+ let found = diff(
66+ "anthropic",
67+ &catalogue(),
68+ &[listed("claude-haiku-6"), listed("claude-haiku-5-5"), listed("claude-haiku-4-5-20251001"), listed("claude-sonnet-5-5")],
69+ );
70+ assert_eq!(found.added.iter().map(|m| m.id.as_str()).collect::<Vec<_>>(), ["claude-haiku-6"]);
71+ // Opus 5.5 was not listed: gone. Haiku 4.5 is listed by its dated id, which it has as an alias.
72+ assert_eq!(found.gone, ["claude-opus-5-5"]);
73+ let seen: Vec<&str> = found.seen.iter().map(|(m, _)| m.as_str()).collect();
74+ assert_eq!(seen, ["claude-haiku-5-5", "claude-haiku-4-5", "claude-haiku-4-5-20251001", "claude-sonnet-5-5"]);
75+ assert!(found.restored.is_empty());
76+}
77+
78+#[test]
79+fn a_dated_id_of_a_model_it_has_is_that_model_and_becomes_an_alias() {
80+ let found = diff("anthropic", &catalogue(), &[listed("claude-sonnet-5-5-20261001")]);
81+ assert!(found.added.is_empty());
82+ assert_eq!(found.seen, [("claude-sonnet-5-5".to_owned(), Some("claude-sonnet-5-5-20261001".to_owned()))]);
83+ assert!(is_dated("claude-sonnet-5-5-20261001", "claude-sonnet-5-5"));
84+ assert!(!is_dated("claude-sonnet-5-5-preview", "claude-sonnet-5-5"));
85+ assert!(!is_dated("claude-sonnet-5", "claude-sonnet-5-5"));
86+}
87+
88+#[test]
89+fn an_empty_or_failed_list_finds_nothing_gone() {
90+ let found = diff("anthropic", &catalogue(), &[]);
91+ assert_eq!(found, Diff::default());
92+ // Another provider's models are never its to judge.
93+ let found = diff("workers-ai", &catalogue(), &[listed("@cf/meta/llama-5")]);
94+ assert!(found.gone.is_empty());
95+ assert_eq!(found.added.len(), 1);
96+}
97+
98+#[test]
99+fn a_deprecated_model_listed_again_is_restored_and_a_retired_one_is_left_alone() {
100+ let mut models = catalogue();
101+ models[3].status = "deprecated".into();
102+ models[3].missing_since = Some("2026-10-01T04:17:00Z".into());
103+ models[4].status = "retired".into();
104+ let found = diff("anthropic", &models, &[listed("claude-sonnet-5-5"), listed("claude-haiku-5-5")]);
105+ assert_eq!(found.restored, ["claude-sonnet-5-5"]);
106+ // Retired Opus is not listed and stays retired; Haiku 4.5 is gone.
107+ assert_eq!(found.gone, ["claude-haiku-4-5", "claude-haiku-4-5-20251001"]);
108+}
109+
110+#[test]
111+fn only_chat_and_embeddings_models_are_added() {
112+ let mut speech = listed("@cf/openai/whisper");
113+ speech.kind = "other".into();
114+ let mut embed = listed("@cf/baai/bge-large-en-v1.5");
115+ embed.kind = "embeddings".into();
116+ let found = diff("workers-ai", &[], &[speech, embed, listed("@cf/meta/llama-5")]);
117+ assert_eq!(found.added.iter().map(|m| m.id.as_str()).collect::<Vec<_>>(), ["@cf/baai/bge-large-en-v1.5", "@cf/meta/llama-5"]);
118+}
119+
120+#[test]
121+fn a_new_model_is_priced_from_the_table_or_its_listing_or_not_at_all() {
122+ // A dated Opus 5.5 is in the table.
123+ let (row, facts) = new_row("anthropic", &listed("claude-opus-5-5-20261101"), 120);
124+ assert!(facts.priced);
125+ assert_eq!((row.input_micros, row.output_micros, row.cache_read_micros), (4_000_000, 20_000_000, 200_000));
126+ assert_eq!((facts.family.as_str(), facts.tier_hint.as_str(), facts.position), ("opus", "frontier", 120));
127+ // A Haiku g1t has no price for: unpriced, a fast model by its name.
128+ let (row, facts) = new_row("anthropic", &listed("claude-haiku-6"), 0);
129+ assert!(!facts.priced);
130+ assert_eq!((row.input_micros, facts.tier_hint.as_str()), (0, "small"));
131+ // Workers AI says its own price; cached tokens cost what input does.
132+ let mut open = listed("@cf/meta/llama-5-70b");
133+ open.price = Some(g1t_contracts::billing::ListedPrice { input_micros: 290_000, output_micros: 2_250_000 });
134+ let (row, facts) = new_row("workers-ai", &open, 0);
135+ assert!(facts.priced);
136+ assert_eq!((row.input_micros, row.output_micros, row.cache_read_micros, row.name.as_str()), (290_000, 2_250_000, 290_000, "llama-5-70b"));
137+ assert_eq!((facts.family.as_str(), facts.tier_hint.as_str()), ("meta", ""));
138+}
139+
140+#[test]
141+fn haiku_5_5_is_known_with_its_long_prompt_prices() {
142+ let prices = known_prices("claude-haiku-5-5").unwrap();
143+ assert_eq!((prices.input_micros, prices.output_micros, prices.cache_read_micros), (100_000, 500_000, 10_000));
144+ assert_eq!((prices.threshold, prices.over_input_micros, prices.over_output_micros), (100_000, 500_000, 2_500_000));
145+ assert!(known_prices("claude-haiku-6").is_none());
146+ // `claude-opus-5` is not `claude-opus-5-5`.
147+ assert_eq!(known_prices("claude-opus-5").unwrap().input_micros, 5_000_000);
148+}
149+
150+#[test]
151+fn a_typical_run_is_priced_request_by_request() {
152+ let models = catalogue();
153+ let haiku = &models[0].prices;
154+ let sonnet = &models[3].prices;
155+ let opus = &models[4].prices;
156+ // 40 requests of 2k in, 1.5k out, 45k cache reads and 4k cache writes.
157+ assert_eq!(typical_run(sonnet), 1_340_000);
158+ assert_eq!(typical_run(opus), 2_680_000);
159+ // Each prompt is 51k, under Haiku 5.5's 100k threshold: its lower prices.
160+ assert_eq!(typical_run(haiku), 76_000);
161+ let mut embeddings = haiku.clone();
162+ embeddings.kind = "embeddings".into();
163+ assert_eq!(typical_run(&embeddings), 0);
164+}
165+
166+#[test]
167+fn a_default_that_can_be_used_is_used() {
168+ let resolved = resolve("tier_small", "claude-haiku-5-5", &catalogue());
169+ assert_eq!(resolved.model.unwrap().model, "claude-haiku-5-5");
170+ assert!(resolved.note.is_none());
171+}
172+
173+#[test]
174+fn a_retired_default_falls_back_to_the_next_model_of_its_tier_and_says_so() {
175+ let mut models = catalogue();
176+ models[0].status = "retired".into();
177+ let resolved = resolve("tier_small", "claude-haiku-5-5", &models);
178+ assert_eq!(resolved.model.unwrap().model, "claude-haiku-4-5");
179+ assert_eq!(resolved.note.as_deref(), Some("Claude Haiku 5.5 is retired; using Claude Haiku 4.5 instead."));
180+ // Deprecated and unpriced models are never handed out either.
181+ models[1].status = "deprecated".into();
182+ models[2].priced = false;
183+ let resolved = resolve("background", "claude-haiku-5-5", &models);
184+ assert!(resolved.model.is_none());
185+ assert_eq!(resolved.note.as_deref(), Some("Claude Haiku 5.5 is retired, and no other model suits it."));
186+ // The gateway's first model takes any Claude.
187+ let resolved = resolve("gateway_first", "claude-haiku-5-5", &models);
188+ assert_eq!(resolved.model.unwrap().model, "claude-sonnet-5-5");
189+}
190+
191+#[test]
192+fn a_new_model_is_never_a_default_until_approved() {
193+ let mut models = catalogue();
194+ models.push(model("claude-haiku-6", "Claude Haiku 6", "small", "new"));
195+ let resolved = resolve("tier_small", "claude-haiku-6", &models);
196+ assert_eq!(resolved.model.unwrap().model, "claude-haiku-5-5");
197+ assert_eq!(resolved.note.as_deref(), Some("Claude Haiku 6 is new; using Claude Haiku 5.5 instead."));
198+ let set = AdminSetModelDefaultArgs { purpose: "tier_small".into(), model: Some("claude-haiku-6".into()), reason: "cheaper".into(), ..Default::default() };
199+ assert_eq!(check_default(&set, &models).unwrap_err(), "Claude Haiku 6 is new: approve it first.");
200+}
201+
202+#[test]
203+fn defaults_are_checked_before_they_are_kept() {
204+ let models = catalogue();
205+ let mut open = model("@cf/openai/gpt-oss-120b", "gpt-oss-120b", "", "available");
206+ open.prices.provider = "workers-ai".into();
207+ let models = [models, vec![open]].concat();
208+ let args = |purpose: &str, model: Option<&str>, tier: Option<&str>, effort: Option<&str>, reason: &str| AdminSetModelDefaultArgs {
209+ purpose: purpose.into(),
210+ model: model.map(Into::into),
211+ tier: tier.map(Into::into),
212+ effort: effort.map(Into::into),
213+ reason: reason.into(),
214+ by: "staff@g1t.sh".into(),
215+ };
216+ assert_eq!(check_default(&args("tier_large", Some("claude-opus-5-5"), None, None, "try it"), &models), Ok((Some("claude-opus-5-5".into()), None, None)));
217+ assert_eq!(check_default(&args("tier_large", Some("claude-opus-5-5"), None, None, " "), &models).unwrap_err(), "Say why, for whoever looks next.");
218+ assert!(check_default(&args("tier_large", Some("@cf/openai/gpt-oss-120b"), None, None, "x"), &models).unwrap_err().contains("not a Claude chat model"));
219+ assert_eq!(check_default(&args("tier_large", Some("claude-nope"), None, None, "x"), &models).unwrap_err(), "claude-nope is not in the catalogue.");
220+ assert_eq!(
221+ check_default(&args("job_plan", None, Some("small"), Some("xhigh"), "x"), &models),
222+ Ok((None, Some("small".into()), Some("xhigh".into())))
223+ );
224+ assert_eq!(check_default(&args("job_update", None, Some("small"), Some(""), "x"), &models), Ok((None, Some("small".into()), None)));
225+ assert!(check_default(&args("job_plan", None, Some("change"), None, "x"), &models).is_err());
226+ assert!(check_default(&args("job_review", None, Some("change"), None, "x"), &models).is_ok());
227+ assert!(check_default(&args("job_plan", None, Some("small"), Some("huge"), "x"), &models).is_err());
228+ assert!(check_default(&args("job_dance", None, Some("small"), None, "x"), &models).is_err());
229+}
230+
231+#[test]
232+fn the_runtime_view_has_each_purpose_and_job() {
233+ let row = |purpose: &str, model: Option<&str>, tier: Option<&str>, effort: Option<&str>| ModelDefault {
234+ purpose: purpose.into(),
235+ model: model.map(Into::into),
236+ tier: tier.map(Into::into),
237+ effort: effort.map(Into::into),
238+ updated_at: "2026-10-08T00:00:00Z".into(),
239+ updated_by: "migration".into(),
240+ reason: String::new(),
241+ };
242+ let defaults = vec![
243+ row("tier_small", Some("claude-haiku-5-5"), None, None),
244+ row("gateway_first", Some("claude-haiku-5-5"), None, None),
245+ row("job_plan", None, Some("small"), Some("high")),
246+ row("job_review", None, Some("change"), None),
247+ row("job_answer", None, Some("nonsense"), None),
248+ ];
249+ let view = defaults_view(&defaults, &catalogue());
250+ assert_eq!(view.models.iter().map(|m| m.purpose.as_str()).collect::<Vec<_>>(), ["tier_small", "gateway_first"]);
251+ assert_eq!(
252+ view.jobs,
253+ [
254+ JobDefault { kind: "review".into(), tier: "change".into(), effort: None },
255+ JobDefault { kind: "plan".into(), tier: "small".into(), effort: Some("high".into()) },
256+ ]
257+ );
258+}
259+
260+#[test]
261+fn the_gateway_lists_staffs_first_claude_first() {
262+ let mut models: Vec<GatewayModel> = catalogue().into_iter().map(|m| m.prices).collect();
263+ put_first(&mut models, Some("claude-sonnet-5-5"));
264+ assert_eq!(models[0].model, "claude-sonnet-5-5");
265+ assert_eq!(models[1].model, "claude-haiku-5-5");
266+ put_first(&mut models, Some("claude-nope"));
267+ assert_eq!(models[0].model, "claude-sonnet-5-5");
268+}
269+
270+#[test]
271+fn prices_staff_confirm_must_make_sense() {
272+ let good = known_prices("claude-haiku-5-5").unwrap();
273+ assert!(check_prices("chat", &good).is_ok());
274+ assert!(check_prices("chat", &ModelPrices { output_micros: 0, ..good.clone() }).is_err());
275+ assert!(check_prices("chat", &ModelPrices { input_micros: -1, ..good.clone() }).is_err());
276+ assert!(check_prices("chat", &ModelPrices { threshold: 0, ..good.clone() }).is_err());
277+ assert!(check_prices("chat", &ModelPrices { input_micros: 2_000_000_000, ..good.clone() }).is_err());
278+ assert!(check_prices("embeddings", &ModelPrices { input_micros: 12_000, ..ModelPrices::default() }).is_ok());
279+}
280+
281+#[test]
282+fn staff_are_told_what_a_check_found() {
283+ let result = DiscoveryResult {
284+ provider: "anthropic".into(),
285+ added: vec!["claude-haiku-6".into()],
286+ deprecated: vec!["claude-haiku-4-5".into()],
287+ ..DiscoveryResult::default()
288+ };
289+ let (subject, lines) = discovery_email(&result);
290+ assert_eq!(subject, "g1t: new Anthropic models: claude-haiku-6");
291+ assert!(lines[0].contains("until you approve it with its prices in sudo"));
292+ assert!(lines[1].contains("no longer lists claude-haiku-4-5"));
293+ assert_eq!(discovery_detail(&result), "anthropic: new: claude-haiku-6; no longer listed: claude-haiku-4-5");
294+}
295+
296+#[test]
297+fn dollars_read_plainly() {
298+ assert_eq!(dollars(100_000), "0.10");
299+ assert_eq!(dollars(2_000_000), "2");
300+ assert_eq!(dollars(12_500_000), "12.50");
301+ assert_eq!(dollars(10_000), "0.01");
302+ assert_eq!(dollars(125_000), "0.125");
303+}
304+
305+#[test]
306+fn the_migration_seeds_every_purpose_as_routing_had_it() {
307+ let sql = include_str!("../migrations/0048_model_catalogue.sql");
308+ for purpose in MODEL_PURPOSES {
309+ assert!(sql.contains(&format!("('{purpose}', 'claude-")), "no seed for {purpose}");
310+ }
311+ for kind in JOB_KINDS {
312+ assert!(sql.contains(&format!("('job_{kind}', NULL, '")), "no seed for job_{kind}");
313+ }
314+ // As AGENT_ROUTING (services/runner/wrangler.jsonc) had them.
315+ let runner = include_str!("../../runner/wrangler.jsonc");
316+ for (purpose, model) in [("tier_small", "claude-haiku-5-5"), ("tier_large", "claude-sonnet-5-5"), ("tier_frontier", "claude-opus-5-5")] {
317+ assert!(sql.contains(&format!("('{purpose}', '{model}'")));
318+ assert!(runner.contains(&format!("\\\"model\\\":\\\"{model}\\\"")), "{model} is not in AGENT_ROUTING");
319+ }
320+ assert!(sql.contains("('job_plan', NULL, 'small', 'high'"));
321+ assert!(sql.contains("('job_review', NULL, 'change', NULL"));
322+ // Every column the discovery insert names exists.
323+ let insert = "model, name, provider, kind, input_micros, output_micros, cache_read_micros, cache_write_micros, cache_write_1h_micros, threshold, over_input_micros, over_output_micros, over_cache_read_micros, over_cache_write_micros, over_cache_write_1h_micros, position, updated_at, family, tier_hint, context_window, max_output, capabilities, status, priced, source, first_seen_at, last_seen_at";
324+ let all = [
325+ include_str!("../migrations/0045_gateway.sql"),
326+ include_str!("../migrations/0047_gateway_formats.sql"),
327+ sql,
328+ ]
329+ .concat();
330+ for column in insert.split(", ") {
331+ assert!(all.contains(&format!("{column} ")) || all.contains(&format!("{column},")), "gateway_models has no {column}");
332+ }
333+}
+67−22
178178 id.starts_with("gw_") && id.len() <= 64 && id.chars().all(|c| c.is_ascii_alphanumeric() || c == '_' || c == '-')
179179 }
180180
181−#[derive(Deserialize)]
182−struct ModelRow {
183− model: String,
184− name: String,
185− provider: String,
181+/// A row of `gateway_models`, the model catalogue (catalogue.rs reads the
182+/// rest of its columns).
183+#[derive(Clone, Deserialize)]
184+pub(crate) struct ModelRow {
185+ pub model: String,
186+ pub name: String,
187+ pub provider: String,
186188 #[serde(default)]
187− kind: Option<String>,
188− input_micros: i64,
189− output_micros: i64,
190− cache_read_micros: i64,
191− cache_write_micros: i64,
189+ pub kind: Option<String>,
190+ pub input_micros: i64,
191+ pub output_micros: i64,
192+ pub cache_read_micros: i64,
193+ pub cache_write_micros: i64,
192194 #[serde(default)]
193− cache_write_1h_micros: Option<i64>,
195+ pub cache_write_1h_micros: Option<i64>,
196+ #[serde(default)]
197+ pub threshold: Option<f64>,
198+ #[serde(default)]
199+ pub over_input_micros: Option<i64>,
200+ #[serde(default)]
201+ pub over_output_micros: Option<i64>,
202+ #[serde(default)]
203+ pub over_cache_read_micros: Option<i64>,
204+ #[serde(default)]
205+ pub over_cache_write_micros: Option<i64>,
194206 #[serde(default)]
195− threshold: Option<f64>,
207+ pub over_cache_write_1h_micros: Option<i64>,
196208 #[serde(default)]
197− over_input_micros: Option<i64>,
209+ pub aliases: Option<String>,
198210 #[serde(default)]
199− over_output_micros: Option<i64>,
211+ pub family: Option<String>,
200212 #[serde(default)]
201− over_cache_read_micros: Option<i64>,
213+ pub tier_hint: Option<String>,
202214 #[serde(default)]
203− over_cache_write_micros: Option<i64>,
215+ pub context_window: Option<f64>,
216+ #[serde(default)]
217+ pub max_output: Option<f64>,
218+ #[serde(default)]
219+ pub capabilities: Option<String>,
220+ #[serde(default)]
221+ pub dimensions: Option<f64>,
222+ #[serde(default)]
223+ pub status: Option<String>,
224+ #[serde(default)]
225+ pub priced: Option<f64>,
226+ #[serde(default)]
227+ pub source: Option<String>,
228+ #[serde(default)]
229+ pub first_seen_at: Option<String>,
230+ #[serde(default)]
231+ pub last_seen_at: Option<String>,
232+ #[serde(default)]
233+ pub missing_since: Option<String>,
234+ #[serde(default)]
235+ pub approved_by: Option<String>,
236+ #[serde(default)]
237+ pub approved_at: Option<String>,
204238 #[serde(default)]
205− over_cache_write_1h_micros: Option<i64>,
239+ pub note: Option<String>,
206240 }
207241
242+/// The models the AI Gateway offers: approved or deprecated, and priced.
243+/// `new` (not approved yet), `retired` and unpriced models are refused
244+/// before they reach a provider, and never charged.
245+pub(crate) const OFFERED_SQL: &str = "status IN ('available', 'deprecated') AND priced = 1";
246+
208247 impl From<ModelRow> for GatewayModel {
209248 fn from(row: ModelRow) -> Self {
210249 GatewayModel {
289328 }
290329
291330 impl Billing {
292− /// `gateway_models`: what the gateway offers on g1t's key, with prices.
331+ /// `gateway_models`: what the gateway offers on g1t's key, with prices,
332+ /// in the catalogue's order, with staff's first Claude (`gateway_first`
333+ /// in `model_defaults`) at the top.
293334 pub(crate) async fn gateway_models(&self) -> Result<Vec<GatewayModel>> {
294− Ok(self
335+ let mut models: Vec<GatewayModel> = self
295336 .db
296− .prepare("SELECT * FROM gateway_models ORDER BY position, model")
337+ .prepare(format!("SELECT * FROM gateway_models WHERE {OFFERED_SQL} ORDER BY position, model"))
297338 .all()
298339 .await?
299340 .results::<ModelRow>()?
300341 .into_iter()
301342 .map(GatewayModel::from)
302− .collect())
343+ .collect();
344+ // Without the defaults (a database before them), the catalogue's order.
345+ let first = self.model_default("gateway_first").await.ok().flatten().and_then(|row| row.model);
346+ crate::catalogue::put_first(&mut models, first.as_deref());
347+ Ok(models)
303348 }
304349
305350 async fn gateway_model(&self, model: &str) -> Result<Option<GatewayModel>> {
306351 Ok(self
307352 .db
308− .prepare("SELECT * FROM gateway_models WHERE model = ?")
353+ .prepare(format!("SELECT * FROM gateway_models WHERE model = ? AND {OFFERED_SQL}"))
309354 .bind(&[model.into()])?
310355 .first::<ModelRow>(None)
311356 .await?
+11−0
2020 mod accounts;
2121 mod ai;
2222 mod budget;
23+mod catalogue;
2324 mod cards;
2425 mod closing;
2526 mod compute;
12561257 if let Err(error) = billing.forget_gateway_requests().await {
12571258 worker::console_error!("clearing old AI Gateway requests failed: {error}");
12581259 }
1260+ // Checks of the providers' model lists keep 90 days (catalogue.rs).
1261+ if let Err(error) = billing.forget_model_checks().await {
1262+ worker::console_error!("clearing old model checks failed: {error}");
1263+ }
12591264 }
12601265 }
12611266
13071312 "record_tokens" => reply(&billing.record_tokens(args(body)?).await?),
13081313 "token_usage" => reply(&billing.token_usage(args(body)?).await?),
13091314 "gateway_models" => reply(&billing.gateway_models().await?),
1315+ "model_defaults" => reply(&billing.model_defaults().await?),
1316+ "record_discovery" => reply(&billing.record_discovery(args(body)?).await?),
1317+ "admin_models" => reply(&billing.admin_models().await?),
1318+ "admin_decide_model" => reply(&billing.admin_decide_model(args(body)?).await?),
1319+ "admin_set_model_default" => reply(&billing.admin_set_model_default(args(body)?).await?),
13101320 "gateway_admit" => reply(&billing.gateway_admit(args(body)?).await?),
13111321 "record_gateway" => reply(&billing.record_gateway(args(body)?).await?),
13121322 "gateway_requests" => reply(&billing.gateway_requests(args(body)?).await?),
15001510 include_str!("../migrations/0045_gateway.sql"),
15011511 include_str!("../migrations/0046_reset_costs.sql"),
15021512 include_str!("../migrations/0047_gateway_formats.sql"),
1513+ include_str!("../migrations/0048_model_catalogue.sql"),
15031514 ];
15041515
15051516 /// The columns of `table` after the migrations: each with whether an
+11−3
917917
918918 /// Emails staff through Cloudflare Email Sending, the `EMAIL` binding.
919919 pub(crate) async fn email_staff(env: &Env, to: &str, subject: &str, lines: &[String]) -> Result<()> {
920− let link = "https://sudo.g1t.sh/costs";
921− let text = format!("{}\n\nCosts & margin: {link}\n\nSent by g1t-billing's margin guard (COSTS_ALERT_EMAIL).\n", lines.join("\n\n"));
920+ email_staff_page(env, to, subject, lines, ("Costs & margin", "https://sudo.g1t.sh/costs"), "g1t-billing's margin guard").await
921+}
922+
923+/// Emails staff, linking to a page of sudo (`page`: its name and address)
924+/// and saying what sent it.
925+pub(crate) async fn email_staff_page(env: &Env, to: &str, subject: &str, lines: &[String], page: (&str, &str), sender: &str) -> Result<()> {
926+ let (name, link) = page;
927+ let text = format!("{}\n\n{name}: {link}\n\nSent by {sender} (COSTS_ALERT_EMAIL).\n", lines.join("\n\n"));
922928 let mut html = String::from("<div style=\"font-family:system-ui,sans-serif;max-width:560px;margin:0 auto;padding:24px 16px;color:#16150f\">");
923929 for line in lines {
924930 html.push_str(&format!("<p style=\"font-size:15px;line-height:1.6\">{}</p>", escape(line)));
925931 }
926932 html.push_str(&format!(
927− "<p><a href=\"{link}\">Open Costs &amp; margin in sudo</a></p><p style=\"font-size:13px;color:#6e6a5e\">Sent by g1t-billing's margin guard (COSTS_ALERT_EMAIL).</p></div>"
933+ "<p><a href=\"{link}\">Open {} in sudo</a></p><p style=\"font-size:13px;color:#6e6a5e\">Sent by {} (COSTS_ALERT_EMAIL).</p></div>",
934+ escape(name),
935+ escape(sender)
928936 ));
929937 let mail = Mail { to, from: "g1t <noreply@g1t.sh>", subject, text, html };
930938 let binding = g1t_kit::js::binding(env, "EMAIL")?;
+6−1
9393 BILLING?: ServiceBinding;
9494 };
9595
96−/** Workers AI's embedding model: 768 dimensions, as the index was made with. */
96+/**
97+ * Workers AI's embedding model: 768 dimensions, as the index was made with.
98+ * Pinned, not a default staff choose in sudo: vectors from another model
99+ * mean nothing beside these, so changing it means a new index, rebuilt.
100+ * The catalogue (billing's `gateway_models`) lists it with its price.
101+ */
97102 const EMBED_MODEL = "@cf/baai/bge-base-en-v1.5";
98103 /** What it costs, in millionths of a dollar per token ($0.067 per million). */
99104 const MICROS_PER_TOKEN = 0.067;
+123−0
1+import assert from "node:assert/strict";
2+import { test } from "node:test";
3+
4+import type { DiscoveryResult, ProviderModel } from "@g1t/contracts";
5+
6+import { type Fetch, anthropicRequest, discover, fromAnthropic, fromWorkersAi, listAnthropic, listWorkersAi } from "./discover.ts";
7+
8+const hosted = { AI_GATEWAY_ID: "g1t", CLOUDFLARE_ACCOUNT_ID: "acct", AI_GATEWAY_TOKEN: "aig-secret", WORKERS_AI_TOKEN: "wai-secret" };
9+
10+function answering(pages: Record<string, unknown>, seen: { url: string; headers: Headers }[] = []): Fetch {
11+ return async (url, init) => {
12+ seen.push({ url, headers: new Headers(init?.headers) });
13+ const key = Object.keys(pages).find((prefix) => url.includes(prefix));
14+ if (!key) return new Response("not found", { status: 404 });
15+ return Response.json(pages[key]);
16+ };
17+}
18+
19+test("Anthropic's list is read through g1t's gateway, with its token and never a model call", () => {
20+ const request = anthropicRequest(hosted, null)!;
21+ assert.equal(request.url, "https://gateway.ai.cloudflare.com/v1/acct/g1t/anthropic/v1/models?limit=1000");
22+ assert.equal(request.headers.get("cf-aig-authorization"), "Bearer aig-secret");
23+ assert.equal(request.headers.get("anthropic-version"), "2023-06-01");
24+ assert.equal(anthropicRequest(hosted, "claude-x")!.url.endsWith("&after_id=claude-x"), true);
25+ // Without a gateway, straight to Anthropic with g1t's key; with neither, no way.
26+ assert.equal(anthropicRequest({ AI_GATEWAY_ID: "", CLOUDFLARE_ACCOUNT_ID: "acct", ANTHROPIC_API_KEY: "sk" }, null)!.url, "https://api.anthropic.com/v1/models?limit=1000");
27+ assert.equal(anthropicRequest({ AI_GATEWAY_ID: "", CLOUDFLARE_ACCOUNT_ID: "acct" }, null), null);
28+});
29+
30+test("an Anthropic model's context window and capabilities come from its list entry", () => {
31+ const model = fromAnthropic({
32+ id: "claude-haiku-6",
33+ display_name: "Claude Haiku 6",
34+ max_input_tokens: 400000,
35+ max_tokens: 64000,
36+ capabilities: { effort: { supported: true }, thinking: { supported: true }, image_input: { supported: false } },
37+ });
38+ assert.deepEqual(model, {
39+ id: "claude-haiku-6",
40+ name: "Claude Haiku 6",
41+ kind: "chat",
42+ contextWindow: 400000,
43+ maxOutput: 64000,
44+ capabilities: ["effort", "thinking", "tools"],
45+ price: null,
46+ });
47+ assert.equal(fromAnthropic({ display_name: "No id" }), null);
48+ // An older list with no capability tree: nothing claimed.
49+ assert.deepEqual(fromAnthropic({ id: "claude-haiku-4-5-20251001" })!.capabilities, []);
50+});
51+
52+test("Anthropic's list is read page by page", async () => {
53+ const seen: { url: string; headers: Headers }[] = [];
54+ const fetcher: Fetch = async (url, init) => {
55+ seen.push({ url, headers: new Headers(init?.headers) });
56+ return url.includes("after_id=claude-b")
57+ ? Response.json({ data: [{ id: "claude-c" }], has_more: false, last_id: "claude-c" })
58+ : Response.json({ data: [{ id: "claude-a" }, { id: "claude-b" }], has_more: true, last_id: "claude-b" });
59+ };
60+ const models = await listAnthropic(hosted, fetcher);
61+ assert.deepEqual(models.map((m) => m.id), ["claude-a", "claude-b", "claude-c"]);
62+ assert.equal(seen.length, 2);
63+ await assert.rejects(listAnthropic(hosted, async () => new Response("nope", { status: 401 })), /answered 401/);
64+});
65+
66+test("a Workers AI model's kind, context and price come from its search entry", () => {
67+ const chat = fromWorkersAi({
68+ name: "@cf/meta/llama-5-70b",
69+ task: { name: "Text Generation" },
70+ properties: [
71+ { property_id: "context_window", value: "131072" },
72+ { property_id: "function_calling", value: "true" },
73+ { property_id: "price", value: [{ unit: "per M input tokens", price: 0.29 }, { unit: "per M output tokens", price: 2.25 }] },
74+ ],
75+ })!;
76+ assert.deepEqual(chat, {
77+ id: "@cf/meta/llama-5-70b",
78+ name: "llama-5-70b",
79+ kind: "chat",
80+ contextWindow: 131072,
81+ maxOutput: 0,
82+ capabilities: ["tools"],
83+ price: { inputMicros: 290000, outputMicros: 2250000 },
84+ });
85+ const embed = fromWorkersAi({ name: "@cf/baai/bge-large-en-v1.5", task: { name: "Text Embeddings" }, properties: [] })!;
86+ assert.equal(embed.kind, "embeddings");
87+ assert.deepEqual(embed.capabilities, ["embeddings"]);
88+ assert.equal(embed.price, null);
89+ assert.equal(fromWorkersAi({ name: "@cf/openai/whisper", task: { name: "Automatic Speech Recognition" } })!.kind, "automatic-speech-recognition");
90+});
91+
92+test("Workers AI's search is read with the Workers AI token", async () => {
93+ const seen: { url: string; headers: Headers }[] = [];
94+ const models = await listWorkersAi(
95+ hosted,
96+ answering({ "ai/models/search": { result: [{ name: "@cf/openai/gpt-oss-120b", task: { name: "Text Generation" } }] } }, seen),
97+ );
98+ assert.deepEqual(models.map((m) => m.id), ["@cf/openai/gpt-oss-120b"]);
99+ assert.equal(seen[0].url, "https://api.cloudflare.com/client/v4/accounts/acct/ai/models/search?per_page=100&page=1");
100+ assert.equal(seen[0].headers.get("authorization"), "Bearer wai-secret");
101+});
102+
103+test("every provider is recorded, a failed one as a failed check without its key", async () => {
104+ const recorded: { provider: string; models: ProviderModel[]; by: string; error: string | null }[] = [];
105+ const record = async (provider: string, models: ProviderModel[], by: string, error: string | null): Promise<DiscoveryResult> => {
106+ recorded.push({ provider, models, by, error });
107+ return { provider, checkedAt: "", by, listed: models.length, added: [], deprecated: [], restored: [], error };
108+ };
109+ const fetcher: Fetch = async (url) =>
110+ url.includes("anthropic")
111+ ? Response.json({ data: [{ id: "claude-haiku-5-5" }], has_more: false })
112+ : new Response("token wai-secret is not allowed", { status: 403 });
113+ const results = await discover(hosted, fetcher, record, "staff@g1t.sh");
114+ assert.deepEqual(results.map((r) => [r.provider, r.listed]), [["anthropic", 1], ["workers-ai", 0]]);
115+ assert.equal(recorded[0].by, "staff@g1t.sh");
116+ assert.equal(recorded[0].error, null);
117+ assert.match(recorded[1].error!, /answered 403/);
118+ assert.doesNotMatch(recorded[1].error!, /wai-secret/);
119+ // An empty list is a failed check, never every model gone.
120+ await discover(hosted, async () => Response.json({ data: [], result: [] }), record, "");
121+ assert.equal(recorded[2].error, "The list was empty.");
122+ assert.equal(recorded[2].by, "schedule");
123+});
+207−0
1+/**
2+ * Discovery: which models g1t's providers offer now, so new ones reach the
3+ * catalogue without anyone typing an id.
4+ *
5+ * Once a day (the cron in wrangler.jsonc) and when staff press "Check for
6+ * new models" in sudo (the `Discovery` entrypoint in index.ts), this lists:
7+ *
8+ * - **Anthropic**, with `GET /v1/models` through g1t's AI Gateway (the
9+ * same path and credentials g1t's runs use), or straight to Anthropic
10+ * with g1t's key where there is no gateway.
11+ * - **Workers AI**, with Cloudflare's model search
12+ * (`GET /accounts/{account}/ai/models/search`), with `WORKERS_AI_TOKEN`
13+ * (Workers AI Read) or else `AI_GATEWAY_TOKEN`.
14+ *
15+ * Listing models is free: nothing here calls a model. Each provider's list
16+ * goes to billing's `record_discovery`, which owns the catalogue: it adds
17+ * new ids as `new`, marks ones no longer listed `deprecated`, and emails
18+ * staff. A provider that cannot be listed is recorded as a failed check
19+ * and changes nothing.
20+ */
21+import type { DiscoveryResult, ProviderModel } from "@g1t/contracts";
22+
23+import type { HostedRouting } from "./route.ts";
24+
25+export type Fetch = (url: string, init?: RequestInit) => Promise<Response>;
26+
27+/** The version of Anthropic's API the list is read with. */
28+export const ANTHROPIC_VERSION = "2023-06-01";
29+/** The most pages read from either provider: far more models than either offers. */
30+const MAX_PAGES = 20;
31+const WORKERS_AI_PAGE = 100;
32+
33+/** Where Anthropic's list is read, with what; null when g1t has no way to reach Anthropic. */
34+export function anthropicRequest(env: HostedRouting, afterId: string | null): { url: string; headers: Headers } | null {
35+ const headers = new Headers({ "anthropic-version": ANTHROPIC_VERSION });
36+ const query = `?limit=1000${afterId ? `&after_id=${encodeURIComponent(afterId)}` : ""}`;
37+ if (env.ANTHROPIC_API_KEY) headers.set("x-api-key", env.ANTHROPIC_API_KEY);
38+ if (env.AI_GATEWAY_ID && env.CLOUDFLARE_ACCOUNT_ID && (env.AI_GATEWAY_TOKEN || env.ANTHROPIC_API_KEY)) {
39+ if (env.AI_GATEWAY_TOKEN) headers.set("cf-aig-authorization", `Bearer ${env.AI_GATEWAY_TOKEN}`);
40+ // Told apart from runs in the gateway's logs.
41+ headers.set("cf-aig-metadata", JSON.stringify({ task: "discovery" }));
42+ return { url: `https://gateway.ai.cloudflare.com/v1/${env.CLOUDFLARE_ACCOUNT_ID}/${env.AI_GATEWAY_ID}/anthropic/v1/models${query}`, headers };
43+ }
44+ if (!env.ANTHROPIC_API_KEY) return null;
45+ return { url: `https://api.anthropic.com/v1/models${query}`, headers };
46+}
47+
48+type Json = Record<string, unknown>;
49+
50+const isObject = (value: unknown): value is Json => typeof value === "object" && value !== null && !Array.isArray(value);
51+const supported = (node: unknown): boolean => isObject(node) && node.supported === true;
52+const count = (value: unknown): number => {
53+ const n = typeof value === "string" ? Number(value) : value;
54+ return typeof n === "number" && Number.isFinite(n) && n > 0 ? Math.floor(n) : 0;
55+};
56+
57+/**
58+ * What an Anthropic model can do, from the capability tree its list
59+ * gives: `effort`, `thinking`, `vision`, and `tools` (every Claude model
60+ * takes tools). Unknown without the tree.
61+ */
62+export function anthropicCapabilities(tree: unknown): string[] {
63+ if (!isObject(tree)) return [];
64+ const out: string[] = [];
65+ if (supported(tree.effort)) out.push("effort");
66+ if (supported(tree.thinking)) out.push("thinking");
67+ out.push("tools");
68+ if (supported(tree.image_input)) out.push("vision");
69+ return out;
70+}
71+
72+/** One model from Anthropic's list, or null when it has no id. */
73+export function fromAnthropic(item: unknown): ProviderModel | null {
74+ if (!isObject(item) || typeof item.id !== "string" || !item.id.trim()) return null;
75+ return {
76+ id: item.id.trim(),
77+ name: typeof item.display_name === "string" ? item.display_name : "",
78+ kind: "chat",
79+ contextWindow: count(item.max_input_tokens),
80+ maxOutput: count(item.max_tokens),
81+ capabilities: anthropicCapabilities(item.capabilities),
82+ price: null,
83+ };
84+}
85+
86+/** Every model Anthropic lists for g1t's key, page by page. Throws with what went wrong. */
87+export async function listAnthropic(env: HostedRouting, fetcher: Fetch): Promise<ProviderModel[]> {
88+ const models: ProviderModel[] = [];
89+ let after: string | null = null;
90+ for (let page = 0; page < MAX_PAGES; page++) {
91+ const request = anthropicRequest(env, after);
92+ if (!request) throw new Error("No way to reach Anthropic: set AI_GATEWAY_TOKEN (or ANTHROPIC_API_KEY) on the models service.");
93+ const answer = await fetcher(request.url, { headers: request.headers });
94+ if (!answer.ok) throw new Error(`Anthropic's model list answered ${answer.status}: ${(await answer.text()).slice(0, 300)}`);
95+ const body = (await answer.json()) as Json;
96+ const data = Array.isArray(body.data) ? body.data : [];
97+ for (const item of data) {
98+ const model = fromAnthropic(item);
99+ if (model) models.push(model);
100+ }
101+ const last = typeof body.last_id === "string" ? body.last_id : null;
102+ if (body.has_more !== true || !last || data.length === 0) return models;
103+ after = last;
104+ }
105+ return models;
106+}
107+
108+/** A Workers AI model's kind, from its task. */
109+export function workersAiKind(task: unknown): string {
110+ const name = isObject(task) && typeof task.name === "string" ? task.name.toLowerCase() : "";
111+ if (name === "text generation") return "chat";
112+ if (name === "text embeddings") return "embeddings";
113+ return name ? name.replace(/\s+/g, "-") : "other";
114+}
115+
116+/** Dollars per million tokens to millionths, from Workers AI's price property. */
117+function listedPrice(value: unknown): ProviderModel["price"] {
118+ if (!Array.isArray(value)) return null;
119+ let input = 0;
120+ let output = 0;
121+ for (const entry of value) {
122+ if (!isObject(entry)) continue;
123+ const unit = typeof entry.unit === "string" ? entry.unit.toLowerCase() : "";
124+ const price = typeof entry.price === "number" ? entry.price : Number(entry.price);
125+ if (!Number.isFinite(price) || price <= 0 || !unit.includes("per m")) continue;
126+ if (unit.includes("input")) input = Math.round(price * 1_000_000);
127+ else if (unit.includes("output")) output = Math.round(price * 1_000_000);
128+ }
129+ return input > 0 ? { inputMicros: input, outputMicros: output } : null;
130+}
131+
132+/** One model from Workers AI's search, or null when it has no `@cf/…` name. */
133+export function fromWorkersAi(item: unknown): ProviderModel | null {
134+ if (!isObject(item) || typeof item.name !== "string" || !item.name.startsWith("@")) return null;
135+ const properties = new Map<string, unknown>();
136+ for (const property of Array.isArray(item.properties) ? item.properties : []) {
137+ if (isObject(property) && typeof property.property_id === "string") properties.set(property.property_id, property.value);
138+ }
139+ const kind = workersAiKind(item.task);
140+ const capabilities: string[] = [];
141+ if (properties.get("function_calling") === "true" || properties.get("function_calling") === true) capabilities.push("tools");
142+ if (kind === "embeddings") capabilities.push("embeddings");
143+ return {
144+ id: item.name,
145+ name: item.name.split("/").pop() ?? item.name,
146+ kind,
147+ contextWindow: count(properties.get("context_window")),
148+ maxOutput: 0,
149+ capabilities,
150+ price: listedPrice(properties.get("price")),
151+ };
152+}
153+
154+/** Every model Workers AI offers on g1t's account. Throws with what went wrong. */
155+export async function listWorkersAi(env: HostedRouting, fetcher: Fetch): Promise<ProviderModel[]> {
156+ const token = env.WORKERS_AI_TOKEN || env.AI_GATEWAY_TOKEN;
157+ if (!token || !env.CLOUDFLARE_ACCOUNT_ID) throw new Error("No Cloudflare token with Workers AI Read: set WORKERS_AI_TOKEN on the models service.");
158+ const models: ProviderModel[] = [];
159+ for (let page = 1; page <= MAX_PAGES; page++) {
160+ const url = `https://api.cloudflare.com/client/v4/accounts/${env.CLOUDFLARE_ACCOUNT_ID}/ai/models/search?per_page=${WORKERS_AI_PAGE}&page=${page}`;
161+ const answer = await fetcher(url, { headers: { authorization: `Bearer ${token}` } });
162+ if (!answer.ok) throw new Error(`Workers AI's model search answered ${answer.status}: ${(await answer.text()).slice(0, 300)}`);
163+ const body = (await answer.json()) as Json;
164+ const result = Array.isArray(body.result) ? body.result : [];
165+ for (const item of result) {
166+ const model = fromWorkersAi(item);
167+ if (model) models.push(model);
168+ }
169+ if (result.length < WORKERS_AI_PAGE) return models;
170+ }
171+ return models;
172+}
173+
174+/** The providers listed, each with how. */
175+export const PROVIDERS: { provider: string; list: (env: HostedRouting, fetcher: Fetch) => Promise<ProviderModel[]> }[] = [
176+ { provider: "anthropic", list: listAnthropic },
177+ { provider: "workers-ai", list: listWorkersAi },
178+];
179+
180+/** Where each provider's list goes: billing's `record_discovery`. */
181+export type Recorder = (provider: string, models: ProviderModel[], by: string, error: string | null) => Promise<DiscoveryResult>;
182+
183+/**
184+ * Lists every provider and records each list with billing, one at a time.
185+ * A provider that cannot be listed is recorded as a failed check (nothing
186+ * in the catalogue changes); the others still are.
187+ */
188+export async function discover(env: HostedRouting, fetcher: Fetch, record: Recorder, by: string): Promise<DiscoveryResult[]> {
189+ const who = by.trim().slice(0, 200) || "schedule";
190+ const results: DiscoveryResult[] = [];
191+ for (const { provider, list } of PROVIDERS) {
192+ let models: ProviderModel[] = [];
193+ let error: string | null = null;
194+ try {
195+ models = await list(env, fetcher);
196+ if (models.length === 0) error = "The list was empty.";
197+ } catch (failure) {
198+ error = failure instanceof Error ? failure.message : String(failure);
199+ }
200+ // Never a key in what is recorded.
201+ for (const secret of [env.AI_GATEWAY_TOKEN, env.ANTHROPIC_API_KEY, env.WORKERS_AI_TOKEN]) {
202+ if (secret && error) error = error.split(secret).join("[secret]");
203+ }
204+ results.push(await record(provider, error ? [] : models, who, error));
205+ }
206+ return results;
207+}
+36−0
2020 * at `/anthropic` in Anthropic's format or `/openai/v1` in OpenAI's, goes
2121 * to `serve.ts`, and is logged and charged to the workspace.
2222 */
23+import { WorkerEntrypoint } from "cloudflare:workers";
24+
2325 import {
26+ type DiscoveryResult,
2427 type GatewayModel,
28+ type ModelDiscoveryApi,
2529 type GatewayProvider,
2630 type ModelUpstream,
2731 type ServiceBinding,
3236 } from "@g1t/contracts";
3337
3438 import { openaiError } from "./chat";
39+import { discover } from "./discover";
3540 import { type AnthropicRequest, StreamTranslator, errorFromChat, estimateTokens, fromChat, toChat } from "./openai";
3641 import { isAnswer, tokenReport } from "./report";
3742 import { type HostedRouting, presentedToken, upstreamRequest } from "./route";
180185 };
181186 }
182187
188+// --- Discovery -----------------------------------------------------------------
189+
190+/** Lists every provider's models and records what changed with billing (discover.ts). */
191+function checkModels(env: Env, by: string): Promise<DiscoveryResult[]> {
192+ const billing = billingClient(env.BILLING);
193+ return discover(env, (url, init) => fetch(url, init), (provider, models, who, error) => billing.recordDiscovery(provider, models, who, error), by);
194+}
195+
196+/**
197+ * "Check for new models" in sudo, which binds this entrypoint. Only a
198+ * service binding reaches it: nothing at models.g1t.sh does.
199+ */
200+export class Discovery extends WorkerEntrypoint<Env> implements ModelDiscoveryApi {
201+ async check(by: string): Promise<DiscoveryResult[]> {
202+ return checkModels(this.env, typeof by === "string" ? by : "");
203+ }
204+}
205+
183206 export default {
207+ /** Once a day: the providers' model lists, against the catalogue. */
208+ async scheduled(_controller: ScheduledController, env: Env, ctx: ExecutionContext): Promise<void> {
209+ ctx.waitUntil(
210+ checkModels(env, "schedule")
211+ .then((results) => {
212+ for (const result of results) {
213+ console.log(`models: ${result.provider} listed ${result.listed}, new ${result.added.length}, gone ${result.deprecated.length}${result.error ? `, failed: ${result.error}` : ""}`);
214+ }
215+ })
216+ .catch((error) => console.error("models: checking the providers' lists failed", error)),
217+ );
218+ },
219+
184220 async fetch(request: Request, env: Env, ctx: ExecutionContext): Promise<Response> {
185221 const url = new URL(request.url);
186222 if (url.pathname === "/" || url.pathname === "") {
+4−0
2222 "AI_GATEWAY_ID": "g1t",
2323 "CLOUDFLARE_ACCOUNT_ID": "1e6f2cffa3f445920836e8ebe446bb58"
2424 },
25+ // Once a day, after billing's daily run: every provider's model list,
26+ // recorded with billing's catalogue (src/discover.ts). sudo's "Check
27+ // for new models" does the same through the `Discovery` entrypoint.
28+ "triggers": { "crons": ["29 5 * * *"] },
2529 // Secrets: AI_GATEWAY_TOKEN; ANTHROPIC_API_KEY only if the gateway does
2630 // not hold g1t's key; WORKERS_AI_TOKEN, a Cloudflare API token with
2731 // Workers AI Read on this account, for the AI Gateway's open models
+14−4
6767 type RouteSignals,
6868 canReachModel,
6969 changeSize,
70+ effortFor,
7071 failuresInARow,
7172 gatewaySession,
7273 leftLowConfidence,
7374 modelEnv,
7475 outcomesOf,
75− parseRouting,
7676 route,
77+ routingReader,
7778 taskOf,
7879 tierVars,
7980 } from "./model-env";
81+
82+/**
83+ * The routing in force: staff's defaults from billing, read at most once a
84+ * minute per isolate, on AGENT_ROUTING (alone when billing cannot be read).
85+ */
86+const routingNow = routingReader();
8087 import { hubContext } from "./hub";
8188 import { hostedOpen } from "./hosted";
8289 import { delegateInput, noModelMessage, notStarted, queued, started } from "./delegate";
162169 * change; `largeLabels`, `frontierLabels` and `smallLabels`, issue
163170 * labels that move work; `frontierAfter`, failures in a row before the
164171 * frontier tier; `learning`, how a repository's own runs move it.
165− * Anything left out takes the default.
172+ * Anything left out takes the default. Staff's defaults in sudo
173+ * (billing's `model_defaults`) replace `tiers`, `tasks` and `effort`
174+ * whenever billing can be read.
166175 */
167176 AGENT_ROUTING?: string;
168177 /**
13881397 requestedBy: string | null,
13891398 input: RouteInput = {},
13901399 ): Promise<Result<Record<string, string>>> {
1391− const routing = parseRouting(this.env.AGENT_ROUTING);
1400+ // Staff's defaults from billing's catalogue, on AGENT_ROUTING.
1401+ const routing = await routingNow(this.env.AGENT_ROUTING, () => billingClient(this.env.BILLING).modelDefaults());
13921402 const task = taskOf(kind);
13931403 const signals: RouteSignals = { ...input };
13941404 if (input.viewer) {
14641474 }
14651475 : modelEnv(this.env, routing, task, tier, direct ? { ...tags, session: direct } : tags);
14661476 // How hard it thinks, by the kind of work, on g1t's tiers.
1467− const effort = named ? undefined : routing.effort[kind];
1477+ const effort = named ? undefined : effortFor(routing, kind, tier);
14681478 if (effort) vars.CLAUDE_CODE_EFFORT_LEVEL = effort;
14691479 // Why this model: shown on the run and at the top of its session.
14701480 vars.AGENT_MODEL_REASON = effort ? `${reason.replace(/\.$/, "")}, at ${effort} effort.` : reason;
+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

This change is too large to show in full.