Skip to content

Commit

AI Gateway: OpenAI's format, open models, and your own providers

models.g1t.sh/openai/v1 answers POST /chat/completions (streamed or not, function tools, response_format, reasoning_effort, cached-token usage), POST /embeddings and GET /models, on the same tokens, scope, admission, charging and log as /anthropic. Any model works in either format: the proxy translates Anthropic <-> OpenAI both ways, streaming and tool calls included, carrying Claude's thinking blocks in the first tool call's id. - Model ids: anthropic/claude-*, workers-ai/@cf/..., bare ids still work. - Open models on g1t's account through Workers AI via the AI Gateway (WORKERS_AI_TOKEN), priced in gateway_models (billing migration 0046). - 0046 also fixes Sonnet 5.5 cache reads ($0.10), prices one-hour cache writes, and adds Claude Haiku 5.5 with long-prompt pricing (threshold and over_ prices); charging splits 5m/1h cache writes. - Your own providers: any model connection (Anthropic or OpenAI key, any compatible endpoint) takes the models in config.gateway_models (ids, prefix*, ns/* stripped); counted, never charged. New update_integration (PATCH /workspaces/{ws}/integrations/{id}) rotates keys write-only. Provider errors are scrubbed of keys. - Log, REST and MCP show format, provider, connection, cache_write_hour. - Web: both base URLs and Served by on the AI Gateway page; AI Gateway models and key replacement on Integrations. - Docs: AI Gateway guide (both formats, models and prices, own providers, errors), MCP and integrations reference, OpenAPI, llms.txt, pricing, billing operations.

syntaqxcommitted Parent502b69bBrowse files
44 files+642−1070/44 viewed
+2−0
359359 &[
360360 Op::ListIntegrations,
361361 Op::ConnectIntegration,
362+ Op::UpdateIntegration,
362363 Op::DisconnectIntegration,
363364 Op::TestIntegration,
364365 Op::GetModelRoutes,
461462 Op::ListEvents => "List repository events",
462463 Op::ListIntegrations => "List integrations",
463464 Op::ConnectIntegration => "Connect an integration",
465+ Op::UpdateIntegration => "Update an integration",
464466 Op::DisconnectIntegration => "Disconnect an integration",
465467 Op::TestIntegration => "Test an integration",
466468 Op::GetContext => "Look up a ticket",
+46−5
169169 ListEvents,
170170 ListIntegrations,
171171 ConnectIntegration,
172+ UpdateIntegration,
172173 DisconnectIntegration,
173174 TestIntegration,
174175 GetContext,
629630 }
630631
631632 impl Op {
632− pub const ALL: [Op; 219] = [
633+ pub const ALL: [Op; 220] = [
633634 Op::Whoami,
634635 Op::GetWorkspace,
635636 Op::CreateWorkspace,
711712 Op::ListEvents,
712713 Op::ListIntegrations,
713714 Op::ConnectIntegration,
715+ Op::UpdateIntegration,
714716 Op::DisconnectIntegration,
715717 Op::TestIntegration,
716718 Op::GetContext,
939941 Op::ListEvents => "list_events",
940942 Op::ListIntegrations => "list_integrations",
941943 Op::ConnectIntegration => "connect_integration",
944+ Op::UpdateIntegration => "update_integration",
942945 Op::DisconnectIntegration => "disconnect_integration",
943946 Op::TestIntegration => "test_integration",
944947 Op::GetContext => "get_context",
12671270 "A workspace's integrations: its own model provider, the alert sources that open issues (Sentry, Datadog, webhooks), and the trackers whose tickets agents can read (Jira, Linear). Secrets are never returned. Members only."
12681271 }
12691272 Op::ConnectIntegration => {
1270− "Connect a workspace to an outside system. provider is a model provider (anthropic, openai, gemini, xai, mistral, deepseek, azure_openai, openrouter, groq, together, fireworks, cerebras, anthropic_endpoint or openai_endpoint: your own key, billed by that provider, and free on g1t while it is being built out; a workspace can connect several and route each kind of work with set_model_routes), or sentry, datadog, webhook, jira or linear. config holds the settings each needs; secret is the API key or token. For datadog and webhook, g1t makes the signing secret and returns it once. Owners only."
1273+ "Connect a workspace to an outside system. provider is a model provider (anthropic, openai, gemini, xai, mistral, deepseek, azure_openai, openrouter, groq, together, fireworks, cerebras, anthropic_endpoint or openai_endpoint: your own key, billed by that provider, and free on g1t while it is being built out; a workspace can connect several and route each kind of work with set_model_routes), or sentry, datadog, webhook, jira or linear. config holds the settings each needs; secret is the API key or token, kept encrypted and never returned (secret_hint shows its last four characters). For a model provider, config.gateway_models chooses which AI Gateway requests go to it by the model they name: ids such as gpt-5.5, or prefixes ending in * such as gpt-* or ollama/* (a /* prefix is taken off before sending); absent, an Anthropic key or Anthropic-compatible endpoint takes claude-* and the others take nothing. Requests on the workspace's own provider are counted and never charged. For datadog and webhook, g1t makes the signing secret and returns it once. Owners only."
1274+ }
1275+ Op::UpdateIntegration => {
1276+ "Change an integration: its name, its config (replaced whole when given) or its secret (a new key replaces the old one, write-only). Use it to rotate a model provider's key or to choose its config.gateway_models, the AI Gateway models it takes. Fields left out are kept. Secrets are never returned. Owners only."
12711277 }
12721278 Op::DisconnectIntegration => {
12731279 "Remove an integration and its secrets. Agents already running on a model provider being removed stop reaching it. Owners only."
15171523 "Who a workspace's invoices are made out to: the billing `email`, `name`, `address`, tax ID (`tax_id_type`, `tax_id`), `po_number` and the invoices' `language`, with the default `payment_method` as far as it is safe to show (its kind, brand, last four digits and expiry). `customer` is false until the workspace has been set up to pay. Tax is worked out from the address: `tax_location` says whether it is enough for that (a country, and in the US a ZIP code), `tax_address_needed_at` is set while g1t is holding a charge for want of one, `tax_id_status` is Stripe's check of the tax ID (`pending`, `verified`, `unverified` or `unavailable`), and `tax_exempt` is `none`, `exempt` or `reverse`. Members of the workspace only."
15181524 }
15191525 Op::ListGatewayRequests => {
1520− "A workspace's recent AI Gateway requests, newest first: each with its `id`, `created_at`, `model`, the access token that sent it (`token_id`, `token_name`), its tokens by kind (`input`, `output`, `cache_read`, `cache_write`), what they cost at the model's price (`cost_micros`) and what the workspace was charged (`charged_micros`, before included usage and AI credit paid for it; 0 on the workspace's own provider key, `own_key`), the HTTP `status` it was answered with, whether it was `streamed`, `duration_ms`, and `error` for one that was refused or failed. Prompts and answers are never kept. `limit` is how many, 50 unless given and 200 at most; pass `next` from one page as `before` for the next. Requests are kept `retention_days` (30). Members of the workspace only."
1526+ "A workspace's recent AI Gateway requests, newest first: each with its `id`, `created_at`, `model`, the access token that sent it (`token_id`, `token_name`), its tokens by kind (`input`, `output`, `cache_read`, `cache_write`, and of those writes `cache_write_hour` to the hour-long cache), the `format` it was sent in (`anthropic` or `openai`), who served it (`provider`: `anthropic` or `workers-ai` on g1t's account, the connection's provider on the workspace's own, and `connection`, that connection's name), what they cost at the model's price (`cost_micros`) and what the workspace was charged (`charged_micros`, before included usage and AI credit paid for it; 0 on the workspace's own provider key, `own_key`), the HTTP `status` it was answered with, whether it was `streamed`, `duration_ms`, and `error` for one that was refused or failed. Prompts and answers are never kept. `limit` is how many, 50 unless given and 200 at most; pass `next` from one page as `before` for the next. Requests are kept `retention_days` (30). Members of the workspace only."
15211527 }
15221528 Op::ListUserTeams => {
15231529 "The teams someone is in within a workspace, as list_teams describes them, leaving out secret teams you cannot see. Members of the workspace only."
22602266 "name": { "type": "string", "description": "What to call it. The provider's name if left out." },
22612267 "config": {
22622268 "type": "object",
2263− "description": "Settings. repo (owner/name) is where alerts open issues; assign puts an agent on each; label names the label (bug). organization is the Sentry org's slug. site is Jira's address; email the account its token belongs to; keys the project or team keys it answers for. base_url and auth_header (x-api-key or authorization) are for your own endpoint; model overrides the model for every kind of work. write_back (default true) tells the outside system when the work lands.",
2269+ "description": "Settings. repo (owner/name) is where alerts open issues; assign puts an agent on each; label names the label (bug). organization is the Sentry org's slug. site is Jira's address; email the account its token belongs to; keys the project or team keys it answers for. base_url and auth_header (x-api-key or authorization) are for your own endpoint; model overrides the model for every kind of work; gateway_models (model ids, or prefixes ending in * such as gpt-* or ollama/*) chooses which AI Gateway requests go to a model provider. write_back (default true) tells the outside system when the work lands.",
22642270 },
2265− "secret": { "type": "string", "description": "The API key or token g1t uses to call it." },
2271+ "secret": { "type": "string", "description": "The API key or token g1t uses to call it. Write-only: kept encrypted, never returned." },
22662272 "signing_secret": { "type": "string", "description": "For sentry: the integration's client secret." },
22672273 }),
22682274 &["workspace", "provider"],
22692275 ),
2276+ Op::UpdateIntegration => object(
2277+ json!({
2278+ "workspace": workspace_schema(),
2279+ "id": { "type": "string", "description": "The integration's id." },
2280+ "name": { "type": "string", "description": "A new name." },
2281+ "config": {
2282+ "type": "object",
2283+ "description": "Its settings, replaced whole: the same fields as connect_integration's config. For a model provider, gateway_models chooses the AI Gateway models it takes.",
2284+ },
2285+ "secret": { "type": "string", "description": "A new API key or token, replacing the old one. Write-only: kept encrypted, never returned." },
2286+ "signing_secret": { "type": "string", "description": "For sentry: a new client secret." },
2287+ }),
2288+ &["workspace", "id"],
2289+ ),
22702290 Op::GetModelRoutes => object(json!({ "workspace": workspace_schema() }), &["workspace"]),
22712291 Op::ListWebhooks => object(hook_owner(json!({})), &[]),
22722292 Op::ListWorkflows => repo_only(),
29102930 | Op::CreateRepo
29112931 | Op::ListIntegrations
29122932 | Op::ConnectIntegration
2933+ | Op::UpdateIntegration
29132934 | Op::DisconnectIntegration
29142935 | Op::TestIntegration
29152936 | Op::GetModelRoutes
40944115 )
40954116 .await
40964117 }
4118+ Op::UpdateIntegration => {
4119+ let config = match &input["config"] {
4120+ Value::Null => Value::Null,
4121+ config => camel_keys(config),
4122+ };
4123+ pass(
4124+ integrations,
4125+ "update",
4126+ &json!({
4127+ "actor": actor(),
4128+ "workspace": workspace(),
4129+ "id": text(input, "id"),
4130+ "name": optional_text(input, "name"),
4131+ "config": config,
4132+ "secret": optional_text(input, "secret"),
4133+ "signingSecret": optional_text(input, "signing_secret"),
4134+ }),
4135+ )
4136+ .await
4137+ }
40974138 Op::DisconnectIntegration | Op::TestIntegration => {
40984139 pass(
40994140 integrations,
+67−4
37133713 },
37143714 "notes": "`signing_secret` is set only when g1t made it, for `datadog` and `webhook`, and is shown only this once. A model provider is tested as it is connected. See [integrations](/guides/integrations/) and [model providers](/guides/models/)."
37153715 },
3716+ "update_integration": {
3717+ "params": {
3718+ "workspace": "flagon-io",
3719+ "id": "con_01kpx5c2d8e4f6g0h2j4k6m8n0"
3720+ },
3721+ "request": {
3722+ "config": {
3723+ "base_url": "https://gpu.flagon.dev/v1",
3724+ "gateway_models": ["ollama/*"]
3725+ },
3726+ "secret": "sk-office-gpu-key"
3727+ },
3728+ "response": {
3729+ "id": "con_01kpx5c2d8e4f6g0h2j4k6m8n0",
3730+ "workspace": "flagon-io",
3731+ "provider": "openai_endpoint",
3732+ "kind": "models",
3733+ "name": "Office GPU",
3734+ "config": {
3735+ "write_back": true,
3736+ "base_url": "https://gpu.flagon.dev/v1",
3737+ "gateway_models": ["ollama/*"]
3738+ },
3739+ "secret_hint": "…-key",
3740+ "webhook_url": null,
3741+ "created_by": "syntaqx",
3742+ "created_at": "2026-10-07T12:10:44.512Z",
3743+ "last_used_at": "2026-10-07T14:00:02.000Z",
3744+ "last_error": null,
3745+ "models": ["llama3.3:70b", "qwen3-coder:30b"]
3746+ },
3747+ "notes": "The secret is write-only: it is kept encrypted and only its last four characters come back, as `secret_hint`. With `gateway_models` set to `ollama/*`, an AI Gateway request for `ollama/qwen3-coder:30b` reaches this endpoint as `qwen3-coder:30b`, on the workspace's own account. See [the AI Gateway guide](/guides/ai-gateway/#your-own-providers)."
3748+ },
37163749 "disconnect_integration": {
37173750 "params": {
37183751 "workspace": "flagon-io",
59665999 "workspace": "flagon-io"
59676000 },
59686001 "query": {
5969− "limit": 2
6002+ "limit": 3
59706003 },
59716004 "response": {
59726005 "requests": [
59806013 "output": 512,
59816014 "cache_read": 12000,
59826015 "cache_write": 0,
5983− "cost_micros": 11200,
5984− "charged_micros": 11200,
6016+ "cache_write_hour": 0,
6017+ "cost_micros": 10000,
6018+ "charged_micros": 10000,
59856019 "status": 200,
59866020 "own_key": false,
6021+ "format": "anthropic",
6022+ "provider": "anthropic",
6023+ "connection": null,
59876024 "streamed": true,
59886025 "duration_ms": 4210,
59896026 "error": null
59906027 },
59916028 {
6029+ "id": "gw_9b3d1f7a5c2e4b6d8f0a1c33",
6030+ "created_at": "2026-10-07T14:00:02.000Z",
6031+ "model": "qwen3-coder:30b",
6032+ "token_id": "tok_01kkr5a2b8c4d6e0f3g5h7j9k1",
6033+ "token_name": "triage-bot",
6034+ "input": 2210,
6035+ "output": 96,
6036+ "cache_read": 0,
6037+ "cache_write": 0,
6038+ "cache_write_hour": 0,
6039+ "cost_micros": 0,
6040+ "charged_micros": 0,
6041+ "status": 200,
6042+ "own_key": true,
6043+ "format": "openai",
6044+ "provider": "openai_endpoint",
6045+ "connection": "Office GPU",
6046+ "streamed": false,
6047+ "duration_ms": 1830,
6048+ "error": null
6049+ },
6050+ {
59926051 "id": "gw_1a7e3c5b9d2f4e6a8c0b2d41",
59936052 "created_at": "2026-10-07T13:58:40.000Z",
59946053 "model": "claude-opus-5-5",
59986057 "output": 0,
59996058 "cache_read": 0,
60006059 "cache_write": 0,
6060+ "cache_write_hour": 0,
60016061 "cost_micros": 0,
60026062 "charged_micros": 0,
60036063 "status": 402,
60046064 "own_key": false,
6065+ "format": "openai",
6066+ "provider": "",
6067+ "connection": null,
60056068 "streamed": false,
60066069 "duration_ms": 38,
60076070 "error": "The flagon-io workspace is out of AI credit and has used this month's included usage, so the AI Gateway refuses requests to g1t's models. An owner can buy AI credit or turn on auto-reload at /flagon-io/-/billing#ai-credit."
60106073 "next": "gw_1a7e3c5b9d2f4e6a8c0b2d41",
60116074 "retention_days": 30
60126075 },
6013− "notes": "To send requests, see the AI Gateway guide: POST https://models.g1t.sh/anthropic/v1/messages with a workspace access token that has `models:write`. `charged_micros` is what the request was charged before included usage and AI credit paid for it; the payment itself is on the statement as AI Gateway. Pass `next` as `before` for the next page; it is null on the last."
6076+ "notes": "To send requests, see the AI Gateway guide: POST https://models.g1t.sh/anthropic/v1/messages (Anthropic's format) or https://models.g1t.sh/openai/v1/chat/completions (OpenAI's) with a workspace access token that has `models:write`. `format` is the format a request was sent in; `provider` who served it (`anthropic` or `workers-ai` on g1t's account, the connection's provider on the workspace's own, empty when it was refused before reaching one), and `connection` the workspace's own connection by name. `cache_write_hour` is the part of `cache_write` written to the hour-long cache. `charged_micros` is what the request was charged before included usage and AI credit paid for it; the payment itself is on the statement as AI Gateway. Pass `next` as `before` for the next page; it is null on the last."
60146077 },
60156078 "list_invoices": {
60166079 "params": {
+6−0
485485 &[],
486486 ),
487487 route(
488+ "PATCH",
489+ "/workspaces/:workspace/integrations/:id",
490+ Op::UpdateIntegration,
491+ &[],
492+ ),
493+ route(
488494 "DELETE",
489495 "/workspaces/:workspace/integrations/:id",
490496 Op::DisconnectIntegration,
+2−0
287287 a("revoke_invite", Op::RevokeWorkspaceInvite, "Revoke a pending invite"),
288288 a("list_integrations", Op::ListIntegrations, "Model providers, alert sources, trackers"),
289289 a("connect_integration", Op::ConnectIntegration, "Connect one"),
290+ a("update_integration", Op::UpdateIntegration, "Change one: rotate its key, choose its AI Gateway models"),
290291 a("disconnect_integration", Op::DisconnectIntegration, "Remove one"),
291292 a("test_integration", Op::TestIntegration, "Check its credentials"),
292293 a("get_model_routes", Op::GetModelRoutes, "Where each kind of work's model requests go"),
417418 | Op::RemoveEmail
418419 | Op::RemoveCollaborator
419420 | Op::DisconnectIntegration
421+ | Op::UpdateIntegration
420422 | Op::DeleteWebhook
421423 | Op::DeleteActionsSecret
422424 | Op::DeleteActionsVariable
+333−74
11 ---
22 title: AI Gateway
3−description: Send your own code's model requests through g1t with a workspace access token, paid from AI credit at the model's price, with a log of every request.
3+description: Send your own code's model requests through g1t in Anthropic's or OpenAI's format, with a workspace access token, to Claude, open models or your own providers, with a log of every request.
44 ---
55
6−The AI Gateway takes model requests from your own code, in Anthropic's
7−Messages format, and sends them to the model. You point an Anthropic SDK,
8−Claude Code or anything else that speaks that format at one base URL and
9−give it a workspace access token as its API key:
6+The AI Gateway takes model requests from your own code and sends them to
7+the model. It speaks two formats, and any model works in either:
8+
9+| Format | Base URL | For |
10+| --- | --- | --- |
11+| Anthropic's Messages API | `https://models.g1t.sh/anthropic` | Anthropic's SDKs, Claude Code, and anything else that speaks that format |
12+| OpenAI's Chat Completions API | `https://models.g1t.sh/openai/v1` | OpenAI's SDKs, and any tool that lets you set an OpenAI-compatible base URL |
1013
11−| | |
12−| --- | --- |
13−| Base URL | `https://models.g1t.sh/anthropic` |
14−| API key | A workspace access token (`g1t_…`) with the `models:write` scope |
14+In both, the API key is a workspace access token (`g1t_…`) with the
15+`models:write` scope.
1516
16−Each request is charged to the workspace at the model's price and paid from
17−the plan's included usage and [AI credit](/guides/usage-and-billing/#ai-credit).
18−While the gateway is in beta there is no markup. If the workspace has
19−connected its own Anthropic key, requests go there instead and cost
20−nothing on g1t. Every request is logged with its model, tokens, cost and
21−status. Prompts and answers are never kept.
17+The model a request names decides where it goes:
18+
19+- **g1t's models**: Claude on Anthropic, and open models on Workers AI.
20+ Each request is charged to the workspace at the model's price and paid
21+ from the plan's included usage and
22+ [AI credit](/guides/usage-and-billing/#ai-credit). While the gateway is in
23+ beta there is no markup.
24+- **Your own providers**: an Anthropic key, an OpenAI key, or any endpoint
25+ that speaks either API, connected under
26+ [Integrations](/guides/models/#connect-a-provider). You choose which models
27+ go to each. Those requests are counted and never charged on g1t.
28+
29+Every request is logged with its format, who served it, its model, tokens,
30+cost and status. Prompts and answers are never kept.
2231
2332 ## Before you start
2433
2635
2736 - **The workspace on the g1t plan**, with AI credit or this month's
2837 included usage left. See [AI credit](/guides/usage-and-billing/#ai-credit).
29−- **The workspace's own Anthropic key**, connected under
30− [Integrations](/guides/models/#connect-a-provider). Then the plan is not
31− needed and nothing is charged.
38+- **One of the workspace's own model providers**, connected under
39+ [Integrations](#your-own-providers). Requests for the models it takes need
40+ no plan and cost nothing on g1t.
3241
3342 ## Make a token
3443
5160
5261 ## Send a request
5362
54−The gateway answers the same routes as Anthropic's API, below the base URL:
63+The token goes in `x-api-key` or in `Authorization: Bearer`, in either
64+format. The gateway answers these routes:
5565
5666 | Route | What it does |
5767 | --- | --- |
5868 | `POST /anthropic/v1/messages` | A message, streamed (`"stream": true`) or whole. Logged and charged. |
59−| `POST /anthropic/v1/messages/count_tokens` | Counts a request's input tokens. Not logged, and costs nothing. |
69+| `POST /anthropic/v1/messages/count_tokens` | Counts a request's input tokens. Not logged, and costs nothing. For a model that does not speak Anthropic's API, an estimate. |
70+| `POST /openai/v1/chat/completions` | A chat completion, streamed or whole. Logged and charged. |
71+| `POST /openai/v1/embeddings` | Embeddings, from an embeddings model. Logged and charged by their input tokens. |
72+| `GET /openai/v1/models` | The models this workspace can use, with g1t's prices. |
6073
61−The token goes in `x-api-key`, or in `Authorization: Bearer`. Request and
62−answer bodies are Anthropic's, unchanged, and so are streamed events.
74+### Anthropic's format
6375
6476 With curl:
6577
6981 -H "anthropic-version: 2023-06-01" \
7082 -H "content-type: application/json" \
7183 -d '{
72− "model": "claude-sonnet-5-5",
84+ "model": "claude-haiku-5-5",
7385 "max_tokens": 1024,
7486 "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }]
7587 }'
8698 });
8799
88100 const message = await client.messages.create({
89− model: "claude-sonnet-5-5",
101+ model: "claude-haiku-5-5",
90102 max_tokens: 1024,
91103 messages: [{ role: "user", content: "Write a commit message for: fix the login redirect" }],
92104 });
105117 )
106118
107119 message = client.messages.create(
108− model="claude-sonnet-5-5",
120+ model="claude-haiku-5-5",
109121 max_tokens=1024,
110122 messages=[{"role": "user", "content": "Write a commit message for: fix the login redirect"}],
111123 )
112124 ```
113125
114−### Claude Code
126+Request and answer bodies are Anthropic's, and so are streamed events. To
127+an open model, such as `workers-ai/@cf/openai/gpt-oss-120b`, the request is
128+translated: messages, system prompt, images, tools and tool results,
129+`tool_choice`, stop sequences, `output_config.effort` (as
130+`reasoning_effort`, `xhigh` and `max` as `high`) and `output_config.format`
131+(as a JSON schema). The answer comes back as an Anthropic message, tool
132+calls included. Server tools have no counterpart there and are refused on
133+g1t's models.
134+
135+### OpenAI's format
136+
137+With curl:
138+
139+```sh
140+curl https://models.g1t.sh/openai/v1/chat/completions \
141+ -H "Authorization: Bearer $G1T_TOKEN" \
142+ -H "content-type: application/json" \
143+ -d '{
144+ "model": "anthropic/claude-haiku-5-5",
145+ "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }]
146+ }'
147+```
148+
149+With OpenAI's TypeScript SDK:
150+
151+```ts
152+import OpenAI from "openai";
153+
154+const client = new OpenAI({
155+ baseURL: "https://models.g1t.sh/openai/v1",
156+ apiKey: process.env.G1T_TOKEN,
157+});
158+
159+const completion = await client.chat.completions.create({
160+ model: "workers-ai/@cf/openai/gpt-oss-120b",
161+ messages: [{ role: "user", content: "Label this issue: the login page is blank on Safari" }],
162+});
163+```
164+
165+With OpenAI's Python SDK:
166+
167+```python
168+import os
169+
170+from openai import OpenAI
171+
172+client = OpenAI(
173+ base_url="https://models.g1t.sh/openai/v1",
174+ api_key=os.environ["G1T_TOKEN"],
175+)
176+
177+completion = client.chat.completions.create(
178+ model="anthropic/claude-sonnet-5-5",
179+ messages=[{"role": "user", "content": "Summarize this diff in one sentence."}],
180+ stream=True,
181+ stream_options={"include_usage": True},
182+)
183+for chunk in completion:
184+ print(chunk.choices[0].delta.content or "" if chunk.choices else "", end="")
185+```
186+
187+Embeddings:
188+
189+```sh
190+curl https://models.g1t.sh/openai/v1/embeddings \
191+ -H "Authorization: Bearer $G1T_TOKEN" \
192+ -H "content-type: application/json" \
193+ -d '{ "model": "workers-ai/@cf/baai/bge-m3", "input": ["fix the login redirect"] }'
194+```
195+
196+What OpenAI's format supports, to any model:
197+
198+| In the request | |
199+| --- | --- |
200+| `messages` | `system`, `developer`, `user`, `assistant` and `tool` messages. User content can be text, `image_url` (a URL or a `data:` URL) and, to Claude, `file` with `file_data` (a PDF as a `data:` URL). |
201+| `tools`, `tool_choice`, `parallel_tool_calls` | Function tools. To Claude, `required` is `any` and a named function is that tool. |
202+| `stream`, `stream_options.include_usage` | Server-sent chunks, ending with `data: [DONE]`. With `include_usage`, a last chunk carries the usage. |
203+| `max_tokens`, `max_completion_tokens` | To Claude, 8,192 when neither is given. |
204+| `temperature`, `top_p`, `stop`, `user` | As given. Some models refuse sampling settings. |
205+| `reasoning_effort` | To Claude, `output_config.effort`; `minimal` and `none` are `low`. |
206+| `response_format` | `json_schema` is a JSON schema the answer follows. `json_object` asks Claude for one JSON object. |
207+| `thinking` | To Claude, passed as it is, for a caller that sets Anthropic's thinking. |
208+
209+To Claude, the answer's `usage` counts cached tokens in `prompt_tokens`,
210+with `prompt_tokens_details.cached_tokens`; Claude's thinking comes back as
211+`reasoning_content`. Claude's thinking blocks must go back with the tool
212+calls they led to, so the gateway carries them in the first tool call's
213+`id`: send the `id` back unchanged in the assistant message and the `tool`
214+message, as OpenAI's SDKs do. `n` above 1 is refused for Claude.
215+
216+### Claude Code and other tools
115217
116−Set two environment variables before you start it:
218+Claude Code speaks Anthropic's format. Set two environment variables
219+before you start it:
117220
118221 ```sh
119222 export ANTHROPIC_BASE_URL=https://models.g1t.sh/anthropic
123226
124227 Claude Code's own small requests go to Claude Haiku 4.5, which the gateway
125228 offers. To choose the main model, also set `ANTHROPIC_MODEL`, such as
126−`claude-opus-5-5`.
229+`claude-opus-5-5`, or `ANTHROPIC_SMALL_FAST_MODEL`, such as
230+`claude-haiku-5-5`.
127231
232+Any other tool that lets you set an OpenAI-compatible base URL and key
233+works the same way: give it `https://models.g1t.sh/openai/v1` and the
234+token, and a model id from [the models list](#models).
235+
128236 ## Models
129237
130−On g1t's models the gateway offers these. Prices are per million tokens,
131−the provider's list price; cache writes are five-minute ones.
238+### Model ids
132239
133−| Model | `model` | Input | Output | Cache reads | Cache writes |
134−| --- | --- | --- | --- | --- | --- |
135−| Claude Opus 5.5 | `claude-opus-5-5` | $4.00 | $20.00 | $0.20 | $5.00 |
136−| Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2.00 | $10.00 | $0.20 | $2.50 |
137−| Claude Haiku 4.5 | `claude-haiku-4-5`, `claude-haiku-4-5-20251001` | $1.00 | $5.00 | $0.10 | $1.25 |
240+A request names a model:
138241
139−A request for any other model is refused with `400` before it reaches the
140−provider, and the error names the models offered. On the workspace's own
141−key, a request can name any model that key can use.
242+| Id | Goes to |
243+| --- | --- |
244+| `anthropic/claude-sonnet-5-5` | Claude on g1t's account, in either format |
245+| `claude-sonnet-5-5` | The same, as Anthropic's API names it |
246+| `workers-ai/@cf/openai/gpt-oss-120b` | An open model on g1t's account, in either format |
247+| `@cf/openai/gpt-oss-120b` | The same |
248+| Any id one of your own providers takes, such as `gpt-5.5` or `ollama/llama3.3` | That provider, with its key. See [your own providers](#your-own-providers). |
142249
143−On g1t's models a request is charged only by its tokens, so what the
144−provider bills some other way is refused with `400` for now:
250+Your own providers come first: when one of them takes a model, the request
251+goes there, even one that names a model g1t offers. An Anthropic key takes
252+`claude-*` unless you choose otherwise, so with one connected, Claude goes to
253+your key.
145254
255+`GET /openai/v1/models` lists what the workspace can use: its own
256+providers' models first, then g1t's, cheapest Claude first. Each has
257+`billed_to` (`workspace` or `g1t`), `connection` (your provider's name) and,
258+on g1t's models, `pricing` in dollars per million tokens:
259+
260+```json
261+{
262+ "object": "list",
263+ "data": [
264+ {
265+ "id": "anthropic/claude-haiku-5-5",
266+ "object": "model",
267+ "created": 0,
268+ "owned_by": "anthropic",
269+ "name": "Claude Haiku 5.5",
270+ "kind": "chat",
271+ "billed_to": "g1t",
272+ "connection": null,
273+ "pricing": {
274+ "currency": "usd",
275+ "input": 0.1,
276+ "output": 0.5,
277+ "cache_read": 0.01,
278+ "cache_write": 0.125,
279+ "cache_write_1h": 0.2,
280+ "long_prompt": { "above_tokens": 100000, "input": 0.5, "output": 2.5, "cache_read": 0.05, "cache_write": 0.625, "cache_write_1h": 1 }
281+ }
282+ }
283+ ]
284+}
285+```
286+
287+### Claude, on Anthropic
288+
289+Prices are per million tokens, Anthropic's list price. Cache writes are
290+five-minute ones; one-hour cache writes (`"ttl": "1h"`) cost twice the
291+input price.
292+
293+| Model | `model` | Input | Output | Cache reads | Cache writes | One-hour cache writes |
294+| --- | --- | --- | --- | --- | --- | --- |
295+| Claude Haiku 5.5 | `claude-haiku-5-5` | $0.10 | $0.50 | $0.01 | $0.125 | $0.20 |
296+| Claude Haiku 5.5, prompts over 100,000 tokens | `claude-haiku-5-5` | $0.50 | $2.50 | $0.05 | $0.625 | $1.00 |
297+| Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2.00 | $10.00 | $0.10 | $2.50 | $4.00 |
298+| Claude Opus 5.5 | `claude-opus-5-5` | $4.00 | $20.00 | $0.20 | $5.00 | $8.00 |
299+| Claude Haiku 4.5 | `claude-haiku-4-5`, `claude-haiku-4-5-20251001` | $1.00 | $5.00 | $0.10 | $1.25 | $2.00 |
300+
301+Claude Haiku 5.5 is the cheapest Claude and the one to start with. It is
302+priced by the prompt's length: a request whose prompt (its input, cache
303+read and cache write tokens) is longer than 100,000 tokens is charged
304+entirely at the higher prices. It takes effort, like Opus: set
305+`output_config.effort` in Anthropic's format, or `reasoning_effort` in
306+OpenAI's.
307+
308+### Open models, on Workers AI
309+
310+Prices are per million tokens, Cloudflare's list price. Workers AI has no
311+prompt-cache price: cached tokens, where a model reports them, cost what
312+input does.
313+
314+| Model | `model` | Input | Output |
315+| --- | --- | --- | --- |
316+| GLM-5.3 Flash | `workers-ai/@cf/zai-org/glm-5.3-flash` | $0.15 | $0.50 |
317+| gpt-oss-20b | `workers-ai/@cf/openai/gpt-oss-20b` | $0.20 | $0.30 |
318+| Llama 4 Scout | `workers-ai/@cf/meta/llama-4-scout-17b-16e-instruct` | $0.27 | $0.85 |
319+| gpt-oss-120b | `workers-ai/@cf/openai/gpt-oss-120b` | $0.35 | $0.75 |
320+| Mistral Small 3.1 | `workers-ai/@cf/mistralai/mistral-small-3.1-24b-instruct` | $0.351 | $0.555 |
321+| DeepSeek V4 Flash | `workers-ai/@cf/deepseek-ai/deepseek-v4-flash-0731` | $0.44 | $1.32 |
322+| Nemotron 3 120B | `workers-ai/@cf/nvidia/nemotron-3-120b-a12b` | $0.50 | $1.50 |
323+| Kimi K2.6 | `workers-ai/@cf/moonshotai/kimi-k2.6` | $0.95 | $4.00 |
324+| DeepSeek V4 Pro | `workers-ai/@cf/deepseek-ai/deepseek-v4-pro-0813` | $1.32 | $3.96 |
325+| GLM-5.3 | `workers-ai/@cf/zai-org/glm-5.3` | $1.40 | $4.40 |
326+
327+Embeddings, through `POST /openai/v1/embeddings` only:
328+
329+| Model | `model` | Input |
330+| --- | --- | --- |
331+| BGE M3 | `workers-ai/@cf/baai/bge-m3` | $0.012 |
332+| BGE Base (English) | `workers-ai/@cf/baai/bge-base-en-v1.5` | $0.067 |
333+
334+Open models cost much less per call than Claude, and suit one-shot work:
335+titles, summaries, labels, triage, embeddings. In a long loop that sends
336+the same context every turn, Claude's cache reads close most of that gap.
337+
338+### What is refused on g1t's models
339+
340+A request for a model nobody offers is refused with `404` before it
341+reaches a provider, and the error names the models offered. On g1t's
342+models a request is charged only by its tokens, so what a provider bills
343+some other way is refused with `400` for now:
344+
146345 | Not offered on g1t's models yet | In the request |
147346 | --- | --- |
148347 | Fast mode | `speed` other than `standard` |
149348 | Inference in one region | `inference_geo` other than `global` |
150349 | Server-side fallbacks | `fallbacks` |
151−| Server tools, such as web search, web fetch and code execution | A tool whose `type` is not your own (`custom` or none) or a client tool (`bash_…`, `text_editor_…`, `computer_…`, `memory_…`) |
350+| Server tools, such as web search, web fetch and code execution | In Anthropic's format, a tool whose `type` is not your own (`custom` or none) or a client tool (`bash_…`, `text_editor_…`, `computer_…`, `memory_…`). In OpenAI's, a tool that is not a `function`, or `web_search_options`. |
152351 | Containers and skills | `container` |
153352
154−All of them work on the workspace's own key, which the provider bills. In
155−Claude Code on g1t's models, its web search fails for this reason; the
156−rest of Claude Code works.
353+All of them work on your own provider, which bills them. In Claude Code on
354+g1t's models, its web search fails for this reason; the rest of Claude Code
355+works.
157356
158−Anthropic's format is the one served today. OpenAI's format and open
159−models are coming later.
357+## Your own providers
358+
359+Connect a model provider under **Integrations**, and choose which models
360+your own code's gateway requests send to it. Requests there use its key,
361+are counted in the log, and are never charged on g1t. Any provider works:
160362
363+| Provider | Takes, unless you choose |
364+| --- | --- |
365+| An Anthropic key, or an Anthropic-compatible endpoint | `claude-*` |
366+| An OpenAI key, or any other provider | Nothing until you choose |
367+| An OpenAI-compatible endpoint: a self-hosted vLLM or Ollama, LiteLLM, another provider | Nothing until you choose |
368+
369+1. Open the workspace's **Integrations** and choose a provider under
370+ **Model providers**. For your own server, choose **OpenAI-compatible
371+ endpoint** or **Anthropic-compatible endpoint** and give its base URL.
372+2. Paste its key. It is sealed when saved and never shown again: the page,
373+ the API and MCP show only its last four characters.
374+3. Under **AI Gateway models**, list the models to send there, separated by
375+ spaces:
376+
377+ | Write | Takes |
378+ | --- | --- |
379+ | `gpt-5.5` | That model only |
380+ | `gpt-*` | Every model whose id starts with `gpt-` |
381+ | `ollama/*` | Every model named `ollama/…`, sent without the prefix: `ollama/llama3.3` arrives as `llama3.3` |
382+ | `*` | Every model |
383+ | Nothing | No gateway requests |
384+
385+4. Select **Connect**. To change the list or replace the key later, open
386+ **Change its AI Gateway models or key** under the provider.
387+
388+The first provider, in the order they were connected, that takes a model
389+gets its requests. Either format reaches either kind of provider: a Claude
390+key answers OpenAI-format requests, and an OpenAI-compatible endpoint
391+answers Claude Code. Agent runs choose their models under
392+[routing](/guides/models/), apart from this list.
393+
394+From code, connect one with
395+[`POST /workspaces/{workspace}/integrations`](/reference/api/integrations/connect-integration/)
396+and change it with
397+[`PATCH /workspaces/{workspace}/integrations/{id}`](/reference/api/integrations/update-integration/),
398+with `config.gateway_models`, or the `workspace` MCP tool's
399+`connect_integration` and `update_integration` actions. Both need
400+`workspace:admin` and act for an owner. The key is write-only: neither
401+returns it.
402+
403+```sh
404+curl -X PATCH https://api.g1t.sh/workspaces/acme/integrations/con_01kpx5c2d8e4f6g0h2j4k6m8n0 \
405+ -H "Authorization: Bearer $G1T_ADMIN_TOKEN" \
406+ -H "content-type: application/json" \
407+ -d '{ "config": { "base_url": "https://gpu.acme.dev/v1", "gateway_models": ["ollama/*"] }, "secret": "…" }'
408+```
409+
410+`config` replaces the provider's settings whole, so send the ones it has
411+with the change.
412+
161413 ## What it costs
162414
163415 | Where it goes | You pay |
164416 | --- | --- |
165417 | g1t's models | Its tokens at the model's price above, with no markup while the gateway is in beta |
166−| The workspace's own Anthropic key | Nothing on g1t. The provider bills you for the model. |
418+| Your own providers | Nothing on g1t. The provider bills you for the model. |
167419
168420 On g1t's models:
169421
170422 - Each request that used tokens is one line on the statement, under
171423 **AI Gateway**, such as *AI Gateway: Claude Sonnet 5.5, 14,352 tokens,
172− token release-notes*.
424+ token release-notes*. A Claude Haiku 5.5 request over 100,000 prompt
425+ tokens says *long-prompt price*.
173426 - The plan's included usage pays first, then AI credit. Trial credit and
174427 g1t's open-source pool never pay for gateway requests.
175428 - It is not an agent run, so the [agent rate](/guides/usage-and-billing/#the-agent-rate)
178431 like any other usage, and shows on **Usage** under the AI Gateway product.
179432 - A workspace with a 100% discount gets it free through the discount; an
180433 enterprise is invoiced for it after use.
181−
182−### Your own key
183−
184−When the workspace has an Anthropic or Anthropic-compatible model provider
185−under [Integrations](/guides/models/), the gateway sends every request to
186−the first one connected, with its key. Those requests are logged with their
187−tokens and marked **Own key**, and g1t charges nothing for them. Remove the
188−provider and requests go to g1t's models again within a few seconds.
189434
190435 ## Limits and errors
191436
196441 usage left. If auto-reload is on, g1t tries it first.
197442 - It is not on the g1t plan.
198443
199−Errors are Anthropic's shape, so SDKs raise their usual errors:
444+Errors are in the format of the route, so SDKs raise their usual errors.
445+Anthropic's:
200446
201447 ```json
202448 { "type": "error", "error": { "type": "billing_error", "message": "The acme workspace is out of AI credit …" } }
203449 ```
204450
205−| Status | `error.type` | Why |
206−| --- | --- | --- |
207−| `400` | `invalid_request_error` | The body is not JSON, the model is not offered, or the request asks for something [not offered on g1t's models yet](#models). |
208−| `401` | `authentication_error` | The token is unknown, expired or deleted. |
209−| `402` | `billing_error` | Out of AI credit, over the spend limit, or not on the plan. The message says what an owner can do. |
210−| `403` | `permission_error` | Not a workspace's token, or it lacks `models:write`. |
211−| `404` | `not_found_error` | A route the gateway does not answer. |
451+OpenAI's:
212452
213−An error from the model provider, such as `429` or `529`, comes back as
214−the provider sent it. Refused and failed requests are logged with their
215−status and why, and cost nothing.
453+```json
454+{ "error": { "message": "The acme workspace is out of AI credit …", "type": "insufficient_quota", "param": null, "code": "insufficient_quota" } }
455+```
456+
457+| Status | Anthropic's `error.type` | OpenAI's `error.type` (`code`) | Why |
458+| --- | --- | --- | --- |
459+| `400` | `invalid_request_error` | `invalid_request_error` | The body is not JSON, the model is of the wrong kind, the request asks for something [not offered on g1t's models yet](#what-is-refused-on-g1ts-models), or it cannot be said to the model (such as `n` above 1 to Claude). |
460+| `401` | `authentication_error` | `authentication_error` (`invalid_api_key`) | The token is unknown, expired or deleted, or your provider refused its key. |
461+| `402` | `billing_error` | `insufficient_quota` (`insufficient_quota`) | Out of AI credit, over the spend limit, or not on the plan. The message says what an owner can do. |
462+| `403` | `permission_error` | `permission_error` | Not a workspace's token, or it lacks `models:write`. |
463+| `404` | `not_found_error` | `invalid_request_error` (`model_not_found`) | No provider offers the model, or a route the gateway does not answer. |
464+
465+An error from the model provider, such as `429` or `529`, comes back with
466+its status and message, in the route's format. A provider's key never
467+appears in an error or the log, even when the provider quotes it. Refused
468+and failed requests are logged with their status and why, and cost
469+nothing. Every answer carries `x-g1t-request-id`, the request's id in the
470+log.
216471
217−A deleted token, a provider added under Integrations, or AI credit just
218−bought takes effect within about ten seconds.
472+A deleted token, a provider added or changed under Integrations, or AI
473+credit just bought takes effect within about ten seconds.
219474
220475 ## See every request
221476
226481 | Column | |
227482 | --- | --- |
228483 | Time | When it was sent. Hover for the exact time, how long it took and whether it streamed. |
229−| Model | The model it named. On the workspace's own key, the one that answered. |
230−| Input, Output, Cache read, Cache write | Its tokens by kind. |
231−| Cost | What it was charged, before included usage and AI credit paid for it, or **Own key**. |
484+| Model | The model it named on g1t's models, or the one that answered on your own provider; below it, the format it was sent in. |
485+| Served by | g1t's account and the provider (*g1t · Anthropic*, *g1t · Workers AI*), or your provider by name. *None* when it was refused first. |
486+| Input, Output | Its tokens by kind. Hover Input for all of them. |
487+| Cache | Cache reads, then cache writes. Hover for how many writes were to the one-hour cache. |
488+| Cost | What it was charged, before included usage and AI credit paid for it, or **Not charged** on your own provider. |
232489 | Status | The status it was answered with. Hover a refusal or failure for why. |
233490 | Token | The name of the token that sent it. |
234491
237494 From code, list them with
238495 [`GET /workspaces/{workspace}/gateway/requests`](/reference/api/billing/list-gateway-requests/),
239496 or the `billing` MCP tool's
240−[`gateway_requests`](/reference/mcp/#billing) action. Both need
241−`models:read`, which the Read only and Agent presets include.
497+[`gateway_requests`](/reference/mcp/#billing) action. Each request has
498+`format`, `provider`, `connection`, `model` and its tokens, with
499+`cache_write_hour` for one-hour cache writes. Both need `models:read`,
500+which the Read only and Agent presets include.
+1−0
208208 | --- | --- |
209209 | `workspace` `list_integrations` | `GET /workspaces/{workspace}/integrations` |
210210 | `workspace` `connect_integration` | `POST /workspaces/{workspace}/integrations` |
211+| `workspace` `update_integration` | `PATCH /workspaces/{workspace}/integrations/{id}` |
211212 | `workspace` `test_integration` | `POST /workspaces/{workspace}/integrations/{id}/test` |
212213 | `workspace` `disconnect_integration` | `DELETE /workspaces/{workspace}/integrations/{id}` |
213214 | `search` `ticket` | `GET /repos/{owner}/{name}/context?reference=` |
+5−2
1818 kind of work, which provider and model it runs on. Each provider bills
1919 you for the model directly. Open to every workspace now.
2020
21−To call models from your own code with a workspace token, paid from the
22−same AI credit, use the [AI Gateway](/guides/ai-gateway/).
21+To call models from your own code with a workspace token, in Anthropic's
22+or OpenAI's format, paid from the same AI credit or sent to these same
23+providers, use the [AI Gateway](/guides/ai-gateway/). Each provider's
24+**AI Gateway models** say which of those requests go to it; see
25+[your own providers](/guides/ai-gateway/#your-own-providers).
2326
2427 ## Auto
2528
+2−1
548548 | [`revoke_invite`](/reference/api/invites/revoke-workspace-invite/) | Revoke a workspace's pending invite. Owners only. | `workspace`, `id` | `workspace:admin` |
549549 | [`list_integrations`](/reference/api/integrations/list-integrations/) | The workspace's connections. Secrets are never returned. Members only. | `workspace` | `workspace:read` |
550550 | [`connect_integration`](/reference/api/integrations/connect-integration/) | Connect a model provider (Anthropic, OpenAI, Gemini, or a compatible endpoint), Sentry, Datadog, a webhook, Jira or Linear, with `config` and `secret`. Owners only. | `workspace`, `provider` | `workspace:admin` |
551+| [`update_integration`](/reference/api/integrations/update-integration/) | Change one: its `name`, its `config` (replaced whole) or its `secret` (write-only, never returned). Rotates a model provider's key, or sets `config.gateway_models`, the [AI Gateway](/guides/ai-gateway/#your-own-providers) models it takes. Owners only. | `workspace`, `id` | `workspace:admin` |
551552 | [`disconnect_integration`](/reference/api/integrations/disconnect-integration/) | Remove it and its secrets. Owners only. | `workspace`, `id` | `workspace:admin` |
552553 | [`test_integration`](/reference/api/integrations/test-integration/) | Check its credentials against the system it connects to. Owners only. | `workspace`, `id` | `workspace:admin` |
553554 | [`get_model_routes`](/reference/api/integrations/get-model-routes/) | Which provider and model each kind of work goes to. Members only. | `workspace` | `workspace:read` |
583584 | [`buy_ai_credit`](/reference/api/billing/buy-ai-credit/) | A payment page (`url`) to buy `amount_cents` of credit, in whole dollars from $10 to $1,000, for a person to open and pay; it returns to the workspace's billing page. Owners, as people. | `workspace`, `amount_cents` | `billing:write` |
584585 | [`invoices`](/reference/api/billing/list-invoices/) | Every invoice (`invoices`, in cents), g1t's itemised usage invoices (`usage_invoices`), and what the next one comes to so far (`upcoming`). | `workspace` | `billing:read` |
585586 | [`billing_details`](/reference/api/billing/get-billing-details/) | Who invoices are made out to, and the payment method on file as far as it is safe to show. | `workspace` | `billing:read` |
586−| [`gateway_requests`](/reference/api/billing/list-gateway-requests/) | The workspace's recent [AI Gateway](/guides/ai-gateway/) requests, newest first: model, tokens by kind, `cost_micros`, `charged_micros`, `status`, `own_key` and the token that sent each. `limit` (50, at most 200) and `before` (the last page's `next`) page through them. Kept 30 days. | `workspace` | `models:read` |
587+| [`gateway_requests`](/reference/api/billing/list-gateway-requests/) | The workspace's recent [AI Gateway](/guides/ai-gateway/) requests, newest first: model, `format` (`anthropic` or `openai`), `provider` and `connection` (who served it), tokens by kind with `cache_write_hour`, `cost_micros`, `charged_micros`, `status`, `own_key` and the token that sent each. `limit` (50, at most 200) and `before` (the last page's `next`) page through them. Kept 30 days. | `workspace` | `models:read` |
587588
588589 ## `notifications`
589590
+0−0

Binary or large file; its contents are not shown.

+31−9
77
88 import type { GatewayRequest } from "@g1t/contracts";
99
10−import { duration, shortCount, statusTone, tokenKinds } from "../lib/gateway";
10+import { cacheKinds, duration, formatLabel, servedBy, shortCount, statusTone, tokenKinds } from "../lib/gateway";
1111 import { money } from "../lib/usage";
1212 import { TimeAgo } from "./ui";
1313 import { Badge } from "./ui/badge";
1818 const NUM = "px-3 py-2 text-right tabular-nums whitespace-nowrap";
1919
2020 /** The columns, so the skeleton has as many as the table. */
21−const COLUMNS = ["Time", "Model", "Input", "Output", "Cache read", "Cache write", "Cost", "Status", "Token"] as const;
21+const COLUMNS = ["Time", "Model", "Served by", "Input", "Output", "Cache", "Cost", "Status", "Token"] as const;
2222
2323 /** One request's row. */
2424 function Row({ request }: { request: GatewayRequest }) {
2525 const status = statusTone(request.status);
26+ const served = servedBy(request);
2627 return (
2728 <tr className="text-sm">
2829 <td className="px-3 py-2 whitespace-nowrap text-muted">
3233 </span>
3334 </Hint>
3435 </td>
35− <td className="px-3 py-2 font-mono text-xs whitespace-nowrap">{request.model}</td>
36+ <td className="max-w-[15rem] px-3 py-2 whitespace-nowrap">
37+ <Hint label={request.model}>
38+ <span tabIndex={0} className="block truncate font-mono text-xs">
39+ {request.model}
40+ </span>
41+ </Hint>
42+ <span className="block text-[0.6875rem] text-faint">{formatLabel(request.format)} format</span>
43+ </td>
44+ <td className="max-w-[12rem] truncate px-3 py-2 whitespace-nowrap">
45+ <Hint label={served.hint}>
46+ <span tabIndex={0} className={request.ownKey ? "" : "text-muted"}>
47+ {served.label}
48+ </span>
49+ </Hint>
50+ </td>
3651 <td className={NUM}>
3752 <Hint label={tokenKinds(request)}>
3853 <span tabIndex={0}>{shortCount(request.input)}</span>
3954 </Hint>
4055 </td>
4156 <td className={NUM}>{shortCount(request.output)}</td>
42− <td className={NUM}>{shortCount(request.cacheRead)}</td>
43− <td className={NUM}>{shortCount(request.cacheWrite)}</td>
4457 <td className={NUM}>
58+ <Hint label={cacheKinds(request)}>
59+ <span tabIndex={0}>
60+ {shortCount(request.cacheRead)}
61+ <span className="text-faint"> / </span>
62+ {shortCount(request.cacheWrite)}
63+ </span>
64+ </Hint>
65+ </td>
66+ <td className={NUM}>
4567 {request.ownKey ? (
46− <Hint label="Sent with the workspace's own provider key: counted, not charged.">
68+ <Hint label="Sent to the workspace's own provider: counted, not charged.">
4769 <span tabIndex={0}>
48− <Badge>Own key</Badge>
70+ <Badge>Not charged</Badge>
4971 </span>
5072 </Hint>
5173 ) : (
6991 }
7092
7193 /** Whether a column holds a number, read right-aligned. */
72−const numeric = (index: number) => index >= 2 && index <= 6;
94+const numeric = (index: number) => index >= 3 && index <= 6;
7395
7496 /** The table's frame: its heading row, around its body. */
7597 function Frame({ children }: { children: ReactNode }) {
114136 <tr key={row} className="text-sm">
115137 {COLUMNS.map((column, index) => (
116138 <td key={column} className={numeric(index) ? NUM : "px-3 py-2"}>
117− <SkeletonLine className={numeric(index) ? "ml-auto w-10" : index === 1 ? "w-36" : "w-16"} />
139+ <SkeletonLine className={numeric(index) ? "ml-auto w-10" : index === 1 ? "w-36" : index === 2 ? "w-24" : "w-16"} />
118140 </td>
119141 ))}
120142 </tr>
+50−1
77 import { type ReactNode, useState } from "react";
88 import { Form, Link } from "react-router";
99
10−import { type Connection, MODEL_TASKS, MODEL_TIERS, type ModelRoute, type ModelTask, type ModelTier, type Provider, PROVIDERS } from "@g1t/contracts";
10+import { type Connection, MODEL_TASKS, MODEL_TIERS, type ModelRoute, type ModelTask, type ModelTier, type Provider, PROVIDERS, gatewayPatterns } from "@g1t/contracts";
1111
1212 import { Field, Input, SubmitButton } from "./ui";
1313 import { Hint } from "./ui/hint";
259259 </Field>
260260 );
261261
262+/** What to suggest a provider's AI Gateway models be, as typed. */
263+const GATEWAY_EXAMPLES: Partial<Record<Provider, string>> = {
264+ openai: "gpt-*",
265+ gemini: "gemini-*",
266+ xai: "grok-*",
267+ mistral: "mistral-* codestral-*",
268+ deepseek: "deepseek-*",
269+ openrouter: "openrouter/*",
270+ groq: "groq/*",
271+ together: "together/*",
272+ fireworks: "fireworks/*",
273+ cerebras: "cerebras/*",
274+ openai_endpoint: "ollama/*",
275+ anthropic_endpoint: "claude-*",
276+ azure_openai: "azure/*",
277+};
278+
279+/**
280+ * Which of the AI Gateway's requests go to a model provider, by the model
281+ * they name: for connecting one, and for changing it after.
282+ */
283+export function GatewayModelsField({ provider, value }: { provider: Provider; value?: string[] }) {
284+ const defaults = gatewayPatterns(provider, {});
285+ return (
286+ <Field
287+ label="AI Gateway models"
288+ hint="Your own code's requests to the AI Gateway that name one of these models come here, counted and never charged by g1t. Model ids, or prefixes ending in *; a prefix ending in /* is taken off before sending, so ollama/llama3.3 arrives as llama3.3. Separate them with spaces. Empty: none."
289+ >
290+ <Input
291+ name="gatewayModels"
292+ defaultValue={(value ?? defaults).join(" ")}
293+ placeholder={GATEWAY_EXAMPLES[provider] ?? "model-id or prefix-*"}
294+ className="font-mono text-[0.8125rem]"
295+ autoComplete="off"
296+ spellCheck={false}
297+ />
298+ </Field>
299+ );
300+}
301+
262302 /** What connecting a model provider asks for. */
263303 export function ModelProviderFields({ provider }: { provider: Provider }): ReactNode {
304+ return (
305+ <>
306+ <ModelProviderKeyFields provider={provider} />
307+ <GatewayModelsField provider={provider} />
308+ </>
309+ );
310+}
311+
312+function ModelProviderKeyFields({ provider }: { provider: Provider }): ReactNode {
264313 const entry = MODEL_CATALOG[provider];
265314 if (provider === "azure_openai") {
266315 return (
+22−1
11 import assert from "node:assert/strict";
22 import { test } from "node:test";
33
4−import { duration, shortCount, statusTone, tokenKinds, totalTokens } from "./gateway.ts";
4+import { cacheKinds, duration, formatLabel, parseGatewayModels, servedBy, shortCount, statusTone, tokenKinds, totalTokens } from "./gateway.ts";
55
66 test("token counts read short", () => {
77 assert.equal(shortCount(812), "812");
2929 assert.equal(duration(4_210), "4.2 s");
3030 assert.equal(duration(125_000), "2 min 5 s");
3131 });
32+
33+test("a request says which format it came in and who served it", () => {
34+ assert.equal(formatLabel("openai"), "OpenAI");
35+ assert.equal(formatLabel("anthropic"), "Anthropic");
36+ assert.deepEqual(servedBy({ ownKey: false, provider: "workers-ai", connection: null }).label, "g1t · Workers AI");
37+ assert.deepEqual(servedBy({ ownKey: false, provider: "anthropic", connection: null }).label, "g1t · Anthropic");
38+ const own = servedBy({ ownKey: true, provider: "openai_endpoint", connection: "Office GPU" });
39+ assert.equal(own.label, "Office GPU");
40+ assert.match(own.hint, /OpenAI-compatible endpoint: counted, not charged/);
41+ assert.equal(servedBy({ ownKey: false, provider: "", connection: null }).label, "None");
42+});
43+
44+test("a connection's AI Gateway models are typed as a list", () => {
45+ assert.deepEqual(parseGatewayModels(" gpt-*, ollama/* claude-haiku-5-5 "), ["gpt-*", "ollama/*", "claude-haiku-5-5"]);
46+ assert.deepEqual(parseGatewayModels(""), []);
47+});
48+
49+test("cache tokens read in words, hour-long writes named", () => {
50+ assert.equal(cacheKinds({ cacheRead: 12_000, cacheWrite: 2_400, cacheWriteHour: 0 }), "12,000 read from the cache, 2,400 written");
51+ assert.equal(cacheKinds({ cacheRead: 0, cacheWrite: 2_400, cacheWriteHour: 2_400 }), "0 read from the cache, 2,400 written, 2,400 of them to the hour-long cache");
52+});
+49−0
1010 /** The base URL an Anthropic SDK or Claude Code is pointed at. */
1111 export const GATEWAY_BASE_URL = "https://models.g1t.sh/anthropic";
1212
13+/** The base URL an OpenAI SDK, or any tool that speaks OpenAI's API, is pointed at. */
14+export const GATEWAY_OPENAI_BASE_URL = "https://models.g1t.sh/openai/v1";
15+
16+/** How a request's format reads. */
17+export function formatLabel(format: GatewayRequest["format"] | undefined): string {
18+ return format === "openai" ? "OpenAI" : "Anthropic";
19+}
20+
21+/** Who served a request, as people read it, and a longer line for its hint. */
22+export function servedBy(request: Pick<GatewayRequest, "ownKey" | "provider" | "connection">): { label: string; hint: string } {
23+ if (request.ownKey) {
24+ const name = request.connection || "Your provider";
25+ return { label: name, hint: `${name}, the workspace's own ${providerName(request.provider)}: counted, not charged.` };
26+ }
27+ if (!request.provider) return { label: "None", hint: "Refused before it reached a model." };
28+ return { label: `g1t · ${providerName(request.provider)}`, hint: `g1t's account at ${providerName(request.provider)}, charged at the model's price.` };
29+}
30+
31+const PROVIDER_NAMES: Record<string, string> = {
32+ anthropic: "Anthropic",
33+ "workers-ai": "Workers AI",
34+ openai: "OpenAI",
35+ openai_endpoint: "OpenAI-compatible endpoint",
36+ anthropic_endpoint: "Anthropic-compatible endpoint",
37+ azure_openai: "Azure OpenAI",
38+ gemini: "Google Gemini",
39+ openrouter: "OpenRouter",
40+};
41+
42+/** A provider's name, as people know it. */
43+export function providerName(provider: string): string {
44+ return PROVIDER_NAMES[provider] ?? (provider || "provider");
45+}
46+
47+/** A connection's AI Gateway models, as a field shows them, from what was typed. */
48+export function parseGatewayModels(text: string): string[] {
49+ return text
50+ .split(/[\s,]+/)
51+ .map((model) => model.trim())
52+ .filter(Boolean);
53+}
54+
1355 /** A count of tokens, short: `812`, `12.4K`, `3.1M`. */
1456 export function shortCount(n: number): string {
1557 if (n < 1_000) return String(n);
3274 return `${n(request.input)} input, ${n(request.output)} output, ${n(request.cacheRead)} cache read, ${n(request.cacheWrite)} cache write`;
3375 }
3476
77+/** A request's cache tokens in words, for a hint: read, written, and how many of the writes last an hour. */
78+export function cacheKinds(request: Pick<GatewayRequest, "cacheRead" | "cacheWrite" | "cacheWriteHour">): string {
79+ const n = (v: number) => v.toLocaleString("en-US");
80+ const hour = request.cacheWriteHour ? `, ${n(request.cacheWriteHour)} of them to the hour-long cache` : "";
81+ return `${n(request.cacheRead)} read from the cache, ${n(request.cacheWrite)} written${hour}`;
82+}
83+
3584 /** How a status reads, and how much it matters. */
3685 export function statusTone(status: number): { label: string; tone: "success" | "warn" | "danger" | "neutral" } {
3786 if (status >= 200 && status < 300) return { label: String(status), tone: "success" };
+1−1
536536 <p className="font-medium">
537537 AI Gateway <span className="ml-1 rounded bg-accent/15 px-1.5 py-0.5 text-xs text-accent">Free during beta</span>
538538 </p>
539− <p className="text-xs text-faint">Your own code calling Claude models through g1t with a workspace token, paid from AI credit; with your own Anthropic key, free</p>
539+ <p className="text-xs text-faint">Your own code calling Claude and open models through g1t, in Anthropic's or OpenAI's format, paid from AI credit; on your own providers, free</p>
540540 </td>
541541 <td className="px-4 py-3 text-muted">What the provider charges</td>
542542 <td className="hidden px-4 py-3 tabular-nums sm:table-cell">{gateway?.markupPercent ?? 0}%</td>
+25−9
1−import { BookOpen, KeyRound } from "lucide-react";
1+import { BookOpen, KeyRound, Plug } from "lucide-react";
22 import { Suspense } from "react";
33 import { Await, Link, data } from "react-router";
44
55 import type { Route } from "./+types/gateway";
66 import { GatewaySkeleton, GatewayTable } from "../../components/gateway";
77 import { ButtonLink, CopyLine, EmptyState } from "../../components/ui";
8−import { GATEWAY_BASE_URL, GATEWAY_DOCS } from "../../lib/gateway";
8+import { GATEWAY_BASE_URL, GATEWAY_DOCS, GATEWAY_OPENAI_BASE_URL } from "../../lib/gateway";
99 import { page } from "../../lib/meta";
1010 import { billing } from "../../lib/services.server";
1111 import { getViewer, roleIn } from "../../lib/session.server";
3737 <div className="space-y-6">
3838 <section className="space-y-3">
3939 <p className="max-w-3xl text-sm text-muted">
40− Send your own code's model requests in Anthropic's Messages format to this base URL, with one of the workspace's
41− access tokens that has the <code className="font-mono text-xs">models:write</code> scope as the API key. On g1t's
42− models each request is charged at the model's price and paid from AI credit; with the workspace's own Anthropic key
43− under Integrations it is only counted. Prompts and answers are never kept.
40+ Send your own code's model requests to one of these base URLs, in Anthropic's or OpenAI's format, with one of the
41+ workspace's access tokens that has the <code className="font-mono text-xs">models:write</code> scope as the API key.
42+ Any model works in either format: Claude and open models on g1t's account are charged at the model's price and paid
43+ from AI credit; models on the workspace's own providers under Integrations are only counted. Prompts and answers are
44+ never kept.
4445 </p>
45− <div className="max-w-xl">
46− <CopyLine text={GATEWAY_BASE_URL} />
47− </div>
46+ <dl className="grid max-w-3xl gap-3 md:grid-cols-2">
47+ <div className="min-w-0">
48+ <dt className="mb-1.5 text-xs text-faint">Anthropic format</dt>
49+ <dd>
50+ <CopyLine text={GATEWAY_BASE_URL} />
51+ </dd>
52+ </div>
53+ <div className="min-w-0">
54+ <dt className="mb-1.5 text-xs text-faint">OpenAI format</dt>
55+ <dd>
56+ <CopyLine text={GATEWAY_OPENAI_BASE_URL} />
57+ </dd>
58+ </div>
59+ </dl>
4860 <div className="flex flex-wrap gap-2">
4961 {owner && (
5062 <ButtonLink variant="quiet" to={`/${slug}/-/tokens`}>
5264 Access tokens
5365 </ButtonLink>
5466 )}
67+ <ButtonLink variant="quiet" to={`/${slug}/-/integrations`}>
68+ <Plug size={14} />
69+ Your own providers
70+ </ButtonLink>
5571 <ButtonLink variant="quiet" to={GATEWAY_DOCS} reloadDocument>
5672 <BookOpen size={14} />
5773 How to use it
+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

+0−0

Binary or large file; its contents are not shown.

This change is too large to show in full.