AI Gateway: OpenAI's format, open models, and your own providers
models.g1t.sh/openai/v1 answers POST /chat/completions (streamed or not, function tools, response_format, reasoning_effort, cached-token usage), POST /embeddings and GET /models, on the same tokens, scope, admission, charging and log as /anthropic. Any model works in either format: the proxy translates Anthropic <-> OpenAI both ways, streaming and tool calls included, carrying Claude's thinking blocks in the first tool call's id. - Model ids: anthropic/claude-*, workers-ai/@cf/..., bare ids still work. - Open models on g1t's account through Workers AI via the AI Gateway (WORKERS_AI_TOKEN), priced in gateway_models (billing migration 0046). - 0046 also fixes Sonnet 5.5 cache reads ($0.10), prices one-hour cache writes, and adds Claude Haiku 5.5 with long-prompt pricing (threshold and over_ prices); charging splits 5m/1h cache writes. - Your own providers: any model connection (Anthropic or OpenAI key, any compatible endpoint) takes the models in config.gateway_models (ids, prefix*, ns/* stripped); counted, never charged. New update_integration (PATCH /workspaces/{ws}/integrations/{id}) rotates keys write-only. Provider errors are scrubbed of keys. - Log, REST and MCP show format, provider, connection, cache_write_hour. - Web: both base URLs and Served by on the AI Gateway page; AI Gateway models and key replacement on Integrations. - Docs: AI Gateway guide (both formats, models and prices, own providers, errors), MCP and integrations reference, OpenAPI, llms.txt, pricing, billing operations.
| 359 | 359 | &[ | |
| 360 | 360 | Op::ListIntegrations, | |
| 361 | 361 | Op::ConnectIntegration, | |
| 362 | + | Op::UpdateIntegration, | |
| 362 | 363 | Op::DisconnectIntegration, | |
| 363 | 364 | Op::TestIntegration, | |
| 364 | 365 | Op::GetModelRoutes, | |
| ⋯ | |||
| 461 | 462 | Op::ListEvents => "List repository events", | |
| 462 | 463 | Op::ListIntegrations => "List integrations", | |
| 463 | 464 | Op::ConnectIntegration => "Connect an integration", | |
| 465 | + | Op::UpdateIntegration => "Update an integration", | |
| 464 | 466 | Op::DisconnectIntegration => "Disconnect an integration", | |
| 465 | 467 | Op::TestIntegration => "Test an integration", | |
| 466 | 468 | Op::GetContext => "Look up a ticket", | |
| 169 | 169 | ListEvents, | |
| 170 | 170 | ListIntegrations, | |
| 171 | 171 | ConnectIntegration, | |
| 172 | + | UpdateIntegration, | |
| 172 | 173 | DisconnectIntegration, | |
| 173 | 174 | TestIntegration, | |
| 174 | 175 | GetContext, | |
| ⋯ | |||
| 629 | 630 | } | |
| 630 | 631 | ||
| 631 | 632 | impl Op { | |
| 632 | − | pub const ALL: [Op; 219] = [ | |
| 633 | + | pub const ALL: [Op; 220] = [ | |
| 633 | 634 | Op::Whoami, | |
| 634 | 635 | Op::GetWorkspace, | |
| 635 | 636 | Op::CreateWorkspace, | |
| ⋯ | |||
| 711 | 712 | Op::ListEvents, | |
| 712 | 713 | Op::ListIntegrations, | |
| 713 | 714 | Op::ConnectIntegration, | |
| 715 | + | Op::UpdateIntegration, | |
| 714 | 716 | Op::DisconnectIntegration, | |
| 715 | 717 | Op::TestIntegration, | |
| 716 | 718 | Op::GetContext, | |
| ⋯ | |||
| 939 | 941 | Op::ListEvents => "list_events", | |
| 940 | 942 | Op::ListIntegrations => "list_integrations", | |
| 941 | 943 | Op::ConnectIntegration => "connect_integration", | |
| 944 | + | Op::UpdateIntegration => "update_integration", | |
| 942 | 945 | Op::DisconnectIntegration => "disconnect_integration", | |
| 943 | 946 | Op::TestIntegration => "test_integration", | |
| 944 | 947 | Op::GetContext => "get_context", | |
| ⋯ | |||
| 1267 | 1270 | "A workspace's integrations: its own model provider, the alert sources that open issues (Sentry, Datadog, webhooks), and the trackers whose tickets agents can read (Jira, Linear). Secrets are never returned. Members only." | |
| 1268 | 1271 | } | |
| 1269 | 1272 | Op::ConnectIntegration => { | |
| 1270 | − | "Connect a workspace to an outside system. provider is a model provider (anthropic, openai, gemini, xai, mistral, deepseek, azure_openai, openrouter, groq, together, fireworks, cerebras, anthropic_endpoint or openai_endpoint: your own key, billed by that provider, and free on g1t while it is being built out; a workspace can connect several and route each kind of work with set_model_routes), or sentry, datadog, webhook, jira or linear. config holds the settings each needs; secret is the API key or token. For datadog and webhook, g1t makes the signing secret and returns it once. Owners only." | |
| 1273 | + | "Connect a workspace to an outside system. provider is a model provider (anthropic, openai, gemini, xai, mistral, deepseek, azure_openai, openrouter, groq, together, fireworks, cerebras, anthropic_endpoint or openai_endpoint: your own key, billed by that provider, and free on g1t while it is being built out; a workspace can connect several and route each kind of work with set_model_routes), or sentry, datadog, webhook, jira or linear. config holds the settings each needs; secret is the API key or token, kept encrypted and never returned (secret_hint shows its last four characters). For a model provider, config.gateway_models chooses which AI Gateway requests go to it by the model they name: ids such as gpt-5.5, or prefixes ending in * such as gpt-* or ollama/* (a /* prefix is taken off before sending); absent, an Anthropic key or Anthropic-compatible endpoint takes claude-* and the others take nothing. Requests on the workspace's own provider are counted and never charged. For datadog and webhook, g1t makes the signing secret and returns it once. Owners only." | |
| 1274 | + | } | |
| 1275 | + | Op::UpdateIntegration => { | |
| 1276 | + | "Change an integration: its name, its config (replaced whole when given) or its secret (a new key replaces the old one, write-only). Use it to rotate a model provider's key or to choose its config.gateway_models, the AI Gateway models it takes. Fields left out are kept. Secrets are never returned. Owners only." | |
| 1271 | 1277 | } | |
| 1272 | 1278 | Op::DisconnectIntegration => { | |
| 1273 | 1279 | "Remove an integration and its secrets. Agents already running on a model provider being removed stop reaching it. Owners only." | |
| ⋯ | |||
| 1517 | 1523 | "Who a workspace's invoices are made out to: the billing `email`, `name`, `address`, tax ID (`tax_id_type`, `tax_id`), `po_number` and the invoices' `language`, with the default `payment_method` as far as it is safe to show (its kind, brand, last four digits and expiry). `customer` is false until the workspace has been set up to pay. Tax is worked out from the address: `tax_location` says whether it is enough for that (a country, and in the US a ZIP code), `tax_address_needed_at` is set while g1t is holding a charge for want of one, `tax_id_status` is Stripe's check of the tax ID (`pending`, `verified`, `unverified` or `unavailable`), and `tax_exempt` is `none`, `exempt` or `reverse`. Members of the workspace only." | |
| 1518 | 1524 | } | |
| 1519 | 1525 | Op::ListGatewayRequests => { | |
| 1520 | − | "A workspace's recent AI Gateway requests, newest first: each with its `id`, `created_at`, `model`, the access token that sent it (`token_id`, `token_name`), its tokens by kind (`input`, `output`, `cache_read`, `cache_write`), what they cost at the model's price (`cost_micros`) and what the workspace was charged (`charged_micros`, before included usage and AI credit paid for it; 0 on the workspace's own provider key, `own_key`), the HTTP `status` it was answered with, whether it was `streamed`, `duration_ms`, and `error` for one that was refused or failed. Prompts and answers are never kept. `limit` is how many, 50 unless given and 200 at most; pass `next` from one page as `before` for the next. Requests are kept `retention_days` (30). Members of the workspace only." | |
| 1526 | + | "A workspace's recent AI Gateway requests, newest first: each with its `id`, `created_at`, `model`, the access token that sent it (`token_id`, `token_name`), its tokens by kind (`input`, `output`, `cache_read`, `cache_write`, and of those writes `cache_write_hour` to the hour-long cache), the `format` it was sent in (`anthropic` or `openai`), who served it (`provider`: `anthropic` or `workers-ai` on g1t's account, the connection's provider on the workspace's own, and `connection`, that connection's name), what they cost at the model's price (`cost_micros`) and what the workspace was charged (`charged_micros`, before included usage and AI credit paid for it; 0 on the workspace's own provider key, `own_key`), the HTTP `status` it was answered with, whether it was `streamed`, `duration_ms`, and `error` for one that was refused or failed. Prompts and answers are never kept. `limit` is how many, 50 unless given and 200 at most; pass `next` from one page as `before` for the next. Requests are kept `retention_days` (30). Members of the workspace only." | |
| 1521 | 1527 | } | |
| 1522 | 1528 | Op::ListUserTeams => { | |
| 1523 | 1529 | "The teams someone is in within a workspace, as list_teams describes them, leaving out secret teams you cannot see. Members of the workspace only." | |
| ⋯ | |||
| 2260 | 2266 | "name": { "type": "string", "description": "What to call it. The provider's name if left out." }, | |
| 2261 | 2267 | "config": { | |
| 2262 | 2268 | "type": "object", | |
| 2263 | − | "description": "Settings. repo (owner/name) is where alerts open issues; assign puts an agent on each; label names the label (bug). organization is the Sentry org's slug. site is Jira's address; email the account its token belongs to; keys the project or team keys it answers for. base_url and auth_header (x-api-key or authorization) are for your own endpoint; model overrides the model for every kind of work. write_back (default true) tells the outside system when the work lands.", | |
| 2269 | + | "description": "Settings. repo (owner/name) is where alerts open issues; assign puts an agent on each; label names the label (bug). organization is the Sentry org's slug. site is Jira's address; email the account its token belongs to; keys the project or team keys it answers for. base_url and auth_header (x-api-key or authorization) are for your own endpoint; model overrides the model for every kind of work; gateway_models (model ids, or prefixes ending in * such as gpt-* or ollama/*) chooses which AI Gateway requests go to a model provider. write_back (default true) tells the outside system when the work lands.", | |
| 2264 | 2270 | }, | |
| 2265 | − | "secret": { "type": "string", "description": "The API key or token g1t uses to call it." }, | |
| 2271 | + | "secret": { "type": "string", "description": "The API key or token g1t uses to call it. Write-only: kept encrypted, never returned." }, | |
| 2266 | 2272 | "signing_secret": { "type": "string", "description": "For sentry: the integration's client secret." }, | |
| 2267 | 2273 | }), | |
| 2268 | 2274 | &["workspace", "provider"], | |
| 2269 | 2275 | ), | |
| 2276 | + | Op::UpdateIntegration => object( | |
| 2277 | + | json!({ | |
| 2278 | + | "workspace": workspace_schema(), | |
| 2279 | + | "id": { "type": "string", "description": "The integration's id." }, | |
| 2280 | + | "name": { "type": "string", "description": "A new name." }, | |
| 2281 | + | "config": { | |
| 2282 | + | "type": "object", | |
| 2283 | + | "description": "Its settings, replaced whole: the same fields as connect_integration's config. For a model provider, gateway_models chooses the AI Gateway models it takes.", | |
| 2284 | + | }, | |
| 2285 | + | "secret": { "type": "string", "description": "A new API key or token, replacing the old one. Write-only: kept encrypted, never returned." }, | |
| 2286 | + | "signing_secret": { "type": "string", "description": "For sentry: a new client secret." }, | |
| 2287 | + | }), | |
| 2288 | + | &["workspace", "id"], | |
| 2289 | + | ), | |
| 2270 | 2290 | Op::GetModelRoutes => object(json!({ "workspace": workspace_schema() }), &["workspace"]), | |
| 2271 | 2291 | Op::ListWebhooks => object(hook_owner(json!({})), &[]), | |
| 2272 | 2292 | Op::ListWorkflows => repo_only(), | |
| ⋯ | |||
| 2910 | 2930 | | Op::CreateRepo | |
| 2911 | 2931 | | Op::ListIntegrations | |
| 2912 | 2932 | | Op::ConnectIntegration | |
| 2933 | + | | Op::UpdateIntegration | |
| 2913 | 2934 | | Op::DisconnectIntegration | |
| 2914 | 2935 | | Op::TestIntegration | |
| 2915 | 2936 | | Op::GetModelRoutes | |
| ⋯ | |||
| 4094 | 4115 | ) | |
| 4095 | 4116 | .await | |
| 4096 | 4117 | } | |
| 4118 | + | Op::UpdateIntegration => { | |
| 4119 | + | let config = match &input["config"] { | |
| 4120 | + | Value::Null => Value::Null, | |
| 4121 | + | config => camel_keys(config), | |
| 4122 | + | }; | |
| 4123 | + | pass( | |
| 4124 | + | integrations, | |
| 4125 | + | "update", | |
| 4126 | + | &json!({ | |
| 4127 | + | "actor": actor(), | |
| 4128 | + | "workspace": workspace(), | |
| 4129 | + | "id": text(input, "id"), | |
| 4130 | + | "name": optional_text(input, "name"), | |
| 4131 | + | "config": config, | |
| 4132 | + | "secret": optional_text(input, "secret"), | |
| 4133 | + | "signingSecret": optional_text(input, "signing_secret"), | |
| 4134 | + | }), | |
| 4135 | + | ) | |
| 4136 | + | .await | |
| 4137 | + | } | |
| 4097 | 4138 | Op::DisconnectIntegration | Op::TestIntegration => { | |
| 4098 | 4139 | pass( | |
| 4099 | 4140 | integrations, | |
| 3713 | 3713 | }, | |
| 3714 | 3714 | "notes": "`signing_secret` is set only when g1t made it, for `datadog` and `webhook`, and is shown only this once. A model provider is tested as it is connected. See [integrations](/guides/integrations/) and [model providers](/guides/models/)." | |
| 3715 | 3715 | }, | |
| 3716 | + | "update_integration": { | |
| 3717 | + | "params": { | |
| 3718 | + | "workspace": "flagon-io", | |
| 3719 | + | "id": "con_01kpx5c2d8e4f6g0h2j4k6m8n0" | |
| 3720 | + | }, | |
| 3721 | + | "request": { | |
| 3722 | + | "config": { | |
| 3723 | + | "base_url": "https://gpu.flagon.dev/v1", | |
| 3724 | + | "gateway_models": ["ollama/*"] | |
| 3725 | + | }, | |
| 3726 | + | "secret": "sk-office-gpu-key" | |
| 3727 | + | }, | |
| 3728 | + | "response": { | |
| 3729 | + | "id": "con_01kpx5c2d8e4f6g0h2j4k6m8n0", | |
| 3730 | + | "workspace": "flagon-io", | |
| 3731 | + | "provider": "openai_endpoint", | |
| 3732 | + | "kind": "models", | |
| 3733 | + | "name": "Office GPU", | |
| 3734 | + | "config": { | |
| 3735 | + | "write_back": true, | |
| 3736 | + | "base_url": "https://gpu.flagon.dev/v1", | |
| 3737 | + | "gateway_models": ["ollama/*"] | |
| 3738 | + | }, | |
| 3739 | + | "secret_hint": "…-key", | |
| 3740 | + | "webhook_url": null, | |
| 3741 | + | "created_by": "syntaqx", | |
| 3742 | + | "created_at": "2026-10-07T12:10:44.512Z", | |
| 3743 | + | "last_used_at": "2026-10-07T14:00:02.000Z", | |
| 3744 | + | "last_error": null, | |
| 3745 | + | "models": ["llama3.3:70b", "qwen3-coder:30b"] | |
| 3746 | + | }, | |
| 3747 | + | "notes": "The secret is write-only: it is kept encrypted and only its last four characters come back, as `secret_hint`. With `gateway_models` set to `ollama/*`, an AI Gateway request for `ollama/qwen3-coder:30b` reaches this endpoint as `qwen3-coder:30b`, on the workspace's own account. See [the AI Gateway guide](/guides/ai-gateway/#your-own-providers)." | |
| 3748 | + | }, | |
| 3716 | 3749 | "disconnect_integration": { | |
| 3717 | 3750 | "params": { | |
| 3718 | 3751 | "workspace": "flagon-io", | |
| ⋯ | |||
| 5966 | 5999 | "workspace": "flagon-io" | |
| 5967 | 6000 | }, | |
| 5968 | 6001 | "query": { | |
| 5969 | − | "limit": 2 | |
| 6002 | + | "limit": 3 | |
| 5970 | 6003 | }, | |
| 5971 | 6004 | "response": { | |
| 5972 | 6005 | "requests": [ | |
| ⋯ | |||
| 5980 | 6013 | "output": 512, | |
| 5981 | 6014 | "cache_read": 12000, | |
| 5982 | 6015 | "cache_write": 0, | |
| 5983 | − | "cost_micros": 11200, | |
| 5984 | − | "charged_micros": 11200, | |
| 6016 | + | "cache_write_hour": 0, | |
| 6017 | + | "cost_micros": 10000, | |
| 6018 | + | "charged_micros": 10000, | |
| 5985 | 6019 | "status": 200, | |
| 5986 | 6020 | "own_key": false, | |
| 6021 | + | "format": "anthropic", | |
| 6022 | + | "provider": "anthropic", | |
| 6023 | + | "connection": null, | |
| 5987 | 6024 | "streamed": true, | |
| 5988 | 6025 | "duration_ms": 4210, | |
| 5989 | 6026 | "error": null | |
| 5990 | 6027 | }, | |
| 5991 | 6028 | { | |
| 6029 | + | "id": "gw_9b3d1f7a5c2e4b6d8f0a1c33", | |
| 6030 | + | "created_at": "2026-10-07T14:00:02.000Z", | |
| 6031 | + | "model": "qwen3-coder:30b", | |
| 6032 | + | "token_id": "tok_01kkr5a2b8c4d6e0f3g5h7j9k1", | |
| 6033 | + | "token_name": "triage-bot", | |
| 6034 | + | "input": 2210, | |
| 6035 | + | "output": 96, | |
| 6036 | + | "cache_read": 0, | |
| 6037 | + | "cache_write": 0, | |
| 6038 | + | "cache_write_hour": 0, | |
| 6039 | + | "cost_micros": 0, | |
| 6040 | + | "charged_micros": 0, | |
| 6041 | + | "status": 200, | |
| 6042 | + | "own_key": true, | |
| 6043 | + | "format": "openai", | |
| 6044 | + | "provider": "openai_endpoint", | |
| 6045 | + | "connection": "Office GPU", | |
| 6046 | + | "streamed": false, | |
| 6047 | + | "duration_ms": 1830, | |
| 6048 | + | "error": null | |
| 6049 | + | }, | |
| 6050 | + | { | |
| 5992 | 6051 | "id": "gw_1a7e3c5b9d2f4e6a8c0b2d41", | |
| 5993 | 6052 | "created_at": "2026-10-07T13:58:40.000Z", | |
| 5994 | 6053 | "model": "claude-opus-5-5", | |
| ⋯ | |||
| 5998 | 6057 | "output": 0, | |
| 5999 | 6058 | "cache_read": 0, | |
| 6000 | 6059 | "cache_write": 0, | |
| 6060 | + | "cache_write_hour": 0, | |
| 6001 | 6061 | "cost_micros": 0, | |
| 6002 | 6062 | "charged_micros": 0, | |
| 6003 | 6063 | "status": 402, | |
| 6004 | 6064 | "own_key": false, | |
| 6065 | + | "format": "openai", | |
| 6066 | + | "provider": "", | |
| 6067 | + | "connection": null, | |
| 6005 | 6068 | "streamed": false, | |
| 6006 | 6069 | "duration_ms": 38, | |
| 6007 | 6070 | "error": "The flagon-io workspace is out of AI credit and has used this month's included usage, so the AI Gateway refuses requests to g1t's models. An owner can buy AI credit or turn on auto-reload at /flagon-io/-/billing#ai-credit." | |
| ⋯ | |||
| 6010 | 6073 | "next": "gw_1a7e3c5b9d2f4e6a8c0b2d41", | |
| 6011 | 6074 | "retention_days": 30 | |
| 6012 | 6075 | }, | |
| 6013 | − | "notes": "To send requests, see the AI Gateway guide: POST https://models.g1t.sh/anthropic/v1/messages with a workspace access token that has `models:write`. `charged_micros` is what the request was charged before included usage and AI credit paid for it; the payment itself is on the statement as AI Gateway. Pass `next` as `before` for the next page; it is null on the last." | |
| 6076 | + | "notes": "To send requests, see the AI Gateway guide: POST https://models.g1t.sh/anthropic/v1/messages (Anthropic's format) or https://models.g1t.sh/openai/v1/chat/completions (OpenAI's) with a workspace access token that has `models:write`. `format` is the format a request was sent in; `provider` who served it (`anthropic` or `workers-ai` on g1t's account, the connection's provider on the workspace's own, empty when it was refused before reaching one), and `connection` the workspace's own connection by name. `cache_write_hour` is the part of `cache_write` written to the hour-long cache. `charged_micros` is what the request was charged before included usage and AI credit paid for it; the payment itself is on the statement as AI Gateway. Pass `next` as `before` for the next page; it is null on the last." | |
| 6014 | 6077 | }, | |
| 6015 | 6078 | "list_invoices": { | |
| 6016 | 6079 | "params": { | |
| 485 | 485 | &[], | |
| 486 | 486 | ), | |
| 487 | 487 | route( | |
| 488 | + | "PATCH", | |
| 489 | + | "/workspaces/:workspace/integrations/:id", | |
| 490 | + | Op::UpdateIntegration, | |
| 491 | + | &[], | |
| 492 | + | ), | |
| 493 | + | route( | |
| 488 | 494 | "DELETE", | |
| 489 | 495 | "/workspaces/:workspace/integrations/:id", | |
| 490 | 496 | Op::DisconnectIntegration, |
| 287 | 287 | a("revoke_invite", Op::RevokeWorkspaceInvite, "Revoke a pending invite"), | |
| 288 | 288 | a("list_integrations", Op::ListIntegrations, "Model providers, alert sources, trackers"), | |
| 289 | 289 | a("connect_integration", Op::ConnectIntegration, "Connect one"), | |
| 290 | + | a("update_integration", Op::UpdateIntegration, "Change one: rotate its key, choose its AI Gateway models"), | |
| 290 | 291 | a("disconnect_integration", Op::DisconnectIntegration, "Remove one"), | |
| 291 | 292 | a("test_integration", Op::TestIntegration, "Check its credentials"), | |
| 292 | 293 | a("get_model_routes", Op::GetModelRoutes, "Where each kind of work's model requests go"), | |
| ⋯ | |||
| 417 | 418 | | Op::RemoveEmail | |
| 418 | 419 | | Op::RemoveCollaborator | |
| 419 | 420 | | Op::DisconnectIntegration | |
| 421 | + | | Op::UpdateIntegration | |
| 420 | 422 | | Op::DeleteWebhook | |
| 421 | 423 | | Op::DeleteActionsSecret | |
| 422 | 424 | | Op::DeleteActionsVariable | |
| 1 | 1 | --- | |
| 2 | 2 | title: AI Gateway | |
| 3 | − | description: Send your own code's model requests through g1t with a workspace access token, paid from AI credit at the model's price, with a log of every request. | |
| 3 | + | description: Send your own code's model requests through g1t in Anthropic's or OpenAI's format, with a workspace access token, to Claude, open models or your own providers, with a log of every request. | |
| 4 | 4 | --- | |
| 5 | 5 | ||
| 6 | − | The AI Gateway takes model requests from your own code, in Anthropic's | |
| 7 | − | Messages format, and sends them to the model. You point an Anthropic SDK, | |
| 8 | − | Claude Code or anything else that speaks that format at one base URL and | |
| 9 | − | give it a workspace access token as its API key: | |
| 6 | + | The AI Gateway takes model requests from your own code and sends them to | |
| 7 | + | the model. It speaks two formats, and any model works in either: | |
| 8 | + | ||
| 9 | + | | Format | Base URL | For | | |
| 10 | + | | --- | --- | --- | | |
| 11 | + | | Anthropic's Messages API | `https://models.g1t.sh/anthropic` | Anthropic's SDKs, Claude Code, and anything else that speaks that format | | |
| 12 | + | | OpenAI's Chat Completions API | `https://models.g1t.sh/openai/v1` | OpenAI's SDKs, and any tool that lets you set an OpenAI-compatible base URL | | |
| 10 | 13 | ||
| 11 | − | | | | | |
| 12 | − | | --- | --- | | |
| 13 | − | | Base URL | `https://models.g1t.sh/anthropic` | | |
| 14 | − | | API key | A workspace access token (`g1t_…`) with the `models:write` scope | | |
| 14 | + | In both, the API key is a workspace access token (`g1t_…`) with the | |
| 15 | + | `models:write` scope. | |
| 15 | 16 | ||
| 16 | − | Each request is charged to the workspace at the model's price and paid from | |
| 17 | − | the plan's included usage and [AI credit](/guides/usage-and-billing/#ai-credit). | |
| 18 | − | While the gateway is in beta there is no markup. If the workspace has | |
| 19 | − | connected its own Anthropic key, requests go there instead and cost | |
| 20 | − | nothing on g1t. Every request is logged with its model, tokens, cost and | |
| 21 | − | status. Prompts and answers are never kept. | |
| 17 | + | The model a request names decides where it goes: | |
| 18 | + | ||
| 19 | + | - **g1t's models**: Claude on Anthropic, and open models on Workers AI. | |
| 20 | + | Each request is charged to the workspace at the model's price and paid | |
| 21 | + | from the plan's included usage and | |
| 22 | + | [AI credit](/guides/usage-and-billing/#ai-credit). While the gateway is in | |
| 23 | + | beta there is no markup. | |
| 24 | + | - **Your own providers**: an Anthropic key, an OpenAI key, or any endpoint | |
| 25 | + | that speaks either API, connected under | |
| 26 | + | [Integrations](/guides/models/#connect-a-provider). You choose which models | |
| 27 | + | go to each. Those requests are counted and never charged on g1t. | |
| 28 | + | ||
| 29 | + | Every request is logged with its format, who served it, its model, tokens, | |
| 30 | + | cost and status. Prompts and answers are never kept. | |
| 22 | 31 | ||
| 23 | 32 | ## Before you start | |
| 24 | 33 | ||
| ⋯ | |||
| 26 | 35 | ||
| 27 | 36 | - **The workspace on the g1t plan**, with AI credit or this month's | |
| 28 | 37 | included usage left. See [AI credit](/guides/usage-and-billing/#ai-credit). | |
| 29 | − | - **The workspace's own Anthropic key**, connected under | |
| 30 | − | [Integrations](/guides/models/#connect-a-provider). Then the plan is not | |
| 31 | − | needed and nothing is charged. | |
| 38 | + | - **One of the workspace's own model providers**, connected under | |
| 39 | + | [Integrations](#your-own-providers). Requests for the models it takes need | |
| 40 | + | no plan and cost nothing on g1t. | |
| 32 | 41 | ||
| 33 | 42 | ## Make a token | |
| 34 | 43 | ||
| ⋯ | |||
| 51 | 60 | ||
| 52 | 61 | ## Send a request | |
| 53 | 62 | ||
| 54 | − | The gateway answers the same routes as Anthropic's API, below the base URL: | |
| 63 | + | The token goes in `x-api-key` or in `Authorization: Bearer`, in either | |
| 64 | + | format. The gateway answers these routes: | |
| 55 | 65 | ||
| 56 | 66 | | Route | What it does | | |
| 57 | 67 | | --- | --- | | |
| 58 | 68 | | `POST /anthropic/v1/messages` | A message, streamed (`"stream": true`) or whole. Logged and charged. | | |
| 59 | − | | `POST /anthropic/v1/messages/count_tokens` | Counts a request's input tokens. Not logged, and costs nothing. | | |
| 69 | + | | `POST /anthropic/v1/messages/count_tokens` | Counts a request's input tokens. Not logged, and costs nothing. For a model that does not speak Anthropic's API, an estimate. | | |
| 70 | + | | `POST /openai/v1/chat/completions` | A chat completion, streamed or whole. Logged and charged. | | |
| 71 | + | | `POST /openai/v1/embeddings` | Embeddings, from an embeddings model. Logged and charged by their input tokens. | | |
| 72 | + | | `GET /openai/v1/models` | The models this workspace can use, with g1t's prices. | | |
| 60 | 73 | ||
| 61 | − | The token goes in `x-api-key`, or in `Authorization: Bearer`. Request and | |
| 62 | − | answer bodies are Anthropic's, unchanged, and so are streamed events. | |
| 74 | + | ### Anthropic's format | |
| 63 | 75 | ||
| 64 | 76 | With curl: | |
| 65 | 77 | ||
| ⋯ | |||
| 69 | 81 | -H "anthropic-version: 2023-06-01" \ | |
| 70 | 82 | -H "content-type: application/json" \ | |
| 71 | 83 | -d '{ | |
| 72 | − | "model": "claude-sonnet-5-5", | |
| 84 | + | "model": "claude-haiku-5-5", | |
| 73 | 85 | "max_tokens": 1024, | |
| 74 | 86 | "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }] | |
| 75 | 87 | }' | |
| ⋯ | |||
| 86 | 98 | }); | |
| 87 | 99 | ||
| 88 | 100 | const message = await client.messages.create({ | |
| 89 | − | model: "claude-sonnet-5-5", | |
| 101 | + | model: "claude-haiku-5-5", | |
| 90 | 102 | max_tokens: 1024, | |
| 91 | 103 | messages: [{ role: "user", content: "Write a commit message for: fix the login redirect" }], | |
| 92 | 104 | }); | |
| ⋯ | |||
| 105 | 117 | ) | |
| 106 | 118 | ||
| 107 | 119 | message = client.messages.create( | |
| 108 | − | model="claude-sonnet-5-5", | |
| 120 | + | model="claude-haiku-5-5", | |
| 109 | 121 | max_tokens=1024, | |
| 110 | 122 | messages=[{"role": "user", "content": "Write a commit message for: fix the login redirect"}], | |
| 111 | 123 | ) | |
| 112 | 124 | ``` | |
| 113 | 125 | ||
| 114 | − | ### Claude Code | |
| 126 | + | Request and answer bodies are Anthropic's, and so are streamed events. To | |
| 127 | + | an open model, such as `workers-ai/@cf/openai/gpt-oss-120b`, the request is | |
| 128 | + | translated: messages, system prompt, images, tools and tool results, | |
| 129 | + | `tool_choice`, stop sequences, `output_config.effort` (as | |
| 130 | + | `reasoning_effort`, `xhigh` and `max` as `high`) and `output_config.format` | |
| 131 | + | (as a JSON schema). The answer comes back as an Anthropic message, tool | |
| 132 | + | calls included. Server tools have no counterpart there and are refused on | |
| 133 | + | g1t's models. | |
| 134 | + | ||
| 135 | + | ### OpenAI's format | |
| 136 | + | ||
| 137 | + | With curl: | |
| 138 | + | ||
| 139 | + | ```sh | |
| 140 | + | curl https://models.g1t.sh/openai/v1/chat/completions \ | |
| 141 | + | -H "Authorization: Bearer $G1T_TOKEN" \ | |
| 142 | + | -H "content-type: application/json" \ | |
| 143 | + | -d '{ | |
| 144 | + | "model": "anthropic/claude-haiku-5-5", | |
| 145 | + | "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }] | |
| 146 | + | }' | |
| 147 | + | ``` | |
| 148 | + | ||
| 149 | + | With OpenAI's TypeScript SDK: | |
| 150 | + | ||
| 151 | + | ```ts | |
| 152 | + | import OpenAI from "openai"; | |
| 153 | + | ||
| 154 | + | const client = new OpenAI({ | |
| 155 | + | baseURL: "https://models.g1t.sh/openai/v1", | |
| 156 | + | apiKey: process.env.G1T_TOKEN, | |
| 157 | + | }); | |
| 158 | + | ||
| 159 | + | const completion = await client.chat.completions.create({ | |
| 160 | + | model: "workers-ai/@cf/openai/gpt-oss-120b", | |
| 161 | + | messages: [{ role: "user", content: "Label this issue: the login page is blank on Safari" }], | |
| 162 | + | }); | |
| 163 | + | ``` | |
| 164 | + | ||
| 165 | + | With OpenAI's Python SDK: | |
| 166 | + | ||
| 167 | + | ```python | |
| 168 | + | import os | |
| 169 | + | ||
| 170 | + | from openai import OpenAI | |
| 171 | + | ||
| 172 | + | client = OpenAI( | |
| 173 | + | base_url="https://models.g1t.sh/openai/v1", | |
| 174 | + | api_key=os.environ["G1T_TOKEN"], | |
| 175 | + | ) | |
| 176 | + | ||
| 177 | + | completion = client.chat.completions.create( | |
| 178 | + | model="anthropic/claude-sonnet-5-5", | |
| 179 | + | messages=[{"role": "user", "content": "Summarize this diff in one sentence."}], | |
| 180 | + | stream=True, | |
| 181 | + | stream_options={"include_usage": True}, | |
| 182 | + | ) | |
| 183 | + | for chunk in completion: | |
| 184 | + | print(chunk.choices[0].delta.content or "" if chunk.choices else "", end="") | |
| 185 | + | ``` | |
| 186 | + | ||
| 187 | + | Embeddings: | |
| 188 | + | ||
| 189 | + | ```sh | |
| 190 | + | curl https://models.g1t.sh/openai/v1/embeddings \ | |
| 191 | + | -H "Authorization: Bearer $G1T_TOKEN" \ | |
| 192 | + | -H "content-type: application/json" \ | |
| 193 | + | -d '{ "model": "workers-ai/@cf/baai/bge-m3", "input": ["fix the login redirect"] }' | |
| 194 | + | ``` | |
| 195 | + | ||
| 196 | + | What OpenAI's format supports, to any model: | |
| 197 | + | ||
| 198 | + | | In the request | | | |
| 199 | + | | --- | --- | | |
| 200 | + | | `messages` | `system`, `developer`, `user`, `assistant` and `tool` messages. User content can be text, `image_url` (a URL or a `data:` URL) and, to Claude, `file` with `file_data` (a PDF as a `data:` URL). | | |
| 201 | + | | `tools`, `tool_choice`, `parallel_tool_calls` | Function tools. To Claude, `required` is `any` and a named function is that tool. | | |
| 202 | + | | `stream`, `stream_options.include_usage` | Server-sent chunks, ending with `data: [DONE]`. With `include_usage`, a last chunk carries the usage. | | |
| 203 | + | | `max_tokens`, `max_completion_tokens` | To Claude, 8,192 when neither is given. | | |
| 204 | + | | `temperature`, `top_p`, `stop`, `user` | As given. Some models refuse sampling settings. | | |
| 205 | + | | `reasoning_effort` | To Claude, `output_config.effort`; `minimal` and `none` are `low`. | | |
| 206 | + | | `response_format` | `json_schema` is a JSON schema the answer follows. `json_object` asks Claude for one JSON object. | | |
| 207 | + | | `thinking` | To Claude, passed as it is, for a caller that sets Anthropic's thinking. | | |
| 208 | + | ||
| 209 | + | To Claude, the answer's `usage` counts cached tokens in `prompt_tokens`, | |
| 210 | + | with `prompt_tokens_details.cached_tokens`; Claude's thinking comes back as | |
| 211 | + | `reasoning_content`. Claude's thinking blocks must go back with the tool | |
| 212 | + | calls they led to, so the gateway carries them in the first tool call's | |
| 213 | + | `id`: send the `id` back unchanged in the assistant message and the `tool` | |
| 214 | + | message, as OpenAI's SDKs do. `n` above 1 is refused for Claude. | |
| 215 | + | ||
| 216 | + | ### Claude Code and other tools | |
| 115 | 217 | ||
| 116 | − | Set two environment variables before you start it: | |
| 218 | + | Claude Code speaks Anthropic's format. Set two environment variables | |
| 219 | + | before you start it: | |
| 117 | 220 | ||
| 118 | 221 | ```sh | |
| 119 | 222 | export ANTHROPIC_BASE_URL=https://models.g1t.sh/anthropic | |
| ⋯ | |||
| 123 | 226 | ||
| 124 | 227 | Claude Code's own small requests go to Claude Haiku 4.5, which the gateway | |
| 125 | 228 | offers. To choose the main model, also set `ANTHROPIC_MODEL`, such as | |
| 126 | − | `claude-opus-5-5`. | |
| 229 | + | `claude-opus-5-5`, or `ANTHROPIC_SMALL_FAST_MODEL`, such as | |
| 230 | + | `claude-haiku-5-5`. | |
| 127 | 231 | ||
| 232 | + | Any other tool that lets you set an OpenAI-compatible base URL and key | |
| 233 | + | works the same way: give it `https://models.g1t.sh/openai/v1` and the | |
| 234 | + | token, and a model id from [the models list](#models). | |
| 235 | + | ||
| 128 | 236 | ## Models | |
| 129 | 237 | ||
| 130 | − | On g1t's models the gateway offers these. Prices are per million tokens, | |
| 131 | − | the provider's list price; cache writes are five-minute ones. | |
| 238 | + | ### Model ids | |
| 132 | 239 | ||
| 133 | − | | Model | `model` | Input | Output | Cache reads | Cache writes | | |
| 134 | − | | --- | --- | --- | --- | --- | --- | | |
| 135 | − | | Claude Opus 5.5 | `claude-opus-5-5` | $4.00 | $20.00 | $0.20 | $5.00 | | |
| 136 | − | | Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2.00 | $10.00 | $0.20 | $2.50 | | |
| 137 | − | | Claude Haiku 4.5 | `claude-haiku-4-5`, `claude-haiku-4-5-20251001` | $1.00 | $5.00 | $0.10 | $1.25 | | |
| 240 | + | A request names a model: | |
| 138 | 241 | ||
| 139 | − | A request for any other model is refused with `400` before it reaches the | |
| 140 | − | provider, and the error names the models offered. On the workspace's own | |
| 141 | − | key, a request can name any model that key can use. | |
| 242 | + | | Id | Goes to | | |
| 243 | + | | --- | --- | | |
| 244 | + | | `anthropic/claude-sonnet-5-5` | Claude on g1t's account, in either format | | |
| 245 | + | | `claude-sonnet-5-5` | The same, as Anthropic's API names it | | |
| 246 | + | | `workers-ai/@cf/openai/gpt-oss-120b` | An open model on g1t's account, in either format | | |
| 247 | + | | `@cf/openai/gpt-oss-120b` | The same | | |
| 248 | + | | Any id one of your own providers takes, such as `gpt-5.5` or `ollama/llama3.3` | That provider, with its key. See [your own providers](#your-own-providers). | | |
| 142 | 249 | ||
| 143 | − | On g1t's models a request is charged only by its tokens, so what the | |
| 144 | − | provider bills some other way is refused with `400` for now: | |
| 250 | + | Your own providers come first: when one of them takes a model, the request | |
| 251 | + | goes there, even one that names a model g1t offers. An Anthropic key takes | |
| 252 | + | `claude-*` unless you choose otherwise, so with one connected, Claude goes to | |
| 253 | + | your key. | |
| 145 | 254 | ||
| 255 | + | `GET /openai/v1/models` lists what the workspace can use: its own | |
| 256 | + | providers' models first, then g1t's, cheapest Claude first. Each has | |
| 257 | + | `billed_to` (`workspace` or `g1t`), `connection` (your provider's name) and, | |
| 258 | + | on g1t's models, `pricing` in dollars per million tokens: | |
| 259 | + | ||
| 260 | + | ```json | |
| 261 | + | { | |
| 262 | + | "object": "list", | |
| 263 | + | "data": [ | |
| 264 | + | { | |
| 265 | + | "id": "anthropic/claude-haiku-5-5", | |
| 266 | + | "object": "model", | |
| 267 | + | "created": 0, | |
| 268 | + | "owned_by": "anthropic", | |
| 269 | + | "name": "Claude Haiku 5.5", | |
| 270 | + | "kind": "chat", | |
| 271 | + | "billed_to": "g1t", | |
| 272 | + | "connection": null, | |
| 273 | + | "pricing": { | |
| 274 | + | "currency": "usd", | |
| 275 | + | "input": 0.1, | |
| 276 | + | "output": 0.5, | |
| 277 | + | "cache_read": 0.01, | |
| 278 | + | "cache_write": 0.125, | |
| 279 | + | "cache_write_1h": 0.2, | |
| 280 | + | "long_prompt": { "above_tokens": 100000, "input": 0.5, "output": 2.5, "cache_read": 0.05, "cache_write": 0.625, "cache_write_1h": 1 } | |
| 281 | + | } | |
| 282 | + | } | |
| 283 | + | ] | |
| 284 | + | } | |
| 285 | + | ``` | |
| 286 | + | ||
| 287 | + | ### Claude, on Anthropic | |
| 288 | + | ||
| 289 | + | Prices are per million tokens, Anthropic's list price. Cache writes are | |
| 290 | + | five-minute ones; one-hour cache writes (`"ttl": "1h"`) cost twice the | |
| 291 | + | input price. | |
| 292 | + | ||
| 293 | + | | Model | `model` | Input | Output | Cache reads | Cache writes | One-hour cache writes | | |
| 294 | + | | --- | --- | --- | --- | --- | --- | --- | | |
| 295 | + | | Claude Haiku 5.5 | `claude-haiku-5-5` | $0.10 | $0.50 | $0.01 | $0.125 | $0.20 | | |
| 296 | + | | Claude Haiku 5.5, prompts over 100,000 tokens | `claude-haiku-5-5` | $0.50 | $2.50 | $0.05 | $0.625 | $1.00 | | |
| 297 | + | | Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2.00 | $10.00 | $0.10 | $2.50 | $4.00 | | |
| 298 | + | | Claude Opus 5.5 | `claude-opus-5-5` | $4.00 | $20.00 | $0.20 | $5.00 | $8.00 | | |
| 299 | + | | Claude Haiku 4.5 | `claude-haiku-4-5`, `claude-haiku-4-5-20251001` | $1.00 | $5.00 | $0.10 | $1.25 | $2.00 | | |
| 300 | + | ||
| 301 | + | Claude Haiku 5.5 is the cheapest Claude and the one to start with. It is | |
| 302 | + | priced by the prompt's length: a request whose prompt (its input, cache | |
| 303 | + | read and cache write tokens) is longer than 100,000 tokens is charged | |
| 304 | + | entirely at the higher prices. It takes effort, like Opus: set | |
| 305 | + | `output_config.effort` in Anthropic's format, or `reasoning_effort` in | |
| 306 | + | OpenAI's. | |
| 307 | + | ||
| 308 | + | ### Open models, on Workers AI | |
| 309 | + | ||
| 310 | + | Prices are per million tokens, Cloudflare's list price. Workers AI has no | |
| 311 | + | prompt-cache price: cached tokens, where a model reports them, cost what | |
| 312 | + | input does. | |
| 313 | + | ||
| 314 | + | | Model | `model` | Input | Output | | |
| 315 | + | | --- | --- | --- | --- | | |
| 316 | + | | GLM-5.3 Flash | `workers-ai/@cf/zai-org/glm-5.3-flash` | $0.15 | $0.50 | | |
| 317 | + | | gpt-oss-20b | `workers-ai/@cf/openai/gpt-oss-20b` | $0.20 | $0.30 | | |
| 318 | + | | Llama 4 Scout | `workers-ai/@cf/meta/llama-4-scout-17b-16e-instruct` | $0.27 | $0.85 | | |
| 319 | + | | gpt-oss-120b | `workers-ai/@cf/openai/gpt-oss-120b` | $0.35 | $0.75 | | |
| 320 | + | | Mistral Small 3.1 | `workers-ai/@cf/mistralai/mistral-small-3.1-24b-instruct` | $0.351 | $0.555 | | |
| 321 | + | | DeepSeek V4 Flash | `workers-ai/@cf/deepseek-ai/deepseek-v4-flash-0731` | $0.44 | $1.32 | | |
| 322 | + | | Nemotron 3 120B | `workers-ai/@cf/nvidia/nemotron-3-120b-a12b` | $0.50 | $1.50 | | |
| 323 | + | | Kimi K2.6 | `workers-ai/@cf/moonshotai/kimi-k2.6` | $0.95 | $4.00 | | |
| 324 | + | | DeepSeek V4 Pro | `workers-ai/@cf/deepseek-ai/deepseek-v4-pro-0813` | $1.32 | $3.96 | | |
| 325 | + | | GLM-5.3 | `workers-ai/@cf/zai-org/glm-5.3` | $1.40 | $4.40 | | |
| 326 | + | ||
| 327 | + | Embeddings, through `POST /openai/v1/embeddings` only: | |
| 328 | + | ||
| 329 | + | | Model | `model` | Input | | |
| 330 | + | | --- | --- | --- | | |
| 331 | + | | BGE M3 | `workers-ai/@cf/baai/bge-m3` | $0.012 | | |
| 332 | + | | BGE Base (English) | `workers-ai/@cf/baai/bge-base-en-v1.5` | $0.067 | | |
| 333 | + | ||
| 334 | + | Open models cost much less per call than Claude, and suit one-shot work: | |
| 335 | + | titles, summaries, labels, triage, embeddings. In a long loop that sends | |
| 336 | + | the same context every turn, Claude's cache reads close most of that gap. | |
| 337 | + | ||
| 338 | + | ### What is refused on g1t's models | |
| 339 | + | ||
| 340 | + | A request for a model nobody offers is refused with `404` before it | |
| 341 | + | reaches a provider, and the error names the models offered. On g1t's | |
| 342 | + | models a request is charged only by its tokens, so what a provider bills | |
| 343 | + | some other way is refused with `400` for now: | |
| 344 | + | ||
| 146 | 345 | | Not offered on g1t's models yet | In the request | | |
| 147 | 346 | | --- | --- | | |
| 148 | 347 | | Fast mode | `speed` other than `standard` | | |
| 149 | 348 | | Inference in one region | `inference_geo` other than `global` | | |
| 150 | 349 | | Server-side fallbacks | `fallbacks` | | |
| 151 | − | | Server tools, such as web search, web fetch and code execution | A tool whose `type` is not your own (`custom` or none) or a client tool (`bash_…`, `text_editor_…`, `computer_…`, `memory_…`) | | |
| 350 | + | | Server tools, such as web search, web fetch and code execution | In Anthropic's format, a tool whose `type` is not your own (`custom` or none) or a client tool (`bash_…`, `text_editor_…`, `computer_…`, `memory_…`). In OpenAI's, a tool that is not a `function`, or `web_search_options`. | | |
| 152 | 351 | | Containers and skills | `container` | | |
| 153 | 352 | ||
| 154 | − | All of them work on the workspace's own key, which the provider bills. In | |
| 155 | − | Claude Code on g1t's models, its web search fails for this reason; the | |
| 156 | − | rest of Claude Code works. | |
| 353 | + | All of them work on your own provider, which bills them. In Claude Code on | |
| 354 | + | g1t's models, its web search fails for this reason; the rest of Claude Code | |
| 355 | + | works. | |
| 157 | 356 | ||
| 158 | − | Anthropic's format is the one served today. OpenAI's format and open | |
| 159 | − | models are coming later. | |
| 357 | + | ## Your own providers | |
| 358 | + | ||
| 359 | + | Connect a model provider under **Integrations**, and choose which models | |
| 360 | + | your own code's gateway requests send to it. Requests there use its key, | |
| 361 | + | are counted in the log, and are never charged on g1t. Any provider works: | |
| 160 | 362 | ||
| 363 | + | | Provider | Takes, unless you choose | | |
| 364 | + | | --- | --- | | |
| 365 | + | | An Anthropic key, or an Anthropic-compatible endpoint | `claude-*` | | |
| 366 | + | | An OpenAI key, or any other provider | Nothing until you choose | | |
| 367 | + | | An OpenAI-compatible endpoint: a self-hosted vLLM or Ollama, LiteLLM, another provider | Nothing until you choose | | |
| 368 | + | ||
| 369 | + | 1. Open the workspace's **Integrations** and choose a provider under | |
| 370 | + | **Model providers**. For your own server, choose **OpenAI-compatible | |
| 371 | + | endpoint** or **Anthropic-compatible endpoint** and give its base URL. | |
| 372 | + | 2. Paste its key. It is sealed when saved and never shown again: the page, | |
| 373 | + | the API and MCP show only its last four characters. | |
| 374 | + | 3. Under **AI Gateway models**, list the models to send there, separated by | |
| 375 | + | spaces: | |
| 376 | + | ||
| 377 | + | | Write | Takes | | |
| 378 | + | | --- | --- | | |
| 379 | + | | `gpt-5.5` | That model only | | |
| 380 | + | | `gpt-*` | Every model whose id starts with `gpt-` | | |
| 381 | + | | `ollama/*` | Every model named `ollama/…`, sent without the prefix: `ollama/llama3.3` arrives as `llama3.3` | | |
| 382 | + | | `*` | Every model | | |
| 383 | + | | Nothing | No gateway requests | | |
| 384 | + | ||
| 385 | + | 4. Select **Connect**. To change the list or replace the key later, open | |
| 386 | + | **Change its AI Gateway models or key** under the provider. | |
| 387 | + | ||
| 388 | + | The first provider, in the order they were connected, that takes a model | |
| 389 | + | gets its requests. Either format reaches either kind of provider: a Claude | |
| 390 | + | key answers OpenAI-format requests, and an OpenAI-compatible endpoint | |
| 391 | + | answers Claude Code. Agent runs choose their models under | |
| 392 | + | [routing](/guides/models/), apart from this list. | |
| 393 | + | ||
| 394 | + | From code, connect one with | |
| 395 | + | [`POST /workspaces/{workspace}/integrations`](/reference/api/integrations/connect-integration/) | |
| 396 | + | and change it with | |
| 397 | + | [`PATCH /workspaces/{workspace}/integrations/{id}`](/reference/api/integrations/update-integration/), | |
| 398 | + | with `config.gateway_models`, or the `workspace` MCP tool's | |
| 399 | + | `connect_integration` and `update_integration` actions. Both need | |
| 400 | + | `workspace:admin` and act for an owner. The key is write-only: neither | |
| 401 | + | returns it. | |
| 402 | + | ||
| 403 | + | ```sh | |
| 404 | + | curl -X PATCH https://api.g1t.sh/workspaces/acme/integrations/con_01kpx5c2d8e4f6g0h2j4k6m8n0 \ | |
| 405 | + | -H "Authorization: Bearer $G1T_ADMIN_TOKEN" \ | |
| 406 | + | -H "content-type: application/json" \ | |
| 407 | + | -d '{ "config": { "base_url": "https://gpu.acme.dev/v1", "gateway_models": ["ollama/*"] }, "secret": "…" }' | |
| 408 | + | ``` | |
| 409 | + | ||
| 410 | + | `config` replaces the provider's settings whole, so send the ones it has | |
| 411 | + | with the change. | |
| 412 | + | ||
| 161 | 413 | ## What it costs | |
| 162 | 414 | ||
| 163 | 415 | | Where it goes | You pay | | |
| 164 | 416 | | --- | --- | | |
| 165 | 417 | | g1t's models | Its tokens at the model's price above, with no markup while the gateway is in beta | | |
| 166 | − | | The workspace's own Anthropic key | Nothing on g1t. The provider bills you for the model. | | |
| 418 | + | | Your own providers | Nothing on g1t. The provider bills you for the model. | | |
| 167 | 419 | ||
| 168 | 420 | On g1t's models: | |
| 169 | 421 | ||
| 170 | 422 | - Each request that used tokens is one line on the statement, under | |
| 171 | 423 | **AI Gateway**, such as *AI Gateway: Claude Sonnet 5.5, 14,352 tokens, | |
| 172 | − | token release-notes*. | |
| 424 | + | token release-notes*. A Claude Haiku 5.5 request over 100,000 prompt | |
| 425 | + | tokens says *long-prompt price*. | |
| 173 | 426 | - The plan's included usage pays first, then AI credit. Trial credit and | |
| 174 | 427 | g1t's open-source pool never pay for gateway requests. | |
| 175 | 428 | - It is not an agent run, so the [agent rate](/guides/usage-and-billing/#the-agent-rate) | |
| ⋯ | |||
| 178 | 431 | like any other usage, and shows on **Usage** under the AI Gateway product. | |
| 179 | 432 | - A workspace with a 100% discount gets it free through the discount; an | |
| 180 | 433 | enterprise is invoiced for it after use. | |
| 181 | − | ||
| 182 | − | ### Your own key | |
| 183 | − | ||
| 184 | − | When the workspace has an Anthropic or Anthropic-compatible model provider | |
| 185 | − | under [Integrations](/guides/models/), the gateway sends every request to | |
| 186 | − | the first one connected, with its key. Those requests are logged with their | |
| 187 | − | tokens and marked **Own key**, and g1t charges nothing for them. Remove the | |
| 188 | − | provider and requests go to g1t's models again within a few seconds. | |
| 189 | 434 | ||
| 190 | 435 | ## Limits and errors | |
| 191 | 436 | ||
| ⋯ | |||
| 196 | 441 | usage left. If auto-reload is on, g1t tries it first. | |
| 197 | 442 | - It is not on the g1t plan. | |
| 198 | 443 | ||
| 199 | − | Errors are Anthropic's shape, so SDKs raise their usual errors: | |
| 444 | + | Errors are in the format of the route, so SDKs raise their usual errors. | |
| 445 | + | Anthropic's: | |
| 200 | 446 | ||
| 201 | 447 | ```json | |
| 202 | 448 | { "type": "error", "error": { "type": "billing_error", "message": "The acme workspace is out of AI credit …" } } | |
| 203 | 449 | ``` | |
| 204 | 450 | ||
| 205 | − | | Status | `error.type` | Why | | |
| 206 | − | | --- | --- | --- | | |
| 207 | − | | `400` | `invalid_request_error` | The body is not JSON, the model is not offered, or the request asks for something [not offered on g1t's models yet](#models). | | |
| 208 | − | | `401` | `authentication_error` | The token is unknown, expired or deleted. | | |
| 209 | − | | `402` | `billing_error` | Out of AI credit, over the spend limit, or not on the plan. The message says what an owner can do. | | |
| 210 | − | | `403` | `permission_error` | Not a workspace's token, or it lacks `models:write`. | | |
| 211 | − | | `404` | `not_found_error` | A route the gateway does not answer. | | |
| 451 | + | OpenAI's: | |
| 212 | 452 | ||
| 213 | − | An error from the model provider, such as `429` or `529`, comes back as | |
| 214 | − | the provider sent it. Refused and failed requests are logged with their | |
| 215 | − | status and why, and cost nothing. | |
| 453 | + | ```json | |
| 454 | + | { "error": { "message": "The acme workspace is out of AI credit …", "type": "insufficient_quota", "param": null, "code": "insufficient_quota" } } | |
| 455 | + | ``` | |
| 456 | + | ||
| 457 | + | | Status | Anthropic's `error.type` | OpenAI's `error.type` (`code`) | Why | | |
| 458 | + | | --- | --- | --- | --- | | |
| 459 | + | | `400` | `invalid_request_error` | `invalid_request_error` | The body is not JSON, the model is of the wrong kind, the request asks for something [not offered on g1t's models yet](#what-is-refused-on-g1ts-models), or it cannot be said to the model (such as `n` above 1 to Claude). | | |
| 460 | + | | `401` | `authentication_error` | `authentication_error` (`invalid_api_key`) | The token is unknown, expired or deleted, or your provider refused its key. | | |
| 461 | + | | `402` | `billing_error` | `insufficient_quota` (`insufficient_quota`) | Out of AI credit, over the spend limit, or not on the plan. The message says what an owner can do. | | |
| 462 | + | | `403` | `permission_error` | `permission_error` | Not a workspace's token, or it lacks `models:write`. | | |
| 463 | + | | `404` | `not_found_error` | `invalid_request_error` (`model_not_found`) | No provider offers the model, or a route the gateway does not answer. | | |
| 464 | + | ||
| 465 | + | An error from the model provider, such as `429` or `529`, comes back with | |
| 466 | + | its status and message, in the route's format. A provider's key never | |
| 467 | + | appears in an error or the log, even when the provider quotes it. Refused | |
| 468 | + | and failed requests are logged with their status and why, and cost | |
| 469 | + | nothing. Every answer carries `x-g1t-request-id`, the request's id in the | |
| 470 | + | log. | |
| 216 | 471 | ||
| 217 | − | A deleted token, a provider added under Integrations, or AI credit just | |
| 218 | − | bought takes effect within about ten seconds. | |
| 472 | + | A deleted token, a provider added or changed under Integrations, or AI | |
| 473 | + | credit just bought takes effect within about ten seconds. | |
| 219 | 474 | ||
| 220 | 475 | ## See every request | |
| 221 | 476 | ||
| ⋯ | |||
| 226 | 481 | | Column | | | |
| 227 | 482 | | --- | --- | | |
| 228 | 483 | | Time | When it was sent. Hover for the exact time, how long it took and whether it streamed. | | |
| 229 | − | | Model | The model it named. On the workspace's own key, the one that answered. | | |
| 230 | − | | Input, Output, Cache read, Cache write | Its tokens by kind. | | |
| 231 | − | | Cost | What it was charged, before included usage and AI credit paid for it, or **Own key**. | | |
| 484 | + | | Model | The model it named on g1t's models, or the one that answered on your own provider; below it, the format it was sent in. | | |
| 485 | + | | Served by | g1t's account and the provider (*g1t · Anthropic*, *g1t · Workers AI*), or your provider by name. *None* when it was refused first. | | |
| 486 | + | | Input, Output | Its tokens by kind. Hover Input for all of them. | | |
| 487 | + | | Cache | Cache reads, then cache writes. Hover for how many writes were to the one-hour cache. | | |
| 488 | + | | Cost | What it was charged, before included usage and AI credit paid for it, or **Not charged** on your own provider. | | |
| 232 | 489 | | Status | The status it was answered with. Hover a refusal or failure for why. | | |
| 233 | 490 | | Token | The name of the token that sent it. | | |
| 234 | 491 | ||
| ⋯ | |||
| 237 | 494 | From code, list them with | |
| 238 | 495 | [`GET /workspaces/{workspace}/gateway/requests`](/reference/api/billing/list-gateway-requests/), | |
| 239 | 496 | or the `billing` MCP tool's | |
| 240 | − | [`gateway_requests`](/reference/mcp/#billing) action. Both need | |
| 241 | − | `models:read`, which the Read only and Agent presets include. | |
| 497 | + | [`gateway_requests`](/reference/mcp/#billing) action. Each request has | |
| 498 | + | `format`, `provider`, `connection`, `model` and its tokens, with | |
| 499 | + | `cache_write_hour` for one-hour cache writes. Both need `models:read`, | |
| 500 | + | which the Read only and Agent presets include. | |
| 208 | 208 | | --- | --- | | |
| 209 | 209 | | `workspace` `list_integrations` | `GET /workspaces/{workspace}/integrations` | | |
| 210 | 210 | | `workspace` `connect_integration` | `POST /workspaces/{workspace}/integrations` | | |
| 211 | + | | `workspace` `update_integration` | `PATCH /workspaces/{workspace}/integrations/{id}` | | |
| 211 | 212 | | `workspace` `test_integration` | `POST /workspaces/{workspace}/integrations/{id}/test` | | |
| 212 | 213 | | `workspace` `disconnect_integration` | `DELETE /workspaces/{workspace}/integrations/{id}` | | |
| 213 | 214 | | `search` `ticket` | `GET /repos/{owner}/{name}/context?reference=` | |
| 18 | 18 | kind of work, which provider and model it runs on. Each provider bills | |
| 19 | 19 | you for the model directly. Open to every workspace now. | |
| 20 | 20 | ||
| 21 | − | To call models from your own code with a workspace token, paid from the | |
| 22 | − | same AI credit, use the [AI Gateway](/guides/ai-gateway/). | |
| 21 | + | To call models from your own code with a workspace token, in Anthropic's | |
| 22 | + | or OpenAI's format, paid from the same AI credit or sent to these same | |
| 23 | + | providers, use the [AI Gateway](/guides/ai-gateway/). Each provider's | |
| 24 | + | **AI Gateway models** say which of those requests go to it; see | |
| 25 | + | [your own providers](/guides/ai-gateway/#your-own-providers). | |
| 23 | 26 | ||
| 24 | 27 | ## Auto | |
| 25 | 28 |
| 548 | 548 | | [`revoke_invite`](/reference/api/invites/revoke-workspace-invite/) | Revoke a workspace's pending invite. Owners only. | `workspace`, `id` | `workspace:admin` | | |
| 549 | 549 | | [`list_integrations`](/reference/api/integrations/list-integrations/) | The workspace's connections. Secrets are never returned. Members only. | `workspace` | `workspace:read` | | |
| 550 | 550 | | [`connect_integration`](/reference/api/integrations/connect-integration/) | Connect a model provider (Anthropic, OpenAI, Gemini, or a compatible endpoint), Sentry, Datadog, a webhook, Jira or Linear, with `config` and `secret`. Owners only. | `workspace`, `provider` | `workspace:admin` | | |
| 551 | + | | [`update_integration`](/reference/api/integrations/update-integration/) | Change one: its `name`, its `config` (replaced whole) or its `secret` (write-only, never returned). Rotates a model provider's key, or sets `config.gateway_models`, the [AI Gateway](/guides/ai-gateway/#your-own-providers) models it takes. Owners only. | `workspace`, `id` | `workspace:admin` | | |
| 551 | 552 | | [`disconnect_integration`](/reference/api/integrations/disconnect-integration/) | Remove it and its secrets. Owners only. | `workspace`, `id` | `workspace:admin` | | |
| 552 | 553 | | [`test_integration`](/reference/api/integrations/test-integration/) | Check its credentials against the system it connects to. Owners only. | `workspace`, `id` | `workspace:admin` | | |
| 553 | 554 | | [`get_model_routes`](/reference/api/integrations/get-model-routes/) | Which provider and model each kind of work goes to. Members only. | `workspace` | `workspace:read` | | |
| ⋯ | |||
| 583 | 584 | | [`buy_ai_credit`](/reference/api/billing/buy-ai-credit/) | A payment page (`url`) to buy `amount_cents` of credit, in whole dollars from $10 to $1,000, for a person to open and pay; it returns to the workspace's billing page. Owners, as people. | `workspace`, `amount_cents` | `billing:write` | | |
| 584 | 585 | | [`invoices`](/reference/api/billing/list-invoices/) | Every invoice (`invoices`, in cents), g1t's itemised usage invoices (`usage_invoices`), and what the next one comes to so far (`upcoming`). | `workspace` | `billing:read` | | |
| 585 | 586 | | [`billing_details`](/reference/api/billing/get-billing-details/) | Who invoices are made out to, and the payment method on file as far as it is safe to show. | `workspace` | `billing:read` | | |
| 586 | − | | [`gateway_requests`](/reference/api/billing/list-gateway-requests/) | The workspace's recent [AI Gateway](/guides/ai-gateway/) requests, newest first: model, tokens by kind, `cost_micros`, `charged_micros`, `status`, `own_key` and the token that sent each. `limit` (50, at most 200) and `before` (the last page's `next`) page through them. Kept 30 days. | `workspace` | `models:read` | | |
| 587 | + | | [`gateway_requests`](/reference/api/billing/list-gateway-requests/) | The workspace's recent [AI Gateway](/guides/ai-gateway/) requests, newest first: model, `format` (`anthropic` or `openai`), `provider` and `connection` (who served it), tokens by kind with `cache_write_hour`, `cost_micros`, `charged_micros`, `status`, `own_key` and the token that sent each. `limit` (50, at most 200) and `before` (the last page's `next`) page through them. Kept 30 days. | `workspace` | `models:read` | | |
| 587 | 588 | ||
| 588 | 589 | ## `notifications` | |
| 589 | 590 | ||
Binary or large file; its contents are not shown.
| 7 | 7 | ||
| 8 | 8 | import type { GatewayRequest } from "@g1t/contracts"; | |
| 9 | 9 | ||
| 10 | − | import { duration, shortCount, statusTone, tokenKinds } from "../lib/gateway"; | |
| 10 | + | import { cacheKinds, duration, formatLabel, servedBy, shortCount, statusTone, tokenKinds } from "../lib/gateway"; | |
| 11 | 11 | import { money } from "../lib/usage"; | |
| 12 | 12 | import { TimeAgo } from "./ui"; | |
| 13 | 13 | import { Badge } from "./ui/badge"; | |
| ⋯ | |||
| 18 | 18 | const NUM = "px-3 py-2 text-right tabular-nums whitespace-nowrap"; | |
| 19 | 19 | ||
| 20 | 20 | /** The columns, so the skeleton has as many as the table. */ | |
| 21 | − | const COLUMNS = ["Time", "Model", "Input", "Output", "Cache read", "Cache write", "Cost", "Status", "Token"] as const; | |
| 21 | + | const COLUMNS = ["Time", "Model", "Served by", "Input", "Output", "Cache", "Cost", "Status", "Token"] as const; | |
| 22 | 22 | ||
| 23 | 23 | /** One request's row. */ | |
| 24 | 24 | function Row({ request }: { request: GatewayRequest }) { | |
| 25 | 25 | const status = statusTone(request.status); | |
| 26 | + | const served = servedBy(request); | |
| 26 | 27 | return ( | |
| 27 | 28 | <tr className="text-sm"> | |
| 28 | 29 | <td className="px-3 py-2 whitespace-nowrap text-muted"> | |
| ⋯ | |||
| 32 | 33 | </span> | |
| 33 | 34 | </Hint> | |
| 34 | 35 | </td> | |
| 35 | − | <td className="px-3 py-2 font-mono text-xs whitespace-nowrap">{request.model}</td> | |
| 36 | + | <td className="max-w-[15rem] px-3 py-2 whitespace-nowrap"> | |
| 37 | + | <Hint label={request.model}> | |
| 38 | + | <span tabIndex={0} className="block truncate font-mono text-xs"> | |
| 39 | + | {request.model} | |
| 40 | + | </span> | |
| 41 | + | </Hint> | |
| 42 | + | <span className="block text-[0.6875rem] text-faint">{formatLabel(request.format)} format</span> | |
| 43 | + | </td> | |
| 44 | + | <td className="max-w-[12rem] truncate px-3 py-2 whitespace-nowrap"> | |
| 45 | + | <Hint label={served.hint}> | |
| 46 | + | <span tabIndex={0} className={request.ownKey ? "" : "text-muted"}> | |
| 47 | + | {served.label} | |
| 48 | + | </span> | |
| 49 | + | </Hint> | |
| 50 | + | </td> | |
| 36 | 51 | <td className={NUM}> | |
| 37 | 52 | <Hint label={tokenKinds(request)}> | |
| 38 | 53 | <span tabIndex={0}>{shortCount(request.input)}</span> | |
| 39 | 54 | </Hint> | |
| 40 | 55 | </td> | |
| 41 | 56 | <td className={NUM}>{shortCount(request.output)}</td> | |
| 42 | − | <td className={NUM}>{shortCount(request.cacheRead)}</td> | |
| 43 | − | <td className={NUM}>{shortCount(request.cacheWrite)}</td> | |
| 44 | 57 | <td className={NUM}> | |
| 58 | + | <Hint label={cacheKinds(request)}> | |
| 59 | + | <span tabIndex={0}> | |
| 60 | + | {shortCount(request.cacheRead)} | |
| 61 | + | <span className="text-faint"> / </span> | |
| 62 | + | {shortCount(request.cacheWrite)} | |
| 63 | + | </span> | |
| 64 | + | </Hint> | |
| 65 | + | </td> | |
| 66 | + | <td className={NUM}> | |
| 45 | 67 | {request.ownKey ? ( | |
| 46 | − | <Hint label="Sent with the workspace's own provider key: counted, not charged."> | |
| 68 | + | <Hint label="Sent to the workspace's own provider: counted, not charged."> | |
| 47 | 69 | <span tabIndex={0}> | |
| 48 | − | <Badge>Own key</Badge> | |
| 70 | + | <Badge>Not charged</Badge> | |
| 49 | 71 | </span> | |
| 50 | 72 | </Hint> | |
| 51 | 73 | ) : ( | |
| ⋯ | |||
| 69 | 91 | } | |
| 70 | 92 | ||
| 71 | 93 | /** Whether a column holds a number, read right-aligned. */ | |
| 72 | − | const numeric = (index: number) => index >= 2 && index <= 6; | |
| 94 | + | const numeric = (index: number) => index >= 3 && index <= 6; | |
| 73 | 95 | ||
| 74 | 96 | /** The table's frame: its heading row, around its body. */ | |
| 75 | 97 | function Frame({ children }: { children: ReactNode }) { | |
| ⋯ | |||
| 114 | 136 | <tr key={row} className="text-sm"> | |
| 115 | 137 | {COLUMNS.map((column, index) => ( | |
| 116 | 138 | <td key={column} className={numeric(index) ? NUM : "px-3 py-2"}> | |
| 117 | − | <SkeletonLine className={numeric(index) ? "ml-auto w-10" : index === 1 ? "w-36" : "w-16"} /> | |
| 139 | + | <SkeletonLine className={numeric(index) ? "ml-auto w-10" : index === 1 ? "w-36" : index === 2 ? "w-24" : "w-16"} /> | |
| 118 | 140 | </td> | |
| 119 | 141 | ))} | |
| 120 | 142 | </tr> | |
| 7 | 7 | import { type ReactNode, useState } from "react"; | |
| 8 | 8 | import { Form, Link } from "react-router"; | |
| 9 | 9 | ||
| 10 | − | import { type Connection, MODEL_TASKS, MODEL_TIERS, type ModelRoute, type ModelTask, type ModelTier, type Provider, PROVIDERS } from "@g1t/contracts"; | |
| 10 | + | import { type Connection, MODEL_TASKS, MODEL_TIERS, type ModelRoute, type ModelTask, type ModelTier, type Provider, PROVIDERS, gatewayPatterns } from "@g1t/contracts"; | |
| 11 | 11 | ||
| 12 | 12 | import { Field, Input, SubmitButton } from "./ui"; | |
| 13 | 13 | import { Hint } from "./ui/hint"; | |
| ⋯ | |||
| 259 | 259 | </Field> | |
| 260 | 260 | ); | |
| 261 | 261 | ||
| 262 | + | /** What to suggest a provider's AI Gateway models be, as typed. */ | |
| 263 | + | const GATEWAY_EXAMPLES: Partial<Record<Provider, string>> = { | |
| 264 | + | openai: "gpt-*", | |
| 265 | + | gemini: "gemini-*", | |
| 266 | + | xai: "grok-*", | |
| 267 | + | mistral: "mistral-* codestral-*", | |
| 268 | + | deepseek: "deepseek-*", | |
| 269 | + | openrouter: "openrouter/*", | |
| 270 | + | groq: "groq/*", | |
| 271 | + | together: "together/*", | |
| 272 | + | fireworks: "fireworks/*", | |
| 273 | + | cerebras: "cerebras/*", | |
| 274 | + | openai_endpoint: "ollama/*", | |
| 275 | + | anthropic_endpoint: "claude-*", | |
| 276 | + | azure_openai: "azure/*", | |
| 277 | + | }; | |
| 278 | + | ||
| 279 | + | /** | |
| 280 | + | * Which of the AI Gateway's requests go to a model provider, by the model | |
| 281 | + | * they name: for connecting one, and for changing it after. | |
| 282 | + | */ | |
| 283 | + | export function GatewayModelsField({ provider, value }: { provider: Provider; value?: string[] }) { | |
| 284 | + | const defaults = gatewayPatterns(provider, {}); | |
| 285 | + | return ( | |
| 286 | + | <Field | |
| 287 | + | label="AI Gateway models" | |
| 288 | + | hint="Your own code's requests to the AI Gateway that name one of these models come here, counted and never charged by g1t. Model ids, or prefixes ending in *; a prefix ending in /* is taken off before sending, so ollama/llama3.3 arrives as llama3.3. Separate them with spaces. Empty: none." | |
| 289 | + | > | |
| 290 | + | <Input | |
| 291 | + | name="gatewayModels" | |
| 292 | + | defaultValue={(value ?? defaults).join(" ")} | |
| 293 | + | placeholder={GATEWAY_EXAMPLES[provider] ?? "model-id or prefix-*"} | |
| 294 | + | className="font-mono text-[0.8125rem]" | |
| 295 | + | autoComplete="off" | |
| 296 | + | spellCheck={false} | |
| 297 | + | /> | |
| 298 | + | </Field> | |
| 299 | + | ); | |
| 300 | + | } | |
| 301 | + | ||
| 262 | 302 | /** What connecting a model provider asks for. */ | |
| 263 | 303 | export function ModelProviderFields({ provider }: { provider: Provider }): ReactNode { | |
| 304 | + | return ( | |
| 305 | + | <> | |
| 306 | + | <ModelProviderKeyFields provider={provider} /> | |
| 307 | + | <GatewayModelsField provider={provider} /> | |
| 308 | + | </> | |
| 309 | + | ); | |
| 310 | + | } | |
| 311 | + | ||
| 312 | + | function ModelProviderKeyFields({ provider }: { provider: Provider }): ReactNode { | |
| 264 | 313 | const entry = MODEL_CATALOG[provider]; | |
| 265 | 314 | if (provider === "azure_openai") { | |
| 266 | 315 | return ( | |
| 1 | 1 | import assert from "node:assert/strict"; | |
| 2 | 2 | import { test } from "node:test"; | |
| 3 | 3 | ||
| 4 | − | import { duration, shortCount, statusTone, tokenKinds, totalTokens } from "./gateway.ts"; | |
| 4 | + | import { cacheKinds, duration, formatLabel, parseGatewayModels, servedBy, shortCount, statusTone, tokenKinds, totalTokens } from "./gateway.ts"; | |
| 5 | 5 | ||
| 6 | 6 | test("token counts read short", () => { | |
| 7 | 7 | assert.equal(shortCount(812), "812"); | |
| ⋯ | |||
| 29 | 29 | assert.equal(duration(4_210), "4.2 s"); | |
| 30 | 30 | assert.equal(duration(125_000), "2 min 5 s"); | |
| 31 | 31 | }); | |
| 32 | + | ||
| 33 | + | test("a request says which format it came in and who served it", () => { | |
| 34 | + | assert.equal(formatLabel("openai"), "OpenAI"); | |
| 35 | + | assert.equal(formatLabel("anthropic"), "Anthropic"); | |
| 36 | + | assert.deepEqual(servedBy({ ownKey: false, provider: "workers-ai", connection: null }).label, "g1t · Workers AI"); | |
| 37 | + | assert.deepEqual(servedBy({ ownKey: false, provider: "anthropic", connection: null }).label, "g1t · Anthropic"); | |
| 38 | + | const own = servedBy({ ownKey: true, provider: "openai_endpoint", connection: "Office GPU" }); | |
| 39 | + | assert.equal(own.label, "Office GPU"); | |
| 40 | + | assert.match(own.hint, /OpenAI-compatible endpoint: counted, not charged/); | |
| 41 | + | assert.equal(servedBy({ ownKey: false, provider: "", connection: null }).label, "None"); | |
| 42 | + | }); | |
| 43 | + | ||
| 44 | + | test("a connection's AI Gateway models are typed as a list", () => { | |
| 45 | + | assert.deepEqual(parseGatewayModels(" gpt-*, ollama/* claude-haiku-5-5 "), ["gpt-*", "ollama/*", "claude-haiku-5-5"]); | |
| 46 | + | assert.deepEqual(parseGatewayModels(""), []); | |
| 47 | + | }); | |
| 48 | + | ||
| 49 | + | test("cache tokens read in words, hour-long writes named", () => { | |
| 50 | + | assert.equal(cacheKinds({ cacheRead: 12_000, cacheWrite: 2_400, cacheWriteHour: 0 }), "12,000 read from the cache, 2,400 written"); | |
| 51 | + | assert.equal(cacheKinds({ cacheRead: 0, cacheWrite: 2_400, cacheWriteHour: 2_400 }), "0 read from the cache, 2,400 written, 2,400 of them to the hour-long cache"); | |
| 52 | + | }); | |
| 10 | 10 | /** The base URL an Anthropic SDK or Claude Code is pointed at. */ | |
| 11 | 11 | export const GATEWAY_BASE_URL = "https://models.g1t.sh/anthropic"; | |
| 12 | 12 | ||
| 13 | + | /** The base URL an OpenAI SDK, or any tool that speaks OpenAI's API, is pointed at. */ | |
| 14 | + | export const GATEWAY_OPENAI_BASE_URL = "https://models.g1t.sh/openai/v1"; | |
| 15 | + | ||
| 16 | + | /** How a request's format reads. */ | |
| 17 | + | export function formatLabel(format: GatewayRequest["format"] | undefined): string { | |
| 18 | + | return format === "openai" ? "OpenAI" : "Anthropic"; | |
| 19 | + | } | |
| 20 | + | ||
| 21 | + | /** Who served a request, as people read it, and a longer line for its hint. */ | |
| 22 | + | export function servedBy(request: Pick<GatewayRequest, "ownKey" | "provider" | "connection">): { label: string; hint: string } { | |
| 23 | + | if (request.ownKey) { | |
| 24 | + | const name = request.connection || "Your provider"; | |
| 25 | + | return { label: name, hint: `${name}, the workspace's own ${providerName(request.provider)}: counted, not charged.` }; | |
| 26 | + | } | |
| 27 | + | if (!request.provider) return { label: "None", hint: "Refused before it reached a model." }; | |
| 28 | + | return { label: `g1t · ${providerName(request.provider)}`, hint: `g1t's account at ${providerName(request.provider)}, charged at the model's price.` }; | |
| 29 | + | } | |
| 30 | + | ||
| 31 | + | const PROVIDER_NAMES: Record<string, string> = { | |
| 32 | + | anthropic: "Anthropic", | |
| 33 | + | "workers-ai": "Workers AI", | |
| 34 | + | openai: "OpenAI", | |
| 35 | + | openai_endpoint: "OpenAI-compatible endpoint", | |
| 36 | + | anthropic_endpoint: "Anthropic-compatible endpoint", | |
| 37 | + | azure_openai: "Azure OpenAI", | |
| 38 | + | gemini: "Google Gemini", | |
| 39 | + | openrouter: "OpenRouter", | |
| 40 | + | }; | |
| 41 | + | ||
| 42 | + | /** A provider's name, as people know it. */ | |
| 43 | + | export function providerName(provider: string): string { | |
| 44 | + | return PROVIDER_NAMES[provider] ?? (provider || "provider"); | |
| 45 | + | } | |
| 46 | + | ||
| 47 | + | /** A connection's AI Gateway models, as a field shows them, from what was typed. */ | |
| 48 | + | export function parseGatewayModels(text: string): string[] { | |
| 49 | + | return text | |
| 50 | + | .split(/[\s,]+/) | |
| 51 | + | .map((model) => model.trim()) | |
| 52 | + | .filter(Boolean); | |
| 53 | + | } | |
| 54 | + | ||
| 13 | 55 | /** A count of tokens, short: `812`, `12.4K`, `3.1M`. */ | |
| 14 | 56 | export function shortCount(n: number): string { | |
| 15 | 57 | if (n < 1_000) return String(n); | |
| ⋯ | |||
| 32 | 74 | return `${n(request.input)} input, ${n(request.output)} output, ${n(request.cacheRead)} cache read, ${n(request.cacheWrite)} cache write`; | |
| 33 | 75 | } | |
| 34 | 76 | ||
| 77 | + | /** A request's cache tokens in words, for a hint: read, written, and how many of the writes last an hour. */ | |
| 78 | + | export function cacheKinds(request: Pick<GatewayRequest, "cacheRead" | "cacheWrite" | "cacheWriteHour">): string { | |
| 79 | + | const n = (v: number) => v.toLocaleString("en-US"); | |
| 80 | + | const hour = request.cacheWriteHour ? `, ${n(request.cacheWriteHour)} of them to the hour-long cache` : ""; | |
| 81 | + | return `${n(request.cacheRead)} read from the cache, ${n(request.cacheWrite)} written${hour}`; | |
| 82 | + | } | |
| 83 | + | ||
| 35 | 84 | /** How a status reads, and how much it matters. */ | |
| 36 | 85 | export function statusTone(status: number): { label: string; tone: "success" | "warn" | "danger" | "neutral" } { | |
| 37 | 86 | if (status >= 200 && status < 300) return { label: String(status), tone: "success" }; | |
| 536 | 536 | <p className="font-medium"> | |
| 537 | 537 | AI Gateway <span className="ml-1 rounded bg-accent/15 px-1.5 py-0.5 text-xs text-accent">Free during beta</span> | |
| 538 | 538 | </p> | |
| 539 | − | <p className="text-xs text-faint">Your own code calling Claude models through g1t with a workspace token, paid from AI credit; with your own Anthropic key, free</p> | |
| 539 | + | <p className="text-xs text-faint">Your own code calling Claude and open models through g1t, in Anthropic's or OpenAI's format, paid from AI credit; on your own providers, free</p> | |
| 540 | 540 | </td> | |
| 541 | 541 | <td className="px-4 py-3 text-muted">What the provider charges</td> | |
| 542 | 542 | <td className="hidden px-4 py-3 tabular-nums sm:table-cell">{gateway?.markupPercent ?? 0}%</td> |
| 1 | − | import { BookOpen, KeyRound } from "lucide-react"; | |
| 1 | + | import { BookOpen, KeyRound, Plug } from "lucide-react"; | |
| 2 | 2 | import { Suspense } from "react"; | |
| 3 | 3 | import { Await, Link, data } from "react-router"; | |
| 4 | 4 | ||
| 5 | 5 | import type { Route } from "./+types/gateway"; | |
| 6 | 6 | import { GatewaySkeleton, GatewayTable } from "../../components/gateway"; | |
| 7 | 7 | import { ButtonLink, CopyLine, EmptyState } from "../../components/ui"; | |
| 8 | − | import { GATEWAY_BASE_URL, GATEWAY_DOCS } from "../../lib/gateway"; | |
| 8 | + | import { GATEWAY_BASE_URL, GATEWAY_DOCS, GATEWAY_OPENAI_BASE_URL } from "../../lib/gateway"; | |
| 9 | 9 | import { page } from "../../lib/meta"; | |
| 10 | 10 | import { billing } from "../../lib/services.server"; | |
| 11 | 11 | import { getViewer, roleIn } from "../../lib/session.server"; | |
| ⋯ | |||
| 37 | 37 | <div className="space-y-6"> | |
| 38 | 38 | <section className="space-y-3"> | |
| 39 | 39 | <p className="max-w-3xl text-sm text-muted"> | |
| 40 | − | Send your own code's model requests in Anthropic's Messages format to this base URL, with one of the workspace's | |
| 41 | − | access tokens that has the <code className="font-mono text-xs">models:write</code> scope as the API key. On g1t's | |
| 42 | − | models each request is charged at the model's price and paid from AI credit; with the workspace's own Anthropic key | |
| 43 | − | under Integrations it is only counted. Prompts and answers are never kept. | |
| 40 | + | Send your own code's model requests to one of these base URLs, in Anthropic's or OpenAI's format, with one of the | |
| 41 | + | workspace's access tokens that has the <code className="font-mono text-xs">models:write</code> scope as the API key. | |
| 42 | + | Any model works in either format: Claude and open models on g1t's account are charged at the model's price and paid | |
| 43 | + | from AI credit; models on the workspace's own providers under Integrations are only counted. Prompts and answers are | |
| 44 | + | never kept. | |
| 44 | 45 | </p> | |
| 45 | − | <div className="max-w-xl"> | |
| 46 | − | <CopyLine text={GATEWAY_BASE_URL} /> | |
| 47 | − | </div> | |
| 46 | + | <dl className="grid max-w-3xl gap-3 md:grid-cols-2"> | |
| 47 | + | <div className="min-w-0"> | |
| 48 | + | <dt className="mb-1.5 text-xs text-faint">Anthropic format</dt> | |
| 49 | + | <dd> | |
| 50 | + | <CopyLine text={GATEWAY_BASE_URL} /> | |
| 51 | + | </dd> | |
| 52 | + | </div> | |
| 53 | + | <div className="min-w-0"> | |
| 54 | + | <dt className="mb-1.5 text-xs text-faint">OpenAI format</dt> | |
| 55 | + | <dd> | |
| 56 | + | <CopyLine text={GATEWAY_OPENAI_BASE_URL} /> | |
| 57 | + | </dd> | |
| 58 | + | </div> | |
| 59 | + | </dl> | |
| 48 | 60 | <div className="flex flex-wrap gap-2"> | |
| 49 | 61 | {owner && ( | |
| 50 | 62 | <ButtonLink variant="quiet" to={`/${slug}/-/tokens`}> | |
| ⋯ | |||
| 52 | 64 | Access tokens | |
| 53 | 65 | </ButtonLink> | |
| 54 | 66 | )} | |
| 67 | + | <ButtonLink variant="quiet" to={`/${slug}/-/integrations`}> | |
| 68 | + | <Plug size={14} /> | |
| 69 | + | Your own providers | |
| 70 | + | </ButtonLink> | |
| 55 | 71 | <ButtonLink variant="quiet" to={GATEWAY_DOCS} reloadDocument> | |
| 56 | 72 | <BookOpen size={14} /> | |
| 57 | 73 | How to use it | |
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
Binary or large file; its contents are not shown.
This change is too large to show in full.