| 1 | --- |
| 2 | title: AI Gateway |
| 3 | description: Send your own code's model requests through g1t in Anthropic's or OpenAI's format, with a workspace access token, to Claude, open models or your own providers, with a log of every request. |
| 4 | --- |
| 5 | |
| 6 | The AI Gateway takes model requests from your own code and sends them to |
| 7 | the model. It speaks two formats, and any model works in either: |
| 8 | |
| 9 | | Format | Base URL | For | |
| 10 | | --- | --- | --- | |
| 11 | | Anthropic's Messages API | `https://models.g1t.sh/anthropic` | Anthropic's SDKs, Claude Code, and anything else that speaks that format | |
| 12 | | OpenAI's Chat Completions API | `https://models.g1t.sh/openai/v1` | OpenAI's SDKs, and any tool that lets you set an OpenAI-compatible base URL | |
| 13 | |
| 14 | In both, the API key is a workspace access token (`g1t_…`) with the |
| 15 | `models:write` scope. |
| 16 | |
| 17 | The model a request names decides where it goes: |
| 18 | |
| 19 | - **g1t's models**: Claude on Anthropic, and open models on Workers AI. |
| 20 | Each request is charged to the workspace at the model's price and paid |
| 21 | from the plan's included usage and |
| 22 | [AI credit](/guides/usage-and-billing/#ai-credit). While the gateway is in |
| 23 | beta there is no markup. |
| 24 | - **Your own providers**: an Anthropic key, an OpenAI key, or any endpoint |
| 25 | that speaks either API, connected under |
| 26 | [Integrations](/guides/models/#connect-a-provider). You choose which models |
| 27 | go to each. Those requests are counted and never charged on g1t. |
| 28 | |
| 29 | Every request is logged with its format, who served it, its model, tokens, |
| 30 | cost and status. Prompts and answers are never kept. |
| 31 | |
| 32 | ## Before you start |
| 33 | |
| 34 | You need one of these: |
| 35 | |
| 36 | - **The workspace on the g1t plan**, with AI credit or this month's |
| 37 | included usage left. See [AI credit](/guides/usage-and-billing/#ai-credit). |
| 38 | - **One of the workspace's own model providers**, connected under |
| 39 | [Integrations](#your-own-providers). Requests for the models it takes need |
| 40 | no plan and cost nothing on g1t. |
| 41 | |
| 42 | ## Make a token |
| 43 | |
| 44 | The gateway takes a workspace's own token, so its usage is the workspace's |
| 45 | and keeps working when the person who set it up leaves. Only owners make |
| 46 | them. |
| 47 | |
| 48 | 1. Open the workspace's **Settings → Access tokens**. |
| 49 | 2. Under **New token**, give it a name, such as `release-notes`. The log |
| 50 | shows each request's token by this name. |
| 51 | 3. The scopes start on the CI preset. Untick what the code does not need, |
| 52 | and under **AI Gateway** tick `models:write`. A token with only |
| 53 | `models:write` can send model requests and nothing else. |
| 54 | 4. Choose an expiry and select **Create token**. Copy the token now: it is |
| 55 | not shown again. |
| 56 | |
| 57 | A personal access token is refused, even with `models:write`: the gateway |
| 58 | has to know which workspace to charge. A token with full access has every |
| 59 | scope, `models:write` included. See [scopes](/guides/authentication/#scopes). |
| 60 | |
| 61 | ## Send a request |
| 62 | |
| 63 | The token goes in `x-api-key` or in `Authorization: Bearer`, in either |
| 64 | format. The gateway answers these routes: |
| 65 | |
| 66 | | Route | What it does | |
| 67 | | --- | --- | |
| 68 | | `POST /anthropic/v1/messages` | A message, streamed (`"stream": true`) or whole. Logged and charged. | |
| 69 | | `POST /anthropic/v1/messages/count_tokens` | Counts a request's input tokens. Not logged, and costs nothing. For a model that does not speak Anthropic's API, an estimate. | |
| 70 | | `POST /openai/v1/chat/completions` | A chat completion, streamed or whole. Logged and charged. | |
| 71 | | `POST /openai/v1/embeddings` | Embeddings, from an embeddings model. Logged and charged by their input tokens. | |
| 72 | | `GET /openai/v1/models` | The models this workspace can use, with g1t's prices. | |
| 73 | |
| 74 | ### Anthropic's format |
| 75 | |
| 76 | With curl: |
| 77 | |
| 78 | ```sh |
| 79 | curl https://models.g1t.sh/anthropic/v1/messages \ |
| 80 | -H "x-api-key: $G1T_TOKEN" \ |
| 81 | -H "anthropic-version: 2023-06-01" \ |
| 82 | -H "content-type: application/json" \ |
| 83 | -d '{ |
| 84 | "model": "claude-haiku-5-5", |
| 85 | "max_tokens": 1024, |
| 86 | "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }] |
| 87 | }' |
| 88 | ``` |
| 89 | |
| 90 | With Anthropic's TypeScript SDK: |
| 91 | |
| 92 | ```ts |
| 93 | import Anthropic from "@anthropic-ai/sdk"; |
| 94 | |
| 95 | const client = new Anthropic({ |
| 96 | baseURL: "https://models.g1t.sh/anthropic", |
| 97 | apiKey: process.env.G1T_TOKEN, |
| 98 | }); |
| 99 | |
| 100 | const message = await client.messages.create({ |
| 101 | model: "claude-haiku-5-5", |
| 102 | max_tokens: 1024, |
| 103 | messages: [{ role: "user", content: "Write a commit message for: fix the login redirect" }], |
| 104 | }); |
| 105 | ``` |
| 106 | |
| 107 | With Anthropic's Python SDK: |
| 108 | |
| 109 | ```python |
| 110 | import os |
| 111 | |
| 112 | import anthropic |
| 113 | |
| 114 | client = anthropic.Anthropic( |
| 115 | base_url="https://models.g1t.sh/anthropic", |
| 116 | api_key=os.environ["G1T_TOKEN"], |
| 117 | ) |
| 118 | |
| 119 | message = client.messages.create( |
| 120 | model="claude-haiku-5-5", |
| 121 | max_tokens=1024, |
| 122 | messages=[{"role": "user", "content": "Write a commit message for: fix the login redirect"}], |
| 123 | ) |
| 124 | ``` |
| 125 | |
| 126 | Request and answer bodies are Anthropic's, and so are streamed events. To |
| 127 | an open model, such as `workers-ai/@cf/openai/gpt-oss-120b`, the request is |
| 128 | translated: messages, system prompt, images, tools and tool results, |
| 129 | `tool_choice`, stop sequences, `output_config.effort` (as |
| 130 | `reasoning_effort`, `xhigh` and `max` as `high`) and `output_config.format` |
| 131 | (as a JSON schema). The answer comes back as an Anthropic message, tool |
| 132 | calls included. Server tools have no counterpart there and are refused on |
| 133 | g1t's models. |
| 134 | |
| 135 | ### OpenAI's format |
| 136 | |
| 137 | With curl: |
| 138 | |
| 139 | ```sh |
| 140 | curl https://models.g1t.sh/openai/v1/chat/completions \ |
| 141 | -H "Authorization: Bearer $G1T_TOKEN" \ |
| 142 | -H "content-type: application/json" \ |
| 143 | -d '{ |
| 144 | "model": "anthropic/claude-haiku-5-5", |
| 145 | "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }] |
| 146 | }' |
| 147 | ``` |
| 148 | |
| 149 | With OpenAI's TypeScript SDK: |
| 150 | |
| 151 | ```ts |
| 152 | import OpenAI from "openai"; |
| 153 | |
| 154 | const client = new OpenAI({ |
| 155 | baseURL: "https://models.g1t.sh/openai/v1", |
| 156 | apiKey: process.env.G1T_TOKEN, |
| 157 | }); |
| 158 | |
| 159 | const completion = await client.chat.completions.create({ |
| 160 | model: "workers-ai/@cf/openai/gpt-oss-120b", |
| 161 | messages: [{ role: "user", content: "Label this issue: the login page is blank on Safari" }], |
| 162 | }); |
| 163 | ``` |
| 164 | |
| 165 | With OpenAI's Python SDK: |
| 166 | |
| 167 | ```python |
| 168 | import os |
| 169 | |
| 170 | from openai import OpenAI |
| 171 | |
| 172 | client = OpenAI( |
| 173 | base_url="https://models.g1t.sh/openai/v1", |
| 174 | api_key=os.environ["G1T_TOKEN"], |
| 175 | ) |
| 176 | |
| 177 | completion = client.chat.completions.create( |
| 178 | model="anthropic/claude-sonnet-5-5", |
| 179 | messages=[{"role": "user", "content": "Summarize this diff in one sentence."}], |
| 180 | stream=True, |
| 181 | stream_options={"include_usage": True}, |
| 182 | ) |
| 183 | for chunk in completion: |
| 184 | print(chunk.choices[0].delta.content or "" if chunk.choices else "", end="") |
| 185 | ``` |
| 186 | |
| 187 | Embeddings: |
| 188 | |
| 189 | ```sh |
| 190 | curl https://models.g1t.sh/openai/v1/embeddings \ |
| 191 | -H "Authorization: Bearer $G1T_TOKEN" \ |
| 192 | -H "content-type: application/json" \ |
| 193 | -d '{ "model": "workers-ai/@cf/baai/bge-m3", "input": ["fix the login redirect"] }' |
| 194 | ``` |
| 195 | |
| 196 | What OpenAI's format supports, to any model: |
| 197 | |
| 198 | | In the request | | |
| 199 | | --- | --- | |
| 200 | | `messages` | `system`, `developer`, `user`, `assistant` and `tool` messages. User content can be text, `image_url` (a URL or a `data:` URL) and, to Claude, `file` with `file_data` (a PDF as a `data:` URL). | |
| 201 | | `tools`, `tool_choice`, `parallel_tool_calls` | Function tools. To Claude, `required` is `any` and a named function is that tool. | |
| 202 | | `stream`, `stream_options.include_usage` | Server-sent chunks, ending with `data: [DONE]`. With `include_usage`, a last chunk carries the usage. | |
| 203 | | `max_tokens`, `max_completion_tokens` | To Claude, 8,192 when neither is given. | |
| 204 | | `temperature`, `top_p`, `stop`, `user` | As given. Some models refuse sampling settings. | |
| 205 | | `reasoning_effort` | To Claude, `output_config.effort`; `minimal` and `none` are `low`. | |
| 206 | | `response_format` | `json_schema` is a JSON schema the answer follows. `json_object` asks Claude for one JSON object. | |
| 207 | | `thinking` | To Claude, passed as it is, for a caller that sets Anthropic's thinking. | |
| 208 | |
| 209 | To Claude, the answer's `usage` counts cached tokens in `prompt_tokens`, |
| 210 | with `prompt_tokens_details.cached_tokens`; Claude's thinking comes back as |
| 211 | `reasoning_content`. Claude's thinking blocks must go back with the tool |
| 212 | calls they led to, so the gateway carries them in the first tool call's |
| 213 | `id`: send the `id` back unchanged in the assistant message and the `tool` |
| 214 | message, as OpenAI's SDKs do. `n` above 1 is refused for Claude. |
| 215 | |
| 216 | ### Claude Code and other tools |
| 217 | |
| 218 | Claude Code speaks Anthropic's format. Set two environment variables |
| 219 | before you start it: |
| 220 | |
| 221 | ```sh |
| 222 | export ANTHROPIC_BASE_URL=https://models.g1t.sh/anthropic |
| 223 | export ANTHROPIC_AUTH_TOKEN=g1t_… |
| 224 | claude |
| 225 | ``` |
| 226 | |
| 227 | Claude Code's own small requests go to Claude Haiku 4.5, which the gateway |
| 228 | offers. To choose the main model, also set `ANTHROPIC_MODEL`, such as |
| 229 | `claude-opus-5-5`, or `ANTHROPIC_SMALL_FAST_MODEL`, such as |
| 230 | `claude-haiku-5-5`. |
| 231 | |
| 232 | Any other tool that lets you set an OpenAI-compatible base URL and key |
| 233 | works the same way: give it `https://models.g1t.sh/openai/v1` and the |
| 234 | token, and a model id from [the models list](#models). |
| 235 | |
| 236 | ## Models |
| 237 | |
| 238 | ### Model ids |
| 239 | |
| 240 | A request names a model: |
| 241 | |
| 242 | | Id | Goes to | |
| 243 | | --- | --- | |
| 244 | | `anthropic/claude-sonnet-5-5` | Claude on g1t's account, in either format | |
| 245 | | `claude-sonnet-5-5` | The same, as Anthropic's API names it | |
| 246 | | `workers-ai/@cf/openai/gpt-oss-120b` | An open model on g1t's account, in either format | |
| 247 | | `@cf/openai/gpt-oss-120b` | The same | |
| 248 | | Any id one of your own providers takes, such as `gpt-5.5` or `ollama/llama3.3` | That provider, with its key. See [your own providers](#your-own-providers). | |
| 249 | |
| 250 | Your own providers come first: when one of them takes a model, the request |
| 251 | goes there, even one that names a model g1t offers. An Anthropic key takes |
| 252 | `claude-*` unless you choose otherwise, so with one connected, Claude goes to |
| 253 | your key. |
| 254 | |
| 255 | `GET /openai/v1/models` lists what the workspace can use: its own |
| 256 | providers' models first, then g1t's, starting with the Claude g1t suggests |
| 257 | starting with (Claude Haiku 5.5 today). g1t adds models as providers |
| 258 | release them, once their prices are confirmed, so the list and the tables |
| 259 | below grow over time; a model a provider stops offering is listed until it |
| 260 | is retired. Each has |
| 261 | `billed_to` (`workspace` or `g1t`), `connection` (your provider's name) and, |
| 262 | on g1t's models, `pricing` in dollars per million tokens: |
| 263 | |
| 264 | ```json |
| 265 | { |
| 266 | "object": "list", |
| 267 | "data": [ |
| 268 | { |
| 269 | "id": "anthropic/claude-haiku-5-5", |
| 270 | "object": "model", |
| 271 | "created": 0, |
| 272 | "owned_by": "anthropic", |
| 273 | "name": "Claude Haiku 5.5", |
| 274 | "kind": "chat", |
| 275 | "billed_to": "g1t", |
| 276 | "connection": null, |
| 277 | "pricing": { |
| 278 | "currency": "usd", |
| 279 | "input": 0.1, |
| 280 | "output": 0.5, |
| 281 | "cache_read": 0.01, |
| 282 | "cache_write": 0.125, |
| 283 | "cache_write_1h": 0.2, |
| 284 | "long_prompt": { "above_tokens": 100000, "input": 0.5, "output": 2.5, "cache_read": 0.05, "cache_write": 0.625, "cache_write_1h": 1 } |
| 285 | } |
| 286 | } |
| 287 | ] |
| 288 | } |
| 289 | ``` |
| 290 | |
| 291 | ### Claude, on Anthropic |
| 292 | |
| 293 | Prices are per million tokens, Anthropic's list price. Cache writes are |
| 294 | five-minute ones; one-hour cache writes (`"ttl": "1h"`) cost twice the |
| 295 | input price. |
| 296 | |
| 297 | | Model | `model` | Input | Output | Cache reads | Cache writes | One-hour cache writes | |
| 298 | | --- | --- | --- | --- | --- | --- | --- | |
| 299 | | Claude Haiku 5.5 | `claude-haiku-5-5` | $0.10 | $0.50 | $0.01 | $0.125 | $0.20 | |
| 300 | | Claude Haiku 5.5, prompts over 100,000 tokens | `claude-haiku-5-5` | $0.50 | $2.50 | $0.05 | $0.625 | $1.00 | |
| 301 | | Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2.00 | $10.00 | $0.10 | $2.50 | $4.00 | |
| 302 | | Claude Opus 5.5 | `claude-opus-5-5` | $4.00 | $20.00 | $0.20 | $5.00 | $8.00 | |
| 303 | | Claude Haiku 4.5 | `claude-haiku-4-5`, `claude-haiku-4-5-20251001` | $1.00 | $5.00 | $0.10 | $1.25 | $2.00 | |
| 304 | |
| 305 | Claude Haiku 5.5 is the cheapest Claude and the one to start with. It is |
| 306 | priced by the prompt's length: a request whose prompt (its input, cache |
| 307 | read and cache write tokens) is longer than 100,000 tokens is charged |
| 308 | entirely at the higher prices. It takes effort, like Opus: set |
| 309 | `output_config.effort` in Anthropic's format, or `reasoning_effort` in |
| 310 | OpenAI's. |
| 311 | |
| 312 | ### Open models, on Workers AI |
| 313 | |
| 314 | Prices are per million tokens, Cloudflare's list price. Workers AI has no |
| 315 | prompt-cache price: cached tokens, where a model reports them, cost what |
| 316 | input does. |
| 317 | |
| 318 | | Model | `model` | Input | Output | |
| 319 | | --- | --- | --- | --- | |
| 320 | | GLM-5.3 Flash | `workers-ai/@cf/zai-org/glm-5.3-flash` | $0.15 | $0.50 | |
| 321 | | gpt-oss-20b | `workers-ai/@cf/openai/gpt-oss-20b` | $0.20 | $0.30 | |
| 322 | | Llama 4 Scout | `workers-ai/@cf/meta/llama-4-scout-17b-16e-instruct` | $0.27 | $0.85 | |
| 323 | | gpt-oss-120b | `workers-ai/@cf/openai/gpt-oss-120b` | $0.35 | $0.75 | |
| 324 | | Mistral Small 3.1 | `workers-ai/@cf/mistralai/mistral-small-3.1-24b-instruct` | $0.351 | $0.555 | |
| 325 | | DeepSeek V4 Flash | `workers-ai/@cf/deepseek-ai/deepseek-v4-flash-0731` | $0.44 | $1.32 | |
| 326 | | Nemotron 3 120B | `workers-ai/@cf/nvidia/nemotron-3-120b-a12b` | $0.50 | $1.50 | |
| 327 | | Kimi K2.6 | `workers-ai/@cf/moonshotai/kimi-k2.6` | $0.95 | $4.00 | |
| 328 | | DeepSeek V4 Pro | `workers-ai/@cf/deepseek-ai/deepseek-v4-pro-0813` | $1.32 | $3.96 | |
| 329 | | GLM-5.3 | `workers-ai/@cf/zai-org/glm-5.3` | $1.40 | $4.40 | |
| 330 | |
| 331 | Embeddings, through `POST /openai/v1/embeddings` only: |
| 332 | |
| 333 | | Model | `model` | Input | |
| 334 | | --- | --- | --- | |
| 335 | | BGE M3 | `workers-ai/@cf/baai/bge-m3` | $0.012 | |
| 336 | | BGE Base (English) | `workers-ai/@cf/baai/bge-base-en-v1.5` | $0.067 | |
| 337 | |
| 338 | Open models cost much less per call than Claude, and suit one-shot work: |
| 339 | titles, summaries, labels, triage, embeddings. In a long loop that sends |
| 340 | the same context every turn, Claude's cache reads close most of that gap. |
| 341 | |
| 342 | ### What is refused on g1t's models |
| 343 | |
| 344 | A request for a model nobody offers is refused with `404` before it |
| 345 | reaches a provider, and the error names the models offered. On g1t's |
| 346 | models a request is charged only by its tokens, so what a provider bills |
| 347 | some other way is refused with `400` for now: |
| 348 | |
| 349 | | Not offered on g1t's models yet | In the request | |
| 350 | | --- | --- | |
| 351 | | Fast mode | `speed` other than `standard` | |
| 352 | | Inference in one region | `inference_geo` other than `global` | |
| 353 | | Server-side fallbacks | `fallbacks` | |
| 354 | | Server tools, such as web search, web fetch and code execution | In Anthropic's format, a tool whose `type` is not your own (`custom` or none) or a client tool (`bash_…`, `text_editor_…`, `computer_…`, `memory_…`). In OpenAI's, a tool that is not a `function`, or `web_search_options`. | |
| 355 | | Containers and skills | `container` | |
| 356 | |
| 357 | All of them work on your own provider, which bills them. In Claude Code on |
| 358 | g1t's models, its web search fails for this reason; the rest of Claude Code |
| 359 | works. |
| 360 | |
| 361 | ## Your own providers |
| 362 | |
| 363 | Connect a model provider under **Integrations**, and choose which models |
| 364 | your own code's gateway requests send to it. Requests there use its key, |
| 365 | are counted in the log, and are never charged on g1t. Any provider works: |
| 366 | |
| 367 | | Provider | Takes, unless you choose | |
| 368 | | --- | --- | |
| 369 | | An Anthropic key, or an Anthropic-compatible endpoint | `claude-*` | |
| 370 | | An OpenAI key, or any other provider | Nothing until you choose | |
| 371 | | An OpenAI-compatible endpoint: a self-hosted vLLM or Ollama, LiteLLM, another provider | Nothing until you choose | |
| 372 | |
| 373 | 1. Open the workspace's **Integrations** and choose a provider under |
| 374 | **Model providers**. For your own server, choose **OpenAI-compatible |
| 375 | endpoint** or **Anthropic-compatible endpoint** and give its base URL. |
| 376 | 2. Paste its key. It is sealed when saved and never shown again: the page, |
| 377 | the API and MCP show only its last four characters. |
| 378 | 3. Under **AI Gateway models**, list the models to send there, separated by |
| 379 | spaces: |
| 380 | |
| 381 | | Write | Takes | |
| 382 | | --- | --- | |
| 383 | | `gpt-5.5` | That model only | |
| 384 | | `gpt-*` | Every model whose id starts with `gpt-` | |
| 385 | | `ollama/*` | Every model named `ollama/…`, sent without the prefix: `ollama/llama3.3` arrives as `llama3.3` | |
| 386 | | `*` | Every model | |
| 387 | | Nothing | No gateway requests | |
| 388 | |
| 389 | 4. Select **Connect**. To change the list or replace the key later, open |
| 390 | **Change its AI Gateway models or key** under the provider. |
| 391 | |
| 392 | The first provider, in the order they were connected, that takes a model |
| 393 | gets its requests. Either format reaches either kind of provider: a Claude |
| 394 | key answers OpenAI-format requests, and an OpenAI-compatible endpoint |
| 395 | answers Claude Code. Agent runs choose their models under |
| 396 | [routing](/guides/models/), apart from this list. |
| 397 | |
| 398 | From code, connect one with |
| 399 | [`POST /workspaces/{workspace}/integrations`](/reference/api/integrations/connect-integration/) |
| 400 | and change it with |
| 401 | [`PATCH /workspaces/{workspace}/integrations/{id}`](/reference/api/integrations/update-integration/), |
| 402 | with `config.gateway_models`, or the `workspace` MCP tool's |
| 403 | `connect_integration` and `update_integration` actions. Both need |
| 404 | `workspace:admin` and act for an owner. The key is write-only: neither |
| 405 | returns it. |
| 406 | |
| 407 | ```sh |
| 408 | curl -X PATCH https://api.g1t.sh/workspaces/acme/integrations/con_01kpx5c2d8e4f6g0h2j4k6m8n0 \ |
| 409 | -H "Authorization: Bearer $G1T_ADMIN_TOKEN" \ |
| 410 | -H "content-type: application/json" \ |
| 411 | -d '{ "config": { "base_url": "https://gpu.acme.dev/v1", "gateway_models": ["ollama/*"] }, "secret": "…" }' |
| 412 | ``` |
| 413 | |
| 414 | `config` replaces the provider's settings whole, so send the ones it has |
| 415 | with the change. |
| 416 | |
| 417 | ## What it costs |
| 418 | |
| 419 | | Where it goes | You pay | |
| 420 | | --- | --- | |
| 421 | | g1t's models | Its tokens at the model's price above, with no markup while the gateway is in beta | |
| 422 | | Your own providers | Nothing on g1t. The provider bills you for the model. | |
| 423 | |
| 424 | On g1t's models: |
| 425 | |
| 426 | - Each request that used tokens is one line on the statement, under |
| 427 | **AI Gateway**, such as *AI Gateway: Claude Sonnet 5.5, 14,352 tokens, |
| 428 | token release-notes*. A Claude Haiku 5.5 request over 100,000 prompt |
| 429 | tokens says *long-prompt price*. |
| 430 | - The plan's included usage pays first, then AI credit. Trial credit and |
| 431 | g1t's open-source pool never pay for gateway requests. |
| 432 | - It is not an agent run, so the [agent rate](/guides/usage-and-billing/#the-agent-rate) |
| 433 | does not apply. |
| 434 | - It counts toward the workspace's [spend limit](/guides/usage-and-billing/#your-spend-limit) |
| 435 | like any other usage, and shows on **Usage** under the AI Gateway product. |
| 436 | - A workspace with a 100% discount gets it free through the discount; an |
| 437 | enterprise is invoiced for it after use. |
| 438 | |
| 439 | ## Limits and errors |
| 440 | |
| 441 | A request on g1t's models is refused before it reaches the model when: |
| 442 | |
| 443 | - The workspace is over its spend limit. |
| 444 | - It is on the plan, and has no AI credit and none of this month's included |
| 445 | usage left. If auto-reload is on, g1t tries it first. |
| 446 | - It is not on the g1t plan. |
| 447 | |
| 448 | Errors are in the format of the route, so SDKs raise their usual errors. |
| 449 | Anthropic's: |
| 450 | |
| 451 | ```json |
| 452 | { "type": "error", "error": { "type": "billing_error", "message": "The acme workspace is out of AI credit …" } } |
| 453 | ``` |
| 454 | |
| 455 | OpenAI's: |
| 456 | |
| 457 | ```json |
| 458 | { "error": { "message": "The acme workspace is out of AI credit …", "type": "insufficient_quota", "param": null, "code": "insufficient_quota" } } |
| 459 | ``` |
| 460 | |
| 461 | | Status | Anthropic's `error.type` | OpenAI's `error.type` (`code`) | Why | |
| 462 | | --- | --- | --- | --- | |
| 463 | | `400` | `invalid_request_error` | `invalid_request_error` | The body is not JSON, the model is of the wrong kind, the request asks for something [not offered on g1t's models yet](#what-is-refused-on-g1ts-models), or it cannot be said to the model (such as `n` above 1 to Claude). | |
| 464 | | `401` | `authentication_error` | `authentication_error` (`invalid_api_key`) | The token is unknown, expired or deleted, or your provider refused its key. | |
| 465 | | `402` | `billing_error` | `insufficient_quota` (`insufficient_quota`) | Out of AI credit, over the spend limit, or not on the plan. The message says what an owner can do. | |
| 466 | | `403` | `permission_error` | `permission_error` | Not a workspace's token, or it lacks `models:write`. | |
| 467 | | `404` | `not_found_error` | `invalid_request_error` (`model_not_found`) | No provider offers the model, or a route the gateway does not answer. | |
| 468 | |
| 469 | An error from the model provider, such as `429` or `529`, comes back with |
| 470 | its status and message, in the route's format. A provider's key never |
| 471 | appears in an error or the log, even when the provider quotes it. Refused |
| 472 | and failed requests are logged with their status and why, and cost |
| 473 | nothing. Every answer carries `x-g1t-request-id`, the request's id in the |
| 474 | log. |
| 475 | |
| 476 | A deleted token, a provider added or changed under Integrations, or AI |
| 477 | credit just bought takes effect within about ten seconds. |
| 478 | |
| 479 | ## See every request |
| 480 | |
| 481 | The **AI Gateway** page lists the workspace's requests, newest first. Open |
| 482 | it from the link under **Usage**, or at `g1t.sh/<workspace>/-/gateway`. |
| 483 | Every member can see it. |
| 484 | |
| 485 | | Column | | |
| 486 | | --- | --- | |
| 487 | | Time | When it was sent. Hover for the exact time, how long it took and whether it streamed. | |
| 488 | | Model | The model it named on g1t's models, or the one that answered on your own provider; below it, the format it was sent in. | |
| 489 | | Served by | g1t's account and the provider (*g1t · Anthropic*, *g1t · Workers AI*), or your provider by name. *None* when it was refused first. | |
| 490 | | Input, Output | Its tokens by kind. Hover Input for all of them. | |
| 491 | | Cache | Cache reads, then cache writes. Hover for how many writes were to the one-hour cache. | |
| 492 | | Cost | What it was charged, before included usage and AI credit paid for it, or **Not charged** on your own provider. | |
| 493 | | Status | The status it was answered with. Hover a refusal or failure for why. | |
| 494 | | Token | The name of the token that sent it. | |
| 495 | |
| 496 | Requests are kept 30 days. |
| 497 | |
| 498 | From code, list them with |
| 499 | [`GET /workspaces/{workspace}/gateway/requests`](/reference/api/billing/list-gateway-requests/), |
| 500 | or the `billing` MCP tool's |
| 501 | [`gateway_requests`](/reference/mcp/#billing) action. Each request has |
| 502 | `format`, `provider`, `connection`, `model` and its tokens, with |
| 503 | `cache_write_hour` for one-hour cache writes. Both need `models:read`, |
| 504 | which the Read only and Agent presets include. |