Skip to content

g1t/apps/docs/src/content/docs/guides/ai-gateway.md

504 lines21,652 bytesCodeBlame
1---
2title: AI Gateway
3description: Send your own code's model requests through g1t in Anthropic's or OpenAI's format, with a workspace access token, to Claude, open models or your own providers, with a log of every request.
4---
5
6The AI Gateway takes model requests from your own code and sends them to
7the model. It speaks two formats, and any model works in either:
8
9| Format | Base URL | For |
10| --- | --- | --- |
11| Anthropic's Messages API | `https://models.g1t.sh/anthropic` | Anthropic's SDKs, Claude Code, and anything else that speaks that format |
12| OpenAI's Chat Completions API | `https://models.g1t.sh/openai/v1` | OpenAI's SDKs, and any tool that lets you set an OpenAI-compatible base URL |
13
14In both, the API key is a workspace access token (`g1t_…`) with the
15`models:write` scope.
16
17The model a request names decides where it goes:
18
19- **g1t's models**: Claude on Anthropic, and open models on Workers AI.
20 Each request is charged to the workspace at the model's price and paid
21 from the plan's included usage and
22 [AI credit](/guides/usage-and-billing/#ai-credit). While the gateway is in
23 beta there is no markup.
24- **Your own providers**: an Anthropic key, an OpenAI key, or any endpoint
25 that speaks either API, connected under
26 [Integrations](/guides/models/#connect-a-provider). You choose which models
27 go to each. Those requests are counted and never charged on g1t.
28
29Every request is logged with its format, who served it, its model, tokens,
30cost and status. Prompts and answers are never kept.
31
32## Before you start
33
34You need one of these:
35
36- **The workspace on the g1t plan**, with AI credit or this month's
37 included usage left. See [AI credit](/guides/usage-and-billing/#ai-credit).
38- **One of the workspace's own model providers**, connected under
39 [Integrations](#your-own-providers). Requests for the models it takes need
40 no plan and cost nothing on g1t.
41
42## Make a token
43
44The gateway takes a workspace's own token, so its usage is the workspace's
45and keeps working when the person who set it up leaves. Only owners make
46them.
47
481. Open the workspace's **Settings → Access tokens**.
492. Under **New token**, give it a name, such as `release-notes`. The log
50 shows each request's token by this name.
513. The scopes start on the CI preset. Untick what the code does not need,
52 and under **AI Gateway** tick `models:write`. A token with only
53 `models:write` can send model requests and nothing else.
544. Choose an expiry and select **Create token**. Copy the token now: it is
55 not shown again.
56
57A personal access token is refused, even with `models:write`: the gateway
58has to know which workspace to charge. A token with full access has every
59scope, `models:write` included. See [scopes](/guides/authentication/#scopes).
60
61## Send a request
62
63The token goes in `x-api-key` or in `Authorization: Bearer`, in either
64format. The gateway answers these routes:
65
66| Route | What it does |
67| --- | --- |
68| `POST /anthropic/v1/messages` | A message, streamed (`"stream": true`) or whole. Logged and charged. |
69| `POST /anthropic/v1/messages/count_tokens` | Counts a request's input tokens. Not logged, and costs nothing. For a model that does not speak Anthropic's API, an estimate. |
70| `POST /openai/v1/chat/completions` | A chat completion, streamed or whole. Logged and charged. |
71| `POST /openai/v1/embeddings` | Embeddings, from an embeddings model. Logged and charged by their input tokens. |
72| `GET /openai/v1/models` | The models this workspace can use, with g1t's prices. |
73
74### Anthropic's format
75
76With curl:
77
78```sh
79curl https://models.g1t.sh/anthropic/v1/messages \
80 -H "x-api-key: $G1T_TOKEN" \
81 -H "anthropic-version: 2023-06-01" \
82 -H "content-type: application/json" \
83 -d '{
84 "model": "claude-haiku-5-5",
85 "max_tokens": 1024,
86 "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }]
87 }'
88```
89
90With Anthropic's TypeScript SDK:
91
92```ts
93import Anthropic from "@anthropic-ai/sdk";
94
95const client = new Anthropic({
96 baseURL: "https://models.g1t.sh/anthropic",
97 apiKey: process.env.G1T_TOKEN,
98});
99
100const message = await client.messages.create({
101 model: "claude-haiku-5-5",
102 max_tokens: 1024,
103 messages: [{ role: "user", content: "Write a commit message for: fix the login redirect" }],
104});
105```
106
107With Anthropic's Python SDK:
108
109```python
110import os
111
112import anthropic
113
114client = anthropic.Anthropic(
115 base_url="https://models.g1t.sh/anthropic",
116 api_key=os.environ["G1T_TOKEN"],
117)
118
119message = client.messages.create(
120 model="claude-haiku-5-5",
121 max_tokens=1024,
122 messages=[{"role": "user", "content": "Write a commit message for: fix the login redirect"}],
123)
124```
125
126Request and answer bodies are Anthropic's, and so are streamed events. To
127an open model, such as `workers-ai/@cf/openai/gpt-oss-120b`, the request is
128translated: messages, system prompt, images, tools and tool results,
129`tool_choice`, stop sequences, `output_config.effort` (as
130`reasoning_effort`, `xhigh` and `max` as `high`) and `output_config.format`
131(as a JSON schema). The answer comes back as an Anthropic message, tool
132calls included. Server tools have no counterpart there and are refused on
133g1t's models.
134
135### OpenAI's format
136
137With curl:
138
139```sh
140curl https://models.g1t.sh/openai/v1/chat/completions \
141 -H "Authorization: Bearer $G1T_TOKEN" \
142 -H "content-type: application/json" \
143 -d '{
144 "model": "anthropic/claude-haiku-5-5",
145 "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }]
146 }'
147```
148
149With OpenAI's TypeScript SDK:
150
151```ts
152import OpenAI from "openai";
153
154const client = new OpenAI({
155 baseURL: "https://models.g1t.sh/openai/v1",
156 apiKey: process.env.G1T_TOKEN,
157});
158
159const completion = await client.chat.completions.create({
160 model: "workers-ai/@cf/openai/gpt-oss-120b",
161 messages: [{ role: "user", content: "Label this issue: the login page is blank on Safari" }],
162});
163```
164
165With OpenAI's Python SDK:
166
167```python
168import os
169
170from openai import OpenAI
171
172client = OpenAI(
173 base_url="https://models.g1t.sh/openai/v1",
174 api_key=os.environ["G1T_TOKEN"],
175)
176
177completion = client.chat.completions.create(
178 model="anthropic/claude-sonnet-5-5",
179 messages=[{"role": "user", "content": "Summarize this diff in one sentence."}],
180 stream=True,
181 stream_options={"include_usage": True},
182)
183for chunk in completion:
184 print(chunk.choices[0].delta.content or "" if chunk.choices else "", end="")
185```
186
187Embeddings:
188
189```sh
190curl https://models.g1t.sh/openai/v1/embeddings \
191 -H "Authorization: Bearer $G1T_TOKEN" \
192 -H "content-type: application/json" \
193 -d '{ "model": "workers-ai/@cf/baai/bge-m3", "input": ["fix the login redirect"] }'
194```
195
196What OpenAI's format supports, to any model:
197
198| In the request | |
199| --- | --- |
200| `messages` | `system`, `developer`, `user`, `assistant` and `tool` messages. User content can be text, `image_url` (a URL or a `data:` URL) and, to Claude, `file` with `file_data` (a PDF as a `data:` URL). |
201| `tools`, `tool_choice`, `parallel_tool_calls` | Function tools. To Claude, `required` is `any` and a named function is that tool. |
202| `stream`, `stream_options.include_usage` | Server-sent chunks, ending with `data: [DONE]`. With `include_usage`, a last chunk carries the usage. |
203| `max_tokens`, `max_completion_tokens` | To Claude, 8,192 when neither is given. |
204| `temperature`, `top_p`, `stop`, `user` | As given. Some models refuse sampling settings. |
205| `reasoning_effort` | To Claude, `output_config.effort`; `minimal` and `none` are `low`. |
206| `response_format` | `json_schema` is a JSON schema the answer follows. `json_object` asks Claude for one JSON object. |
207| `thinking` | To Claude, passed as it is, for a caller that sets Anthropic's thinking. |
208
209To Claude, the answer's `usage` counts cached tokens in `prompt_tokens`,
210with `prompt_tokens_details.cached_tokens`; Claude's thinking comes back as
211`reasoning_content`. Claude's thinking blocks must go back with the tool
212calls they led to, so the gateway carries them in the first tool call's
213`id`: send the `id` back unchanged in the assistant message and the `tool`
214message, as OpenAI's SDKs do. `n` above 1 is refused for Claude.
215
216### Claude Code and other tools
217
218Claude Code speaks Anthropic's format. Set two environment variables
219before you start it:
220
221```sh
222export ANTHROPIC_BASE_URL=https://models.g1t.sh/anthropic
223export ANTHROPIC_AUTH_TOKEN=g1t_…
224claude
225```
226
227Claude Code's own small requests go to Claude Haiku 4.5, which the gateway
228offers. To choose the main model, also set `ANTHROPIC_MODEL`, such as
229`claude-opus-5-5`, or `ANTHROPIC_SMALL_FAST_MODEL`, such as
230`claude-haiku-5-5`.
231
232Any other tool that lets you set an OpenAI-compatible base URL and key
233works the same way: give it `https://models.g1t.sh/openai/v1` and the
234token, and a model id from [the models list](#models).
235
236## Models
237
238### Model ids
239
240A request names a model:
241
242| Id | Goes to |
243| --- | --- |
244| `anthropic/claude-sonnet-5-5` | Claude on g1t's account, in either format |
245| `claude-sonnet-5-5` | The same, as Anthropic's API names it |
246| `workers-ai/@cf/openai/gpt-oss-120b` | An open model on g1t's account, in either format |
247| `@cf/openai/gpt-oss-120b` | The same |
248| Any id one of your own providers takes, such as `gpt-5.5` or `ollama/llama3.3` | That provider, with its key. See [your own providers](#your-own-providers). |
249
250Your own providers come first: when one of them takes a model, the request
251goes there, even one that names a model g1t offers. An Anthropic key takes
252`claude-*` unless you choose otherwise, so with one connected, Claude goes to
253your key.
254
255`GET /openai/v1/models` lists what the workspace can use: its own
256providers' models first, then g1t's, starting with the Claude g1t suggests
257starting with (Claude Haiku 5.5 today). g1t adds models as providers
258release them, once their prices are confirmed, so the list and the tables
259below grow over time; a model a provider stops offering is listed until it
260is retired. Each has
261`billed_to` (`workspace` or `g1t`), `connection` (your provider's name) and,
262on g1t's models, `pricing` in dollars per million tokens:
263
264```json
265{
266 "object": "list",
267 "data": [
268 {
269 "id": "anthropic/claude-haiku-5-5",
270 "object": "model",
271 "created": 0,
272 "owned_by": "anthropic",
273 "name": "Claude Haiku 5.5",
274 "kind": "chat",
275 "billed_to": "g1t",
276 "connection": null,
277 "pricing": {
278 "currency": "usd",
279 "input": 0.1,
280 "output": 0.5,
281 "cache_read": 0.01,
282 "cache_write": 0.125,
283 "cache_write_1h": 0.2,
284 "long_prompt": { "above_tokens": 100000, "input": 0.5, "output": 2.5, "cache_read": 0.05, "cache_write": 0.625, "cache_write_1h": 1 }
285 }
286 }
287 ]
288}
289```
290
291### Claude, on Anthropic
292
293Prices are per million tokens, Anthropic's list price. Cache writes are
294five-minute ones; one-hour cache writes (`"ttl": "1h"`) cost twice the
295input price.
296
297| Model | `model` | Input | Output | Cache reads | Cache writes | One-hour cache writes |
298| --- | --- | --- | --- | --- | --- | --- |
299| Claude Haiku 5.5 | `claude-haiku-5-5` | $0.10 | $0.50 | $0.01 | $0.125 | $0.20 |
300| Claude Haiku 5.5, prompts over 100,000 tokens | `claude-haiku-5-5` | $0.50 | $2.50 | $0.05 | $0.625 | $1.00 |
301| Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2.00 | $10.00 | $0.10 | $2.50 | $4.00 |
302| Claude Opus 5.5 | `claude-opus-5-5` | $4.00 | $20.00 | $0.20 | $5.00 | $8.00 |
303| Claude Haiku 4.5 | `claude-haiku-4-5`, `claude-haiku-4-5-20251001` | $1.00 | $5.00 | $0.10 | $1.25 | $2.00 |
304
305Claude Haiku 5.5 is the cheapest Claude and the one to start with. It is
306priced by the prompt's length: a request whose prompt (its input, cache
307read and cache write tokens) is longer than 100,000 tokens is charged
308entirely at the higher prices. It takes effort, like Opus: set
309`output_config.effort` in Anthropic's format, or `reasoning_effort` in
310OpenAI's.
311
312### Open models, on Workers AI
313
314Prices are per million tokens, Cloudflare's list price. Workers AI has no
315prompt-cache price: cached tokens, where a model reports them, cost what
316input does.
317
318| Model | `model` | Input | Output |
319| --- | --- | --- | --- |
320| GLM-5.3 Flash | `workers-ai/@cf/zai-org/glm-5.3-flash` | $0.15 | $0.50 |
321| gpt-oss-20b | `workers-ai/@cf/openai/gpt-oss-20b` | $0.20 | $0.30 |
322| Llama 4 Scout | `workers-ai/@cf/meta/llama-4-scout-17b-16e-instruct` | $0.27 | $0.85 |
323| gpt-oss-120b | `workers-ai/@cf/openai/gpt-oss-120b` | $0.35 | $0.75 |
324| Mistral Small 3.1 | `workers-ai/@cf/mistralai/mistral-small-3.1-24b-instruct` | $0.351 | $0.555 |
325| DeepSeek V4 Flash | `workers-ai/@cf/deepseek-ai/deepseek-v4-flash-0731` | $0.44 | $1.32 |
326| Nemotron 3 120B | `workers-ai/@cf/nvidia/nemotron-3-120b-a12b` | $0.50 | $1.50 |
327| Kimi K2.6 | `workers-ai/@cf/moonshotai/kimi-k2.6` | $0.95 | $4.00 |
328| DeepSeek V4 Pro | `workers-ai/@cf/deepseek-ai/deepseek-v4-pro-0813` | $1.32 | $3.96 |
329| GLM-5.3 | `workers-ai/@cf/zai-org/glm-5.3` | $1.40 | $4.40 |
330
331Embeddings, through `POST /openai/v1/embeddings` only:
332
333| Model | `model` | Input |
334| --- | --- | --- |
335| BGE M3 | `workers-ai/@cf/baai/bge-m3` | $0.012 |
336| BGE Base (English) | `workers-ai/@cf/baai/bge-base-en-v1.5` | $0.067 |
337
338Open models cost much less per call than Claude, and suit one-shot work:
339titles, summaries, labels, triage, embeddings. In a long loop that sends
340the same context every turn, Claude's cache reads close most of that gap.
341
342### What is refused on g1t's models
343
344A request for a model nobody offers is refused with `404` before it
345reaches a provider, and the error names the models offered. On g1t's
346models a request is charged only by its tokens, so what a provider bills
347some other way is refused with `400` for now:
348
349| Not offered on g1t's models yet | In the request |
350| --- | --- |
351| Fast mode | `speed` other than `standard` |
352| Inference in one region | `inference_geo` other than `global` |
353| Server-side fallbacks | `fallbacks` |
354| Server tools, such as web search, web fetch and code execution | In Anthropic's format, a tool whose `type` is not your own (`custom` or none) or a client tool (`bash_…`, `text_editor_…`, `computer_…`, `memory_…`). In OpenAI's, a tool that is not a `function`, or `web_search_options`. |
355| Containers and skills | `container` |
356
357All of them work on your own provider, which bills them. In Claude Code on
358g1t's models, its web search fails for this reason; the rest of Claude Code
359works.
360
361## Your own providers
362
363Connect a model provider under **Integrations**, and choose which models
364your own code's gateway requests send to it. Requests there use its key,
365are counted in the log, and are never charged on g1t. Any provider works:
366
367| Provider | Takes, unless you choose |
368| --- | --- |
369| An Anthropic key, or an Anthropic-compatible endpoint | `claude-*` |
370| An OpenAI key, or any other provider | Nothing until you choose |
371| An OpenAI-compatible endpoint: a self-hosted vLLM or Ollama, LiteLLM, another provider | Nothing until you choose |
372
3731. Open the workspace's **Integrations** and choose a provider under
374 **Model providers**. For your own server, choose **OpenAI-compatible
375 endpoint** or **Anthropic-compatible endpoint** and give its base URL.
3762. Paste its key. It is sealed when saved and never shown again: the page,
377 the API and MCP show only its last four characters.
3783. Under **AI Gateway models**, list the models to send there, separated by
379 spaces:
380
381 | Write | Takes |
382 | --- | --- |
383 | `gpt-5.5` | That model only |
384 | `gpt-*` | Every model whose id starts with `gpt-` |
385 | `ollama/*` | Every model named `ollama/…`, sent without the prefix: `ollama/llama3.3` arrives as `llama3.3` |
386 | `*` | Every model |
387 | Nothing | No gateway requests |
388
3894. Select **Connect**. To change the list or replace the key later, open
390 **Change its AI Gateway models or key** under the provider.
391
392The first provider, in the order they were connected, that takes a model
393gets its requests. Either format reaches either kind of provider: a Claude
394key answers OpenAI-format requests, and an OpenAI-compatible endpoint
395answers Claude Code. Agent runs choose their models under
396[routing](/guides/models/), apart from this list.
397
398From code, connect one with
399[`POST /workspaces/{workspace}/integrations`](/reference/api/integrations/connect-integration/)
400and change it with
401[`PATCH /workspaces/{workspace}/integrations/{id}`](/reference/api/integrations/update-integration/),
402with `config.gateway_models`, or the `workspace` MCP tool's
403`connect_integration` and `update_integration` actions. Both need
404`workspace:admin` and act for an owner. The key is write-only: neither
405returns it.
406
407```sh
408curl -X PATCH https://api.g1t.sh/workspaces/acme/integrations/con_01kpx5c2d8e4f6g0h2j4k6m8n0 \
409 -H "Authorization: Bearer $G1T_ADMIN_TOKEN" \
410 -H "content-type: application/json" \
411 -d '{ "config": { "base_url": "https://gpu.acme.dev/v1", "gateway_models": ["ollama/*"] }, "secret": "…" }'
412```
413
414`config` replaces the provider's settings whole, so send the ones it has
415with the change.
416
417## What it costs
418
419| Where it goes | You pay |
420| --- | --- |
421| g1t's models | Its tokens at the model's price above, with no markup while the gateway is in beta |
422| Your own providers | Nothing on g1t. The provider bills you for the model. |
423
424On g1t's models:
425
426- Each request that used tokens is one line on the statement, under
427 **AI Gateway**, such as *AI Gateway: Claude Sonnet 5.5, 14,352 tokens,
428 token release-notes*. A Claude Haiku 5.5 request over 100,000 prompt
429 tokens says *long-prompt price*.
430- The plan's included usage pays first, then AI credit. Trial credit and
431 g1t's open-source pool never pay for gateway requests.
432- It is not an agent run, so the [agent rate](/guides/usage-and-billing/#the-agent-rate)
433 does not apply.
434- It counts toward the workspace's [spend limit](/guides/usage-and-billing/#your-spend-limit)
435 like any other usage, and shows on **Usage** under the AI Gateway product.
436- A workspace with a 100% discount gets it free through the discount; an
437 enterprise is invoiced for it after use.
438
439## Limits and errors
440
441A request on g1t's models is refused before it reaches the model when:
442
443- The workspace is over its spend limit.
444- It is on the plan, and has no AI credit and none of this month's included
445 usage left. If auto-reload is on, g1t tries it first.
446- It is not on the g1t plan.
447
448Errors are in the format of the route, so SDKs raise their usual errors.
449Anthropic's:
450
451```json
452{ "type": "error", "error": { "type": "billing_error", "message": "The acme workspace is out of AI credit …" } }
453```
454
455OpenAI's:
456
457```json
458{ "error": { "message": "The acme workspace is out of AI credit …", "type": "insufficient_quota", "param": null, "code": "insufficient_quota" } }
459```
460
461| Status | Anthropic's `error.type` | OpenAI's `error.type` (`code`) | Why |
462| --- | --- | --- | --- |
463| `400` | `invalid_request_error` | `invalid_request_error` | The body is not JSON, the model is of the wrong kind, the request asks for something [not offered on g1t's models yet](#what-is-refused-on-g1ts-models), or it cannot be said to the model (such as `n` above 1 to Claude). |
464| `401` | `authentication_error` | `authentication_error` (`invalid_api_key`) | The token is unknown, expired or deleted, or your provider refused its key. |
465| `402` | `billing_error` | `insufficient_quota` (`insufficient_quota`) | Out of AI credit, over the spend limit, or not on the plan. The message says what an owner can do. |
466| `403` | `permission_error` | `permission_error` | Not a workspace's token, or it lacks `models:write`. |
467| `404` | `not_found_error` | `invalid_request_error` (`model_not_found`) | No provider offers the model, or a route the gateway does not answer. |
468
469An error from the model provider, such as `429` or `529`, comes back with
470its status and message, in the route's format. A provider's key never
471appears in an error or the log, even when the provider quotes it. Refused
472and failed requests are logged with their status and why, and cost
473nothing. Every answer carries `x-g1t-request-id`, the request's id in the
474log.
475
476A deleted token, a provider added or changed under Integrations, or AI
477credit just bought takes effect within about ten seconds.
478
479## See every request
480
481The **AI Gateway** page lists the workspace's requests, newest first. Open
482it from the link under **Usage**, or at `g1t.sh/<workspace>/-/gateway`.
483Every member can see it.
484
485| Column | |
486| --- | --- |
487| Time | When it was sent. Hover for the exact time, how long it took and whether it streamed. |
488| Model | The model it named on g1t's models, or the one that answered on your own provider; below it, the format it was sent in. |
489| Served by | g1t's account and the provider (*g1t · Anthropic*, *g1t · Workers AI*), or your provider by name. *None* when it was refused first. |
490| Input, Output | Its tokens by kind. Hover Input for all of them. |
491| Cache | Cache reads, then cache writes. Hover for how many writes were to the one-hour cache. |
492| Cost | What it was charged, before included usage and AI credit paid for it, or **Not charged** on your own provider. |
493| Status | The status it was answered with. Hover a refusal or failure for why. |
494| Token | The name of the token that sent it. |
495
496Requests are kept 30 days.
497
498From code, list them with
499[`GET /workspaces/{workspace}/gateway/requests`](/reference/api/billing/list-gateway-requests/),
500or the `billing` MCP tool's
501[`gateway_requests`](/reference/mcp/#billing) action. Each request has
502`format`, `provider`, `connection`, `model` and its tokens, with
503`cache_write_hour` for one-hour cache writes. Both need `models:read`,
504which the Read only and Agent presets include.