| 1 | --- |
| 2 | title: Model providers |
| 3 | description: Let Auto choose the model for each job, or choose it yourself; connect Anthropic, OpenAI, Gemini or any compatible endpoint, and pay for models where you choose. |
| 4 | --- |
| 5 | |
| 6 | Each workspace decides where its agents' model spend goes: |
| 7 | |
| 8 | - **g1t's hosted models.** **Auto** chooses the model for each job, g1t pays |
| 9 | the provider, and your workspace is charged the provider's price, with no |
| 10 | markup, plus the [agent rate](/guides/usage-and-billing/#the-agent-rate). |
| 11 | The plan's included usage and [the trial](/guides/usage-and-billing/#the-trial) |
| 12 | pay for it first. While payments are in test mode, they are open only |
| 13 | to a few invited workspaces, g1t's own among them; a card check or a |
| 14 | trial does not open them. Every other workspace connects its own |
| 15 | provider, and an agent assigned without one is refused with a message |
| 16 | that says so. Once payments go live, they are open to all. |
| 17 | - **Your own providers.** Connect as many as you use, then choose, for each |
| 18 | kind of work, which provider and model it runs on. Each provider bills |
| 19 | you for the model directly. Open to every workspace now. |
| 20 | |
| 21 | To call models from your own code with a workspace token, in Anthropic's |
| 22 | or OpenAI's format, paid from the same AI credit or sent to these same |
| 23 | providers, use the [AI Gateway](/guides/ai-gateway/). Each provider's |
| 24 | **AI Gateway models** say which of those requests go to it; see |
| 25 | [your own providers](/guides/ai-gateway/#your-own-providers). |
| 26 | |
| 27 | ## Auto |
| 28 | |
| 29 | On g1t's models you do not have to pick a model. **Auto**, the default, |
| 30 | sends each job to the least costly model that can do it, from three tiers: |
| 31 | **fast** (Claude Haiku 5.5 today), **standard** (Claude Sonnet 5.5) and |
| 32 | **most capable** (Claude Opus 5.5). It decides by the kind of job, the |
| 33 | size of the change it reads, the issue's labels, whether the last attempt |
| 34 | at the same work failed, and what has worked in the repository before: |
| 35 | |
| 36 | - Catching up, answering a question, planning, and reviewing a small |
| 37 | change that touches no sensitive path start on the fast model. |
| 38 | - Making and revising changes, and most reviews, start on the standard |
| 39 | model. |
| 40 | - Planning runs at high effort (the model thinks longer before it |
| 41 | answers), answering at medium and catching up at low, on models that |
| 42 | take an effort level. |
| 43 | - A review of a very large change, work on an issue labelled |
| 44 | `architecture`, and work that failed twice in a row go to the most |
| 45 | capable model. One failure moves the next attempt up one tier. |
| 46 | - When the cheaper model finished nearly all of a repository's recent runs |
| 47 | of the same kind, Auto moves that work down a tier there; when a model |
| 48 | keeps failing, up. |
| 49 | |
| 50 | The models behind the tiers are today's. g1t keeps up with new models |
| 51 | as providers release them: it checks for new ones every day, and when g1t |
| 52 | moves a tier to a new model, your runs use it within a minute, with |
| 53 | nothing for you to change. Nobody picks a model; Auto keeps choosing by the |
| 54 | work. A model a provider retires is never used again: the next model for |
| 55 | that tier runs instead, and the run says so. |
| 56 | |
| 57 | Every run says which model it used and why, in one line on its run and in |
| 58 | its pull request's session, such as *Used a fast model (Claude Haiku 5.5): |
| 59 | small change, 3 files and 80 lines.* The full rules are in |
| 60 | [which model runs](/guides/working-with-g1t/#which-model-runs). |
| 61 | |
| 62 | ## Providers |
| 63 | |
| 64 | Labs and platforms need only a key; g1t knows where they are. |
| 65 | |
| 66 | | | Provider | What you give | |
| 67 | | --- | --- | --- | |
| 68 | | Labs | **Anthropic** | An API key | |
| 69 | | | **OpenAI** | An API key | |
| 70 | | | **Google Gemini** | An API key from Google AI Studio | |
| 71 | | | **xAI** (Grok) | An API key | |
| 72 | | | **Mistral** | An API key | |
| 73 | | | **DeepSeek** | An API key | |
| 74 | | Platforms | **Azure OpenAI** | Your resource's endpoint, its key, and a deployment name | |
| 75 | | | **OpenRouter** | An API key: hundreds of models from every lab | |
| 76 | | | **Groq** | An API key | |
| 77 | | | **Together AI** | An API key | |
| 78 | | | **Fireworks AI** | An API key | |
| 79 | | | **Cerebras** | An API key | |
| 80 | | Any endpoint | **Anthropic-compatible** | A base URL, and a key if it needs one: your own Cloudflare AI Gateway, LiteLLM, Bedrock or Vertex behind a proxy | |
| 81 | | | **OpenAI-compatible** | A base URL including its version, a model, and a key if it needs one: vLLM, Ollama behind a tunnel, LiteLLM | |
| 82 | |
| 83 | An endpoint behind an authenticated Cloudflare AI Gateway also takes the |
| 84 | gateway's token, sent as `cf-aig-authorization`. |
| 85 | |
| 86 | g1t's agents run Claude Code, which speaks Anthropic's API. Every provider |
| 87 | but Anthropic and an Anthropic-compatible endpoint speaks OpenAI's, so g1t's model proxy translates each |
| 88 | request, and the streamed answer back, tool calls included, and meets each |
| 89 | provider's quirks: the token limits DeepSeek and Groq set, how Mistral |
| 90 | names a required tool, Azure's `api-key` header, and the thought signatures |
| 91 | Gemini needs back with each tool call. Agents work the same either way. |
| 92 | How well they work depends on the model: it has to be good at using tools |
| 93 | over many steps. |
| 94 | |
| 95 | ## Connect a provider |
| 96 | |
| 97 | 1. Open the workspace's **Settings → Integrations**. You need to be an owner. |
| 98 | 2. Under **Model providers**, choose one, and give its key (and address, for |
| 99 | an endpoint). |
| 100 | 3. **Connect**. g1t checks the key at once and lists the provider's models. |
| 101 | **Test** checks it again later. |
| 102 | |
| 103 | A workspace can connect any number, including several of the same kind. |
| 104 | |
| 105 | ## Choose which model does which work |
| 106 | |
| 107 | Under **Which model does which work**, each kind of work has a choice: |
| 108 | |
| 109 | | Kind of work | | |
| 110 | | --- | --- | |
| 111 | | Everything | Used for any kind of work that does not choose for itself. | |
| 112 | | Making changes | Writing the change for an issue, and revising it. | |
| 113 | | Reviewing | The second agent that reviews each change. | |
| 114 | | Planning | Turning an outcome into issues. | |
| 115 | | Catching up | Bringing a change up to date with `main`. | |
| 116 | |
| 117 | Each can go to g1t's models, or to any of your providers on any of its |
| 118 | models. |
| 119 | |
| 120 | - **On g1t's models**, choose **Auto** (the default), or a tier for every |
| 121 | run of that kind: **Fast**, **Standard** or **Most capable**. A run on a |
| 122 | chosen tier says the workspace chose it. |
| 123 | - **On an Anthropic provider**, leave the model empty for g1t's choice of |
| 124 | Claude: Auto picks the tier's model for each job, as on g1t's models, on |
| 125 | your key. Name a model to run that one every time. |
| 126 | - **On any other provider**, name the model. |
| 127 | |
| 128 | For example: make changes on Claude through your Anthropic key, review on |
| 129 | GPT through your OpenAI key, catch up on g1t's models on Fast, and plan on |
| 130 | Most capable. |
| 131 | |
| 132 | **Save routing**, and the next runs use it. Without any routing, work goes |
| 133 | to g1t's models where they are open to the workspace, and otherwise to the |
| 134 | first provider you connected. |
| 135 | |
| 136 | A pull request's session says which model ran, and through which provider. |
| 137 | |
| 138 | ## What it costs |
| 139 | |
| 140 | On g1t's models, each run is charged the model at the provider's price |
| 141 | (what AI Gateway priced its requests at, with no markup), the agent rate |
| 142 | on its tokens, and its sandbox time. Auto keeps the first of those down: |
| 143 | the fast model costs about half what the standard one does, and the most |
| 144 | capable up to twice as much, so most of what it saves comes from sending small |
| 145 | jobs to the fast model and from finishing hard ones instead of retrying |
| 146 | them on the same model. |
| 147 | |
| 148 | On your own providers, they bill you for the models. g1t charges each |
| 149 | run's [sandbox time](/guides/usage-and-billing/#sandbox-time), at what it |
| 150 | costs g1t plus 20%, by the second, and the |
| 151 | [agent rate](/guides/usage-and-billing/#the-agent-rate) on the tokens the |
| 152 | run used, from Oct 22, 2026: $0.25 per million, as on g1t's models. Tokens |
| 153 | are counted by g1t's model proxy as answers pass, and by the agent in the |
| 154 | sandbox; the more of the two is charged. On **Usage** it is the line |
| 155 | **Agent rate, your own model key**, with its tokens weighted as the |
| 156 | pricing page says. |
| 157 | |
| 158 | Both count toward the workspace's usage limit like any other charge. |
| 159 | **Usage** also shows the agent's tokens by model. |
| 160 | |
| 161 | ## Your keys never reach a sandbox |
| 162 | |
| 163 | An agent works in a sandbox with internet access, on code and text that |
| 164 | anyone could have written. g1t assumes a sandbox can be talked into |
| 165 | printing its environment, so no key is ever in it: |
| 166 | |
| 167 | 1. When a run starts, g1t gives the sandbox a token for that run only. |
| 168 | 2. The sandbox sends its model requests to `https://models.g1t.sh` with that |
| 169 | token in place of a key. |
| 170 | 3. g1t's model proxy looks the token up, adds the key for the provider the |
| 171 | work is routed to, translates if the provider speaks OpenAI's API, and |
| 172 | forwards the request. Answers stream straight back. |
| 173 | |
| 174 | As each answer passes, the proxy reads how many tokens it used (input, |
| 175 | output, and cache reads and writes) and counts them for the run, under the |
| 176 | person it was for. They show on **Usage** by model, and the |
| 177 | [agent rate](/guides/usage-and-billing/#the-agent-rate) is charged on them. |
| 178 | On g1t's models, the model itself is charged at what AI Gateway priced it |
| 179 | at, never from these counts. |
| 180 | |
| 181 | The token stops working within seconds of the run finishing, however it |
| 182 | ends, and within seconds if you disconnect the provider. A run whose end |
| 183 | g1t never hears about loses it three hours after it starts. Keys are sealed when you save them, and used |
| 184 | only by the proxy. g1t's own runs work the same way, with g1t's key. |
| 185 | |
| 186 | ## From the API |
| 187 | |
| 188 | | MCP tool and action | Route | |
| 189 | | --- | --- | |
| 190 | | `workspace` `connect_integration` | `POST /workspaces/{workspace}/integrations` with `provider` one of `anthropic`, `openai`, `gemini`, `xai`, `mistral`, `deepseek`, `azure_openai`, `openrouter`, `groq`, `together`, `fireworks`, `cerebras`, `anthropic_endpoint`, `openai_endpoint` | |
| 191 | | `workspace` `get_model_routes` | `GET /workspaces/{workspace}/model-routes` | |
| 192 | | `workspace` `set_model_routes` | `PUT /workspaces/{workspace}/model-routes` | |
| 193 | |
| 194 | ```sh |
| 195 | curl -X PUT https://api.g1t.sh/workspaces/acme/model-routes \ |
| 196 | -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \ |
| 197 | -d '{"routes": [ |
| 198 | {"task": "default", "connection_id": "con_…anthropic", "model": null}, |
| 199 | {"task": "review", "connection_id": "con_…openai", "model": "gpt-5"} |
| 200 | ]}' |
| 201 | ``` |
| 202 | |
| 203 | `task` is `default`, `implement`, `review`, `plan` or `update`. |
| 204 | `connection_id` is null for g1t's hosted models. `model` is null for the |
| 205 | provider's default, or for an Anthropic provider, g1t's choice of Claude. |
| 206 | On g1t's hosted models, `model` is `small` (Fast), `large` (Standard) or |
| 207 | `frontier` (Most capable), or null for Auto: |
| 208 | |
| 209 | ```sh |
| 210 | curl -X PUT https://api.g1t.sh/workspaces/acme/model-routes \ |
| 211 | -H "Authorization: Bearer $G1T_TOKEN" -H "Content-Type: application/json" \ |
| 212 | -d '{"routes": [ |
| 213 | {"task": "default", "connection_id": null, "model": null}, |
| 214 | {"task": "plan", "connection_id": null, "model": "frontier"} |
| 215 | ]}' |
| 216 | ``` |