Compare changes
Choose two branches to see what one has that the other does not, then open a pull request for it.
1 commit
3 files+23−50/3 viewed
| 287 | 287 | 1. g1t's model proxy counts each answer's tokens as it passes: input, | |
| 288 | 288 | output, and prompt-cache reads and writes. The sandbox reports what its | |
| 289 | 289 | agent counted too, and the rate is charged on the more of the two. | |
| 290 | − | 2. Each kind of token counts at its weight. Today every token counts | |
| 291 | − | once: input ×1, output ×1, cache reads ×1, cache writes ×1. The weights | |
| 290 | + | 2. Each kind of token counts at its weight: input ×1, output ×1, cache | |
| 291 | + | writes ×1, and cache reads ×0.1, as model providers price them. The weights | |
| 292 | 292 | are on [g1t.sh/pricing](https://g1t.sh/pricing) under the rate, and a | |
| 293 | 293 | change to them is a dated price change like any other. | |
| 294 | 294 | 3. The run is charged when it reports, and again for tokens counted after |
| 525 | 525 | and own keys alike, is four price-book meters (migration | |
| 526 | 526 | `0042_agent_rate_weights.sql`): `agent_token_weight_input`, `_output`, | |
| 527 | 527 | `_cache_read` and `_cache_write`, each a weight in millionths in | |
| 528 | − | `cost_micros` (1,000,000 counts a token once). All start at 1, which is what | |
| 529 | − | the rate always counted. A cached agent run reads most of its context from | |
| 528 | + | `cost_micros` (1,000,000 counts a token once). They started at 1, which is | |
| 529 | + | what the rate always counted; since 2026-10-08 cache reads count a tenth | |
| 530 | + | (100,000; migration `0044_cache_reads_count_a_tenth.sql`), as model providers | |
| 531 | + | price them. Input, output and cache writes count once. A cached agent run reads most of its context from | |
| 530 | 532 | cache (about 90% of its tokens on a typical Sonnet implement run), so the | |
| 531 | 533 | cache-read weight is the lever: at 1 the rate adds about 44% to such a run's | |
| 532 | 534 | model cost; at 0.1, far less. | |
| 533 | 535 | ||
| 534 | − | To count cache reads at a tenth: | |
| 536 | + | To change a weight (cache reads went to a tenth this way): | |
| 535 | 537 | ||
| 536 | 538 | 1. Add the version, effective at once (a lower weight is a fall): | |
| 537 | 539 |
| 1 | + | -- Cache reads count a tenth toward the g1t agent rate, as providers price | |
| 2 | + | -- them: a cached agent run reads most of its context from cache, and at a | |
| 3 | + | -- full weight the rate would add about 44% to a typical Sonnet run for | |
| 4 | + | -- tokens that cost the provider a tenth. A lower weight is a price fall, so | |
| 5 | + | -- it applies at once; the agent rate itself starts on 2026-10-22. | |
| 6 | + | ||
| 7 | + | INSERT OR IGNORE INTO price_versions (id, meter, version, cost_micros, markup_percent, effective_at, reason, created_by, created_at, applied_at) VALUES | |
| 8 | + | ('pv_agent_token_weight_cache_read_2', 'agent_token_weight_cache_read', 2, 100000, 0, '2026-10-08T00:00:00Z', | |
| 9 | + | 'Cache reads count a tenth toward the agent rate, as model providers price them', 'migration', '2026-10-08T04:30:00Z', '2026-10-08T04:30:00Z'); | |
| 10 | + | ||
| 11 | + | UPDATE prices SET cost_micros = 100000, updated_at = '2026-10-08T04:30:00Z' | |
| 12 | + | WHERE meter = 'agent_token_weight_cache_read' AND cost_micros = 1000000; | |
| 13 | + | ||
| 14 | + | INSERT OR IGNORE INTO price_changes (id, meter, old_cost_micros, new_cost_micros, markup_percent, reason, created_at) VALUES | |
| 15 | + | ('prc_agent_token_weight_cache_read_2', 'agent_token_weight_cache_read', 1000000, 100000, 0, | |
| 16 | + | 'Lower: cache reads count a tenth toward the agent rate (input, output and cache writes still count once)', '2026-10-08T04:30:00Z'); |