Skip to content

Compare changes

Choose two branches to see what one has that the other does not, then open a pull request for it.

Open a pull request

1 commit

3 files+23−50/3 viewed
+2−2
287287 1. g1t's model proxy counts each answer's tokens as it passes: input,
288288 output, and prompt-cache reads and writes. The sandbox reports what its
289289 agent counted too, and the rate is charged on the more of the two.
290−2. Each kind of token counts at its weight. Today every token counts
291− once: input ×1, output ×1, cache reads ×1, cache writes ×1. The weights
290+2. Each kind of token counts at its weight: input ×1, output ×1, cache
291+ writes ×1, and cache reads ×0.1, as model providers price them. The weights
292292 are on [g1t.sh/pricing](https://g1t.sh/pricing) under the rate, and a
293293 change to them is a dated price change like any other.
294294 3. The run is charged when it reports, and again for tokens counted after
+5−3
525525 and own keys alike, is four price-book meters (migration
526526 `0042_agent_rate_weights.sql`): `agent_token_weight_input`, `_output`,
527527 `_cache_read` and `_cache_write`, each a weight in millionths in
528−`cost_micros` (1,000,000 counts a token once). All start at 1, which is what
529−the rate always counted. A cached agent run reads most of its context from
528+`cost_micros` (1,000,000 counts a token once). They started at 1, which is
529+what the rate always counted; since 2026-10-08 cache reads count a tenth
530+(100,000; migration `0044_cache_reads_count_a_tenth.sql`), as model providers
531+price them. Input, output and cache writes count once. A cached agent run reads most of its context from
530532 cache (about 90% of its tokens on a typical Sonnet implement run), so the
531533 cache-read weight is the lever: at 1 the rate adds about 44% to such a run's
532534 model cost; at 0.1, far less.
533535
534−To count cache reads at a tenth:
536+To change a weight (cache reads went to a tenth this way):
535537
536538 1. Add the version, effective at once (a lower weight is a fall):
537539
+16−0
1+-- Cache reads count a tenth toward the g1t agent rate, as providers price
2+-- them: a cached agent run reads most of its context from cache, and at a
3+-- full weight the rate would add about 44% to a typical Sonnet run for
4+-- tokens that cost the provider a tenth. A lower weight is a price fall, so
5+-- it applies at once; the agent rate itself starts on 2026-10-22.
6+
7+INSERT OR IGNORE INTO price_versions (id, meter, version, cost_micros, markup_percent, effective_at, reason, created_by, created_at, applied_at) VALUES
8+ ('pv_agent_token_weight_cache_read_2', 'agent_token_weight_cache_read', 2, 100000, 0, '2026-10-08T00:00:00Z',
9+ 'Cache reads count a tenth toward the agent rate, as model providers price them', 'migration', '2026-10-08T04:30:00Z', '2026-10-08T04:30:00Z');
10+
11+UPDATE prices SET cost_micros = 100000, updated_at = '2026-10-08T04:30:00Z'
12+ WHERE meter = 'agent_token_weight_cache_read' AND cost_micros = 1000000;
13+
14+INSERT OR IGNORE INTO price_changes (id, meter, old_cost_micros, new_cost_micros, markup_percent, reason, created_at) VALUES
15+ ('prc_agent_token_weight_cache_read_2', 'agent_token_weight_cache_read', 1000000, 100000, 0,
16+ 'Lower: cache reads count a tenth toward the agent rate (input, output and cache writes still count once)', '2026-10-08T04:30:00Z');