Billing: the agent rate's token weights are price-book data
A cached agent run reads most of its context from cache, and the agent rate counted those reads like any other token. How much each kind counts (input, output, cache reads, cache writes) is now four price-book meters, agent_token_weight_*, a weight in millionths, versioned and dated like any price. All start at 1, so nothing changes until the owner decides: a lower weight is a fall and applies at once. The same weights apply on g1t's models and a workspace's own model key. runs.agent_tokens keeps the weighted tokens charged, so a later count charges only what is new; the proxy's count and the harness's are each weighted by kind and the more is charged. Usage's agent-rate lines count weighted tokens and name the weights; Pricing shows them under the rate. Migration 0042.
| 1469 | 1469 | "Choose what happens when a team is asked to review a pull request. Off, everyone in it is asked. On (`enabled`), g1t picks `count` people from it (1 to 10, never the pull request's author) and asks them, and the team stays shown as asked beside them: `round_robin` picks whoever this team asked least recently, `load_balance` whoever has the fewest pull requests waiting on their review. `skip_busy` leaves out anyone with `busy_at` or more waiting; `include_child_teams` also picks from its child teams' people; `excluded` lists usernames never picked; `notify_team` also tells the rest of the team. Fields left out keep their current value. Owners of the workspace and the team's maintainers. People only. Returns the team." | |
| 1470 | 1470 | } | |
| 1471 | 1471 | Op::GetUsage => { | |
| 1472 | − | "A workspace's usage over a range of days, at price, and what paid for it. `from` and `until` are UTC days, `YYYY-MM-DD`, with `until` included and at most 400 days in all; left out, the current month so far. `products` narrows it to product families (agent, sandboxes, gateway, deployments, git_storage, packages, security, search) and `projects` to repositories (\"owner/name\"). Returns `totals`: `price_micros` less `discount_micros`, `included_micros` and `credits_micros` is `charged_micros`, what is left for the workspace to pay; `pending_micros` is metered this month and charged when it closes; `cost_micros` is what it cost g1t. Then `days` (each day and product with usage), `products` (every family, with its meters: quantity, unit, amount, a `daily` amount for each day of the range, any `allowance` and the split `by_project`), `projects` (every repository with usage in the range), and the AI credit and other credit left now. With `group_by` (`product`, `project` or `day`), `groups` adds up the range that way. Amounts are whole millionths of a dollar. Members of the workspace only." | |
| 1472 | + | "A workspace's usage over a range of days, at price, and what paid for it. `from` and `until` are UTC days, `YYYY-MM-DD`, with `until` included and at most 400 days in all; left out, the current month so far. `products` narrows it to product families (agent, sandboxes, gateway, deployments, git_storage, packages, security, search) and `projects` to repositories (\"owner/name\"). Returns `totals`: `price_micros` less `discount_micros`, `included_micros` and `credits_micros` is `charged_micros`, what is left for the workspace to pay; `pending_micros` is metered this month and charged when it closes; `cost_micros` is what it cost g1t. Then `days` (each day and product with usage), `products` (every family, with its meters: quantity, unit, amount, a `daily` amount for each day of the range, any `allowance`, the split `by_project`, and a `note` where the quantity needs one: the agent rate's meters, `agent_rate` and `agent_rate_own` (on the workspace's own model key), count weighted tokens and name the weights), `projects` (every repository with usage in the range), `models` (the agent's input, output, cache-read and cache-write tokens by model, most first), and the AI credit and other credit left now. With `group_by` (`product`, `project` or `day`), `groups` adds up the range that way. Amounts are whole millionths of a dollar. Members of the workspace only." | |
| 1473 | 1473 | } | |
| 1474 | 1474 | Op::GetBudget => { | |
| 1475 | 1475 | "A workspace's budget: its monthly spend limit (`amount_micros`; `automatic` is true while the owners have not set one, and it is then $200 or twice last month's spend), what was charged this month (`spent_micros`), the most the owners may set it to themselves (`max_amount_micros`), its `alerts` (percent of the limit, each emailed to the owners once a month), whether usage pauses at the limit (`pause_at_limit`), the `webhook` told of each alert, and `state`: `ok`, `warning` or `stopped`, with a `message` when work is stopped or close to it. Members of the workspace only." |
| 5778 | 5778 | "flagon-io/g1t", | |
| 5779 | 5779 | "flagon-io/hello" | |
| 5780 | 5780 | ], | |
| 5781 | + | "models": [ | |
| 5782 | + | { | |
| 5783 | + | "model": "claude-sonnet-5-5", | |
| 5784 | + | "input": 310000, | |
| 5785 | + | "output": 420000, | |
| 5786 | + | "cache_read": 7800000, | |
| 5787 | + | "cache_write": 590000 | |
| 5788 | + | } | |
| 5789 | + | ], | |
| 5781 | 5790 | "included": { | |
| 5782 | 5791 | "used": 20, | |
| 5783 | 5792 | "of": 20, |
Binary or large file; its contents are not shown.
| 483 | 483 | <> | |
| 484 | 484 | <span className="flex min-w-0 items-center gap-2"> | |
| 485 | 485 | <span className="size-2 shrink-0 rounded-sm" style={{ background: color }} /> | |
| 486 | − | <span className="truncate">{meter.label}</span> | |
| 486 | + | <span className="min-w-0"> | |
| 487 | + | <span className="block truncate">{meter.label}</span> | |
| 488 | + | {meter.note && <span className="block truncate text-[0.6875rem] text-faint" title={meter.note}>{meter.note}</span>} | |
| 489 | + | </span> | |
| 487 | 490 | {parts.length > 0 && <ChevronDown size={13} className="shrink-0 text-faint transition-transform group-open:rotate-180" />} | |
| 488 | 491 | </span> | |
| 489 | 492 | <span className="hidden sm:block"> |
| 43 | 43 | } | |
| 44 | 44 | ||
| 45 | 45 | /** Price-book rows the table shows in rows of their own, or not at all. */ | |
| 46 | − | const SHOWN_APART = new Set(["app_month", "agent_models", "agent_tokens", "agent_tokens_own", "gateway_models", "card_fee_percent", "card_fee_fixed"]); | |
| 46 | + | const SHOWN_APART = new Set([ | |
| 47 | + | "app_month", | |
| 48 | + | "agent_models", | |
| 49 | + | "agent_tokens", | |
| 50 | + | "agent_tokens_own", | |
| 51 | + | "agent_token_weight_input", | |
| 52 | + | "agent_token_weight_output", | |
| 53 | + | "agent_token_weight_cache_read", | |
| 54 | + | "agent_token_weight_cache_write", | |
| 55 | + | "gateway_models", | |
| 56 | + | "card_fee_percent", | |
| 57 | + | "card_fee_fixed", | |
| 58 | + | ]); | |
| 59 | + | ||
| 60 | + | /** How each kind of token counts toward the agent rate, from the price book: `input ×1, …`. */ | |
| 61 | + | function tokenWeights(prices: { meter: string; costMicros: number }[]): string { | |
| 62 | + | const weight = (kind: string) => { | |
| 63 | + | const row = prices.find((p) => p.meter === `agent_token_weight_${kind}`); | |
| 64 | + | return row ? Number((row.costMicros / 1_000_000).toFixed(3)) : 1; | |
| 65 | + | }; | |
| 66 | + | return `input ×${weight("input")}, output ×${weight("output")}, cache reads ×${weight("cache_read")}, cache writes ×${weight("cache_write")}`; | |
| 67 | + | } | |
| 47 | 68 | ||
| 48 | 69 | /** What the page says when billing cannot be reached: the published defaults. */ | |
| 49 | 70 | const DEFAULT_FREE: Required<FreeTier> = { | |
| ⋯ | |||
| 470 | 491 | <td className="px-4 py-3"> | |
| 471 | 492 | <p className="font-medium">g1t agent rate</p> | |
| 472 | 493 | <p className="text-xs text-faint">Context, memory, routing and orchestration, on every token an agent run uses</p> | |
| 494 | + | <p className="text-xs text-faint">Tokens count by kind: {tokenWeights(book?.prices ?? [])}</p> | |
| 473 | 495 | </td> | |
| 474 | 496 | <td className="px-4 py-3 text-muted">A flat rate</td> | |
| 475 | 497 | <td className="hidden px-4 py-3 tabular-nums sm:table-cell">—</td> | |
| 3237 | 3237 | #[serde(default)] | |
| 3238 | 3238 | pub allowance: Option<Allowance>, | |
| 3239 | 3239 | pub by_project: Vec<ProjectUsage>, | |
| 3240 | + | /// How the quantity is counted, when that needs saying: for the agent | |
| 3241 | + | /// rate, its tokens are weighted by kind, and this names the weights. | |
| 3242 | + | #[serde(default)] | |
| 3243 | + | pub note: Option<String>, | |
| 3240 | 3244 | } | |
| 3241 | 3245 | ||
| 3242 | 3246 | /// A part of a product, such as the agent's runs, reviews and plans. |
| 1468 | 1468 | daily: number[]; | |
| 1469 | 1469 | allowance?: UsageAllowance | null; | |
| 1470 | 1470 | byProject: ProjectUsage[]; | |
| 1471 | + | /** How the quantity is counted, when that needs saying (the agent rate's token weights). */ | |
| 1472 | + | note?: string | null; | |
| 1471 | 1473 | }; | |
| 1472 | 1474 | ||
| 1473 | 1475 | export type FeatureUsage = { key: string; label: string; micros: number; count: number }; |
| 1 | + | -- How much each kind of token counts toward the g1t agent rate, as | |
| 2 | + | -- price-book meters, on g1t's models and a workspace's own model key alike. | |
| 3 | + | -- | |
| 4 | + | -- A weight is `cost_micros` in millionths: 1,000,000 counts a token once, | |
| 5 | + | -- 100,000 counts it a tenth. Every weight starts at 1, which is what the | |
| 6 | + | -- agent rate counted until now (input, output, cache reads and cache | |
| 7 | + | -- writes, each once), so nothing changes until a new version is decided. | |
| 8 | + | -- A cached agent run reads most of its context from cache, so cache reads | |
| 9 | + | -- dominate its tokens; counting them at 0.1 is the lever | |
| 10 | + | -- (docs/BILLING_OPERATIONS.md, "The agent rate's token weights"). | |
| 11 | + | -- | |
| 12 | + | -- Weights change as any price does: a new version in price_versions with | |
| 13 | + | -- its effective date. A lower weight is a fall and applies at once; a | |
| 14 | + | -- higher one is a rise and waits out the notice. | |
| 15 | + | ||
| 16 | + | INSERT OR IGNORE INTO prices (meter, title, unit, cost_micros, markup_percent, source, updated_at) VALUES | |
| 17 | + | ('agent_token_weight_input', 'Agent rate weight: input tokens', 'weight', 1000000, 0, 'list', '2026-10-08T00:00:00Z'), | |
| 18 | + | ('agent_token_weight_output', 'Agent rate weight: output tokens', 'weight', 1000000, 0, 'list', '2026-10-08T00:00:00Z'), | |
| 19 | + | ('agent_token_weight_cache_read', 'Agent rate weight: cache reads', 'weight', 1000000, 0, 'list', '2026-10-08T00:00:00Z'), | |
| 20 | + | ('agent_token_weight_cache_write', 'Agent rate weight: cache writes', 'weight', 1000000, 0, 'list', '2026-10-08T00:00:00Z'); | |
| 21 | + | ||
| 22 | + | INSERT OR IGNORE INTO price_versions (id, meter, version, cost_micros, markup_percent, effective_at, reason, created_by, created_at, applied_at) VALUES | |
| 23 | + | ('pv_agent_token_weight_input_1', 'agent_token_weight_input', 1, 1000000, 0, '2026-10-08T00:00:00Z', | |
| 24 | + | 'Every input token counts once toward the agent rate', 'migration', '2026-10-08T00:00:00Z', '2026-10-08T00:00:00Z'), | |
| 25 | + | ('pv_agent_token_weight_output_1', 'agent_token_weight_output', 1, 1000000, 0, '2026-10-08T00:00:00Z', | |
| 26 | + | 'Every output token counts once toward the agent rate', 'migration', '2026-10-08T00:00:00Z', '2026-10-08T00:00:00Z'), | |
| 27 | + | ('pv_agent_token_weight_cache_read_1', 'agent_token_weight_cache_read', 1, 1000000, 0, '2026-10-08T00:00:00Z', | |
| 28 | + | 'Every cache-read token counts once toward the agent rate', 'migration', '2026-10-08T00:00:00Z', '2026-10-08T00:00:00Z'), | |
| 29 | + | ('pv_agent_token_weight_cache_write_1', 'agent_token_weight_cache_write', 1, 1000000, 0, '2026-10-08T00:00:00Z', | |
| 30 | + | 'Every cache-write token counts once toward the agent rate', 'migration', '2026-10-08T00:00:00Z', '2026-10-08T00:00:00Z'); |
| 36 | 36 | ||
| 37 | 37 | use g1t_contracts::billing::{ | |
| 38 | 38 | AccountArgs, AiCredit, AiReload, BuyAiCreditArgs, CardFee, Checkout, ConfirmAiCreditArgs, CreditKind, EntryKind, PlanKind, | |
| 39 | − | SetAiReloadArgs, MICROS_PER_DOLLAR, | |
| 39 | + | RunTokens, SetAiReloadArgs, MICROS_PER_DOLLAR, | |
| 40 | 40 | }; | |
| 41 | 41 | use g1t_contracts::time::rfc3339; | |
| 42 | 42 | use g1t_contracts::{FailureCode, Outcome, Role}; | |
| ⋯ | |||
| 802 | 802 | ||
| 803 | 803 | // --- The agent rate -------------------------------------------------------- | |
| 804 | 804 | ||
| 805 | + | /// How much each kind of token counts toward the agent rate: the price | |
| 806 | + | /// book's `agent_token_weight_*` meters, a token's weight in millionths | |
| 807 | + | /// (1,000,000 is 1). A meter missing counts 1. | |
| 808 | + | pub(crate) async fn token_weights(&self) -> Result<TokenWeights> { | |
| 809 | + | let mut weights = TokenWeights::default(); | |
| 810 | + | for (meter, slot) in [ | |
| 811 | + | ("agent_token_weight_input", &mut weights.input), | |
| 812 | + | ("agent_token_weight_output", &mut weights.output), | |
| 813 | + | ("agent_token_weight_cache_read", &mut weights.cache_read), | |
| 814 | + | ("agent_token_weight_cache_write", &mut weights.cache_write), | |
| 815 | + | ] { | |
| 816 | + | if let Some((cost, _)) = self.price(meter).await? { | |
| 817 | + | *slot = weight_of(cost); | |
| 818 | + | } | |
| 819 | + | } | |
| 820 | + | Ok(weights) | |
| 821 | + | } | |
| 822 | + | ||
| 805 | 823 | /// Charges a run's agent rate for the tokens counted since it was last | |
| 806 | 824 | /// charged, once each: when the run reports and again when it is | |
| 807 | 825 | /// settled, so tokens counted late are charged too. `reported` is what | |
| 808 | − | /// the run's harness counted; the rate is charged on no fewer. | |
| 826 | + | /// the run's harness counted; the rate is charged on no fewer. Tokens | |
| 827 | + | /// are weighted by kind (`token_weights`), and `runs.agent_tokens` | |
| 828 | + | /// keeps the weighted tokens charged so far. | |
| 809 | 829 | /// | |
| 810 | 830 | /// On the workspace's own provider the model is not g1t's to charge, | |
| 811 | 831 | /// but the agent rate is, on its own meter (`agent_tokens_own`) and its | |
| 812 | 832 | /// own line (`<run>/agent-own`), so Usage and the statement show it as | |
| 813 | 833 | /// the agent rate on the workspace's own model key. | |
| 814 | − | pub(crate) async fn charge_agent_rate(&self, run_id: &str, run: &RunRow, reported: u64) -> Result<()> { | |
| 834 | + | pub(crate) async fn charge_agent_rate(&self, run_id: &str, run: &RunRow, reported: Option<RunTokens>) -> Result<()> { | |
| 815 | 835 | if self.stripe.is_none() { | |
| 816 | 836 | return Ok(()); | |
| 817 | 837 | } | |
| ⋯ | |||
| 830 | 850 | else { | |
| 831 | 851 | return Ok(()); | |
| 832 | 852 | }; | |
| 853 | + | #[derive(Deserialize)] | |
| 854 | + | struct Kinds { | |
| 855 | + | input: Option<f64>, | |
| 856 | + | output: Option<f64>, | |
| 857 | + | cache_read: Option<f64>, | |
| 858 | + | cache_write: Option<f64>, | |
| 859 | + | } | |
| 833 | 860 | let counted = match &row.session_id { | |
| 834 | 861 | Some(session) => self | |
| 835 | 862 | .db | |
| 836 | 863 | .prepare( | |
| 837 | − | "SELECT SUM(input + output + cache_read + cache_write) AS micros FROM token_usage | |
| 838 | − | WHERE workspace = ? AND session = ? AND day >= ?", | |
| 864 | + | "SELECT SUM(input) AS input, SUM(output) AS output, SUM(cache_read) AS cache_read, SUM(cache_write) AS cache_write | |
| 865 | + | FROM token_usage WHERE workspace = ? AND session = ? AND day >= ?", | |
| 839 | 866 | ) | |
| 840 | 867 | .bind(&[run.workspace.as_str().into(), session.as_str().into(), row.created_at[..10].into()])? | |
| 841 | − | .first::<Sum>(None) | |
| 868 | + | .first::<Kinds>(None) | |
| 842 | 869 | .await? | |
| 843 | − | .and_then(|s| s.micros) | |
| 844 | − | .unwrap_or(0.0) as u64, | |
| 845 | − | None => 0, | |
| 870 | + | .map(|k| { | |
| 871 | + | let n = |v: Option<f64>| v.unwrap_or(0.0).max(0.0) as u64; | |
| 872 | + | RunTokens { input: n(k.input), output: n(k.output), cache_read: n(k.cache_read), cache_write: n(k.cache_write) } | |
| 873 | + | }) | |
| 874 | + | .unwrap_or_default(), | |
| 875 | + | None => RunTokens::default(), | |
| 846 | 876 | }; | |
| 877 | + | let weights = self.token_weights().await?; | |
| 847 | 878 | let charged = row.agent_tokens.unwrap_or(0.0) as u64; | |
| 848 | − | let Some(total) = tokens_to_charge(counted, reported, charged) else { | |
| 879 | + | let Some(total) = tokens_to_charge( | |
| 880 | + | weighted(&counted, &weights), | |
| 881 | + | reported.map_or(0, |tokens| weighted(&tokens, &weights)), | |
| 882 | + | charged, | |
| 883 | + | ) else { | |
| 849 | 884 | return Ok(()); | |
| 850 | 885 | }; | |
| 851 | 886 | // Claimed first: two callers never charge the same tokens. | |
| ⋯ | |||
| 879 | 914 | _ => format!("work on {}#{}", run.repo, run.number), | |
| 880 | 915 | }; | |
| 881 | 916 | let label = if own { "g1t agent rate, your own model key" } else { "g1t agent rate" }; | |
| 882 | − | let description = format!("{label}: {} tokens for {what}{terms_note}{}", crate::features::thousands(tokens), drawn.note()); | |
| 917 | + | let counted_as = if weights.is_flat() { "tokens".to_owned() } else { format!("weighted tokens ({})", weights.describe()) }; | |
| 918 | + | let description = format!("{label}: {} {counted_as} for {what}{terms_note}{}", crate::features::thousands(tokens), drawn.note()); | |
| 883 | 919 | // g1t's own charge, even on the workspace's provider: it counts | |
| 884 | 920 | // toward limits and spend like any other. | |
| 885 | 921 | let line = RunRow { billed_to: None, ..run.clone() }; | |
| ⋯ | |||
| 897 | 933 | } | |
| 898 | 934 | } | |
| 899 | 935 | ||
| 936 | + | /// How much each kind of token counts toward the agent rate. All 1 by | |
| 937 | + | /// default: every token counts once. | |
| 938 | + | #[derive(Clone, Copy, Debug, PartialEq)] | |
| 939 | + | pub(crate) struct TokenWeights { | |
| 940 | + | pub input: f64, | |
| 941 | + | pub output: f64, | |
| 942 | + | pub cache_read: f64, | |
| 943 | + | pub cache_write: f64, | |
| 944 | + | } | |
| 945 | + | ||
| 946 | + | impl Default for TokenWeights { | |
| 947 | + | fn default() -> Self { | |
| 948 | + | TokenWeights { input: 1.0, output: 1.0, cache_read: 1.0, cache_write: 1.0 } | |
| 949 | + | } | |
| 950 | + | } | |
| 951 | + | ||
| 952 | + | impl TokenWeights { | |
| 953 | + | /// Every token counts once. | |
| 954 | + | pub(crate) fn is_flat(&self) -> bool { | |
| 955 | + | *self == TokenWeights::default() | |
| 956 | + | } | |
| 957 | + | ||
| 958 | + | /// `input ×1, output ×1, cache reads ×0.1, cache writes ×1`. | |
| 959 | + | pub(crate) fn describe(&self) -> String { | |
| 960 | + | let w = |v: f64| { | |
| 961 | + | let text = format!("{v:.3}"); | |
| 962 | + | text.trim_end_matches('0').trim_end_matches('.').to_owned() | |
| 963 | + | }; | |
| 964 | + | format!( | |
| 965 | + | "input ×{}, output ×{}, cache reads ×{}, cache writes ×{}", | |
| 966 | + | w(self.input), | |
| 967 | + | w(self.output), | |
| 968 | + | w(self.cache_read), | |
| 969 | + | w(self.cache_write) | |
| 970 | + | ) | |
| 971 | + | } | |
| 972 | + | } | |
| 973 | + | ||
| 974 | + | /// A weight from its price-book figure, in millionths; never below 0. | |
| 975 | + | pub(crate) fn weight_of(micros: f64) -> f64 { | |
| 976 | + | if micros.is_finite() { (micros / 1_000_000.0).max(0.0) } else { 1.0 } | |
| 977 | + | } | |
| 978 | + | ||
| 979 | + | /// A run's tokens as the agent rate counts them, each kind at its weight, | |
| 980 | + | /// rounded down to a whole token. | |
| 981 | + | pub(crate) fn weighted(tokens: &RunTokens, weights: &TokenWeights) -> u64 { | |
| 982 | + | let sum = tokens.input as f64 * weights.input | |
| 983 | + | + tokens.output as f64 * weights.output | |
| 984 | + | + tokens.cache_read as f64 * weights.cache_read | |
| 985 | + | + tokens.cache_write as f64 * weights.cache_write; | |
| 986 | + | sum.max(0.0).floor() as u64 | |
| 987 | + | } | |
| 988 | + | ||
| 900 | 989 | /// The price-book meter a run's agent rate is on: its own for runs on the | |
| 901 | 990 | /// workspace's own model key, so it can be priced and shown apart. | |
| 902 | 991 | pub(crate) fn agent_rate_meter(own_provider: bool) -> &'static str { | |
| ⋯ | |||
| 1010 | 1099 | assert_eq!(agent_rate_reference("run_1", false, 900, 1_200), "run_1/agent/1200"); | |
| 1011 | 1100 | } | |
| 1012 | 1101 | ||
| 1102 | + | fn tokens(input: u64, output: u64, cache_read: u64, cache_write: u64) -> RunTokens { | |
| 1103 | + | RunTokens { input, output, cache_read, cache_write } | |
| 1104 | + | } | |
| 1105 | + | ||
| 1106 | + | #[test] | |
| 1107 | + | fn every_token_counts_once_by_default() { | |
| 1108 | + | let flat = TokenWeights::default(); | |
| 1109 | + | assert!(flat.is_flat()); | |
| 1110 | + | assert_eq!(weighted(&tokens(1_000, 500, 90_000, 5_000), &flat), 96_500); | |
| 1111 | + | assert_eq!(weighted(&RunTokens::default(), &flat), 0); | |
| 1112 | + | } | |
| 1113 | + | ||
| 1114 | + | #[test] | |
| 1115 | + | fn cache_reads_can_count_for_a_tenth() { | |
| 1116 | + | let tenth = TokenWeights { cache_read: weight_of(100_000.0), ..TokenWeights::default() }; | |
| 1117 | + | assert!(!tenth.is_flat()); | |
| 1118 | + | // 1,000 + 500 + 9,000 + 5,000. | |
| 1119 | + | assert_eq!(weighted(&tokens(1_000, 500, 90_000, 5_000), &tenth), 15_500); | |
| 1120 | + | // Rounded down to a whole token. | |
| 1121 | + | assert_eq!(weighted(&tokens(0, 0, 15, 0), &tenth), 1); | |
| 1122 | + | assert_eq!(tenth.describe(), "input ×1, output ×1, cache reads ×0.1, cache writes ×1"); | |
| 1123 | + | assert_eq!(TokenWeights::default().describe(), "input ×1, output ×1, cache reads ×1, cache writes ×1"); | |
| 1124 | + | } | |
| 1125 | + | ||
| 1126 | + | #[test] | |
| 1127 | + | fn a_weight_is_never_below_nothing() { | |
| 1128 | + | assert_eq!(weight_of(1_000_000.0), 1.0); | |
| 1129 | + | assert_eq!(weight_of(1_250_000.0), 1.25); | |
| 1130 | + | assert_eq!(weight_of(-5.0), 0.0); | |
| 1131 | + | assert_eq!(weight_of(f64::NAN), 1.0); | |
| 1132 | + | } | |
| 1133 | + | ||
| 1013 | 1134 | #[test] | |
| 1014 | 1135 | fn the_rate_covers_the_more_of_what_was_counted_and_reported_once() { | |
| 1015 | 1136 | // The proxy counted nothing (no session, or its reports were lost): | |
| 658 | 658 | token_hash: run.token_hash, | |
| 659 | 659 | billed_to: run.billed_to, | |
| 660 | 660 | }; | |
| 661 | − | self.charge_agent_rate(&run.id, &row, 0).await?; | |
| 661 | + | self.charge_agent_rate(&run.id, &row, None).await?; | |
| 662 | 662 | } | |
| 663 | 663 | Ok(()) | |
| 664 | 664 | } | |
| ⋯ | |||
| 715 | 715 | } | |
| 716 | 716 | // Tokens counted after the run reported are charged their agent | |
| 717 | 717 | // rate now (ai.rs). | |
| 718 | − | self.charge_agent_rate(&run.id, &row, 0).await?; | |
| 718 | + | self.charge_agent_rate(&run.id, &row, None).await?; | |
| 719 | 719 | if let Some(why) = &short { | |
| 720 | 720 | worker::console_warn!("run {} settled at no less than reported: {why}", run.id); | |
| 721 | 721 | } | |
| 817 | 817 | if claimed.is_none() { | |
| 818 | 818 | return Ok(Outcome::Ok(false)); | |
| 819 | 819 | } | |
| 820 | − | // What the harness counted, the floor of what the agent rate is | |
| 821 | − | // charged on when the proxy counted fewer (or none). | |
| 822 | − | let reported = a.tokens.map_or(0, |tokens| tokens.total()); | |
| 823 | 820 | // On the workspace's own provider, the model was paid for there and | |
| 824 | 821 | // the run's sandbox time is recorded on its own: only the agent rate | |
| 825 | 822 | // is charged, on its own meter. | |
| 826 | 823 | if run.own_provider() { | |
| 827 | − | self.charge_agent_rate(&a.run_id, &run, reported).await?; | |
| 824 | + | self.charge_agent_rate(&a.run_id, &run, a.tokens).await?; | |
| 828 | 825 | return Ok(Outcome::Ok(true)); | |
| 829 | 826 | } | |
| 830 | 827 | // Its cost plus the margin, on the account's terms; then the plan's | |
| ⋯ | |||
| 862 | 859 | self.record_drawn(&a.run_id, &drawn).await?; | |
| 863 | 860 | self.record_discount(&a.run_id, discount).await?; | |
| 864 | 861 | self.count_spend(&run.workspace, charge_micros(a.cost_usd, 0), charge - drawn.total(), &drawn).await; | |
| 865 | − | self.charge_agent_rate(&a.run_id, &run, reported).await?; | |
| 862 | + | self.charge_agent_rate(&a.run_id, &run, a.tokens).await?; | |
| 866 | 863 | Ok(Outcome::Ok(true)) | |
| 867 | 864 | } | |
| 868 | 865 | } | |
| ⋯ | |||
| 1462 | 1459 | include_str!("../migrations/0039_discounts_not_comped.sql"), | |
| 1463 | 1460 | include_str!("../migrations/0040_ai_credit.sql"), | |
| 1464 | 1461 | include_str!("../migrations/0041_agent_rate_own_key.sql"), | |
| 1462 | + | include_str!("../migrations/0042_agent_rate_weights.sql"), | |
| 1465 | 1463 | ]; | |
| 1466 | 1464 | ||
| 1467 | 1465 | /// The columns of `table` after the migrations: each with whether an | |
| ⋯ | |||
| 1528 | 1526 | assert!(row("('agent_tokens_own',").contains("'million tokens', 0, 0, 'list'")); | |
| 1529 | 1527 | assert_eq!(ai::agent_rate_meter(true), "agent_tokens_own"); | |
| 1530 | 1528 | } | |
| 1529 | + | ||
| 1530 | + | #[test] | |
| 1531 | + | fn the_agent_rate_counts_every_token_once_until_a_weight_is_decided() { | |
| 1532 | + | let sql = include_str!("../migrations/0042_agent_rate_weights.sql"); | |
| 1533 | + | for kind in ["input", "output", "cache_read", "cache_write"] { | |
| 1534 | + | let meter = format!("('agent_token_weight_{kind}', 'Agent rate weight"); | |
| 1535 | + | let row = sql.lines().find(|l| l.contains(&meter)).unwrap_or_else(|| panic!("no {kind} weight")); | |
| 1536 | + | assert!(row.contains("'weight', 1000000, 0, 'list'"), "{kind}: {row}"); | |
| 1537 | + | let version = format!("('pv_agent_token_weight_{kind}_1', 'agent_token_weight_{kind}', 1, 1000000,"); | |
| 1538 | + | assert!(sql.contains(&version), "{kind} has no applied version"); | |
| 1539 | + | } | |
| 1540 | + | assert_eq!(ai::weight_of(1_000_000.0), 1.0); | |
| 1541 | + | } | |
| 1531 | 1542 | #[test] | |
| 1532 | 1543 | fn every_checkout_insert_fills_the_table() { | |
| 1533 | 1544 | // The plan's, the activation's, a card check's, a prepayment's and | |
| 179 | 179 | daily: vec![0; days.len()], | |
| 180 | 180 | allowance: None, | |
| 181 | 181 | by_project: vec![], | |
| 182 | + | note: None, | |
| 182 | 183 | }) | |
| 183 | 184 | .collect(); | |
| 184 | 185 | let mut by_day: BTreeMap<(String, String), i64> = BTreeMap::new(); | |
| ⋯ | |||
| 426 | 427 | trial = Some(crate::credits::left(grant.granted_micros, grant.used_micros)); | |
| 427 | 428 | } | |
| 428 | 429 | } | |
| 430 | + | // The agent rate's tokens are weighted by kind: say how. | |
| 431 | + | let weights = self.token_weights().await?; | |
| 432 | + | for product in &mut products { | |
| 433 | + | for meter in &mut product.meters { | |
| 434 | + | if meter.key == "agent_rate" || meter.key == "agent_rate_own" { | |
| 435 | + | meter.note = Some(agent_rate_note(&weights)); | |
| 436 | + | } | |
| 437 | + | } | |
| 438 | + | } | |
| 429 | 439 | let ai: i64 = credits.grants.iter().filter(|g| g.scope == "models").map(|g| g.left_micros).sum(); | |
| 430 | 440 | let models = self.tokens_by_model(&workspace, &from, &until).await?; | |
| 431 | 441 | Ok(Outcome::Ok(UsageReport { | |
| ⋯ | |||
| 447 | 457 | } | |
| 448 | 458 | } | |
| 449 | 459 | ||
| 460 | + | /// What the agent rate's meters say of their tokens. | |
| 461 | + | pub(crate) fn agent_rate_note(weights: &crate::ai::TokenWeights) -> String { | |
| 462 | + | if weights.is_flat() { | |
| 463 | + | "Weighted tokens: every token counts once (input ×1, output ×1, cache reads ×1, cache writes ×1)".to_owned() | |
| 464 | + | } else { | |
| 465 | + | format!("Weighted tokens: {}", weights.describe()) | |
| 466 | + | } | |
| 467 | + | } | |
| 468 | + | ||
| 450 | 469 | impl Billing { | |
| 451 | 470 | /// Agent tokens by model over the days `from` to `until`, most first: | |
| 452 | 471 | /// what the model proxy counted, on g1t's models and the workspace's | |