repos: git operations are written even when no request follows, and a pull request's working copy counts for its repository's workspace
Sandboxes were not the gap with Cloudflare's count: every sandbox (agent runs, checks, reviews, builds, workflow jobs, and git an agent runs itself) clones, fetches and pushes through g1t's git endpoints with a run credential and is metered there. Only a backup reads the store directly, and its runner reports it. Two things were wrong instead: - An isolate wrote its counts only when a request found the last write 5 s old, so what was counted since waited for a later request on the same isolate and was lost when none came. A clone's fetch, its last request and the one that is an operation, was the one most often stranded. Now a request plans the write that follows in its own wait_until, waiting until it is due: still at most one write per isolate every 5 s, nothing on the request path. - A pull request's working copy (pulls/<pull id>, where agents clone and push) was counted for a workspace called "pulls". Its counts now go to the workspace of the repository it came from, looked up when they are written (one query per write, kept 10 minutes). No migration: operation_mapping is unchanged. Docs: ARTIFACTS.md (R1, M15, the 2026-10-07 findings), git and usage-and-billing guides.
| 186 | 186 | ||
| 187 | 187 | ### Git operations | |
| 188 | 188 | ||
| 189 | − | Each clone, fetch and push is a git operation. Every workspace has 50,000 | |
| 189 | + | Each clone, fetch and push is a git operation, your agents' included: | |
| 190 | + | their sandboxes use the same git endpoints you do, and a pull request's | |
| 191 | + | working copy counts for its repository's workspace. Every workspace has 50,000 | |
| 190 | 192 | a month included. Past that, a workspace on the g1t plan pays $0.18 per | |
| 191 | 193 | 1,000, and a free workspace is never charged: past 50,000 in a month, its | |
| 192 | 194 | git requests past 60 in an hour are answered `429` with when to try again, |
| 337 | 337 | what counts follows what Cloudflare confirms it bills; this page changes | |
| 338 | 338 | with it. | |
| 339 | 339 | ||
| 340 | + | Your agents' git counts the same as yours. Agent runs, checks, reviews, | |
| 341 | + | builds and workflow jobs clone, fetch and push through the same git | |
| 342 | + | endpoints you use, and so does any git command an agent runs itself. Each | |
| 343 | + | counts for the workspace whose repository it is. A pull request's working | |
| 344 | + | copy, where an agent clones and pushes its changes, counts for the | |
| 345 | + | workspace of the repository the pull request is in. g1t's own nightly | |
| 346 | + | backups are never counted for your workspace. | |
| 347 | + | ||
| 340 | 348 | | | Each month (UTC) | | |
| 341 | 349 | | --- | --- | | |
| 342 | 350 | | Free, on every workspace | 50,000 | |
| 235 | 235 | | M12 | Errors | Typed `ArtifactsError`, `rateLimited` events | `ALREADY_EXISTS`/`NOT_FOUND` tolerated; everything else returns 500; no retry, no breaker | Brief Artifacts errors become user-visible failures | R5 | | |
| 236 | 236 | | M13 | Durability | Synchronous replication, asynchronous snapshots; no SLA; no export | No copy outside Artifacts | Beta incident or account issue with no recovery path | R11 | | |
| 237 | 237 | | M14 | Data location | Jurisdiction per namespace, fixed | `PLAN.md` promises residency per workspace; only `g1t` exists | EU customers cannot be offered residency | R7 | | |
| 238 | − | | M15 | Direct credentials | Tokens are bearer, repo-scoped | `git_access` hands out raw write tokens; pushes with them skip branch protection and push protection (`refs_open` only stops caching) | Policy bypass if a token leaks out of a sandbox | Keep TTL minimal on this path | | |
| 238 | + | | M15 | Direct credentials | Tokens are bearer, repo-scoped | `git_access` hands out raw write tokens (no caller is deployed today: sandboxes use g1t's git endpoints, and only backups get a read token); pushes with them skip branch protection and push protection (`refs_open` only stops caching) | Policy bypass if a token leaks out of a sandbox | Keep TTL minimal on this path | | |
| 239 | 239 | | M16 | Hot repository | 2,000 git requests per 10 s per repository; DO soft limit 1,000 req/s | Agents clone the same repository many times per pull request | Not near the limit today; a monorepo with many agents could be | R6 | | |
| 240 | 240 | ||
| 241 | 241 | ## 5. Capacity model: 3,000 workspaces | |
| 363 | 363 | ||
| 364 | 364 | | # | Status | What | | |
| 365 | 365 | | --- | --- | --- | | |
| 366 | − | | R1 | Built; Cloudflare's answer still needed | Every interaction with the store is metered raw (`meters.rs` → `artifacts_meters`, per day, namespace, repository, workspace and meter, with bytes where known): client git (`git.info_refs`, `git.ls_refs`, `git.fetch`, `git.receive_pack`), g1t's own git (`internal.git.*`: landing, catch-up, mirrors, branch listings, fork retirement), every binding call (`binding.get`, `binding.create_token`, `binding.log`, `binding.read_tree`, …, each retry included), and answers g1t served from its own cache (`cache.*`, never operations). Which meters are operations is data: `operation_mapping` (`cost_operations` for g1t's bill, `billable_operations` for workspaces), read every 5 minutes, changed with `set_operation_mapping` without a deploy. Default: `git.fetch`, `git.receive_pack`, `internal.git.fetch`, `internal.git.receive_pack`, `binding.create`, `binding.fork`, `binding.delete` = 1, everything else 0. `git_operations` (what billing reads) is filled from the meters × `billable_operations`, by the hour. RPCs: `artifacts_usage { from, to, workspace?, by_repo? }` (raw meters and the mapping, for the reconciler), `operation_mapping`, `set_operation_mapping { meter, cost_operations, billable_operations, note? }`. Script: `scripts/ops/artifacts-usage.mjs`. | | |
| 366 | + | | R1 | Built; Cloudflare's answer still needed | Every interaction with the store is metered raw (`meters.rs` → `artifacts_meters`, per day, namespace, repository, workspace and meter, with bytes where known): client git (`git.info_refs`, `git.ls_refs`, `git.fetch`, `git.receive_pack`), g1t's own git (`internal.git.*`: landing, catch-up, mirrors, branch listings, fork retirement), every binding call (`binding.get`, `binding.create_token`, `binding.log`, `binding.read_tree`, …, each retry included), and answers g1t served from its own cache (`cache.*`, never operations). Which meters are operations is data: `operation_mapping` (`cost_operations` for g1t's bill, `billable_operations` for workspaces), read every 5 minutes, changed with `set_operation_mapping` without a deploy. Default: `git.fetch`, `git.receive_pack`, `internal.git.fetch`, `internal.git.receive_pack`, `binding.create`, `binding.fork`, `binding.delete` = 1, everything else 0. `git_operations` (what billing reads) is filled from the meters × `billable_operations`, by the hour. RPCs: `artifacts_usage { from, to, workspace?, by_repo? }` (raw meters and the mapping, for the reconciler), `operation_mapping`, `set_operation_mapping { meter, cost_operations, billable_operations, note? }`. Script: `scripts/ops/artifacts-usage.mjs`. Sandboxes' git goes through `git_http` and is metered there; a pull request's working copy counts for its repository's workspace; a request that counts something plans the write that follows in its own `wait_until`, so counts are not stranded in an idle isolate (2026-10-07, see "2026-10-07: where the gap came from"). | | |
| 367 | 367 | | R13 | Built | Nothing on the request path writes D1 for counting. Meters add up per isolate and are written in one batch from `ctx.wait_until` after every request (and at the end of the cron and queue handlers); a failed write is kept for the next. The free-workspace slow-down decides from counts the isolate read back after its last write plus what it added since (`git_ops::standing`, at most 10 minutes old) and billing's plan answer kept 5 minutes. The 63–98 ms `kept` step's D1 upsert is gone. A workspace whose counts this isolate never read is not slowed: nothing slows anyone on a guess. | | |
| 368 | 368 | | R2 | Built; the fork storage test is yours to run | `pull.merged` and `pull.closed` set the fork's `retire_after` (`FORK_RETENTION_DAYS`, 1 day in production; 7 when unset); `pull.reopened` clears it, or makes the fork again. The hourly sweep (`23 * * * *`, 25 a run) keeps the fork's head in its repository as `refs/pull/<pull id>/head` (only missing objects travel; an empty pack when merged), records `retired_at` and `retired_head`, then deletes the fork from the store (a failed delete puts the row back). Reads of a retired fork (the pull request's changes, divergence, tree, blob, log, branches) are answered from the repository with the fork's branch mapped to the kept head (`forks.rs` `Viewed`). Anything that writes or uses git on it (git over HTTPS, `git_access`, catch-up, land, `delete_branch`) makes it again first (`revive`: fork, then move its branch to the head) and schedules it to go again. Work never emits `pull.reopened` today; the handler is ready for it. Script: `scripts/ops/fork-storage-test.mjs`. | | |
| 369 | 369 | | R3 | Built | Credentials g1t uses itself: TTL 3,600 s, reused for 50 minutes (isolate and KV, key `cred2:<key>:<scope>:internal`). `git_access` hands out its own: TTL 300 s, reused 180 s (`…:handout`); `refs_open` still uses 300 s. Every internal path (land, catch-up, mirrors, branch listing, commits, deleting a branch) now reuses kept credentials instead of minting each time. The remote is worked out as `https://<account>.artifacts.cloudflare.net/git/<namespace>/<name>.git` (the documented format, `api/git-protocol`), learned per namespace from the first `info()` an isolate makes, which runs alongside `createToken` and so costs no time; after that a mint is `get` and `createToken`. Optional `ARTIFACTS_REMOTE_BASE` skips even the first `info()`. | | |
| 396 | 396 | `cost_operations` for those meters to 1 (`set_operation_mapping`), and decide whether | |
| 397 | 397 | `billable_operations` follows (cost pass-through says yes). | |
| 398 | 398 | - Every ratio well under 1: Cloudflare counts per clone or fetch session, not per request. | |
| 399 | − | Ratios over 1: something reaches Artifacts that g1t does not meter, such as sandboxes pushing | |
| 400 | − | directly with handed-out credentials. | |
| 399 | + | Ratios over 1: something reaches Artifacts that g1t does not meter, or metered counts were | |
| 400 | + | lost before they were written. Sandboxes are not it: every sandbox but a backup's clones, | |
| 401 | + | fetches and pushes through g1t's git endpoints (see "2026-10-07: where the gap came from"). | |
| 401 | 402 | - Only days after the meters were deployed compare; before that only `git_operations` exists. | |
| 402 | 403 | - Binding calls do appear: Cloudflare's events include `read` and `token_create` actions (and | |
| 403 | 404 | `namespace_*`) besides the five documented ones. If Cloudflare says they are billed, map the | |
| 436 | 437 | - 476 client errors on 2026-10-06 are unexplained; the fetch fix below accounts for some (every | |
| 437 | 438 | failed negotiation was one). | |
| 438 | 439 | ||
| 440 | + | **2026-10-07: where the gap came from.** Cloudflare counted 581 operations (pull 535, push 39, | |
| 441 | + | create 3, fork 4) against g1t's 458 (`git.fetch` 417, `git.receive_pack` 33, ...). The suspicion | |
| 442 | + | was that agents' sandboxes clone and push straight to the store with `git_access` credentials. | |
| 443 | + | They do not. Read from the code: | |
| 444 | + | ||
| 445 | + | - Every sandbox's remote is `https://g1t.sh/<path>.git` (`services/runner/src/index.ts`, | |
| 446 | + | `bump.ts`; Actions' checkout from `services/actions/src/plan.rs`), with a run credential passed | |
| 447 | + | per command (`crates/runner/src/main.rs` `auth_option`, `clone.rs`). Agent runs, answers, | |
| 448 | + | checks, reviews, plans, updates, the merge queue, merge checks, bumps, deploys and workflow | |
| 449 | + | jobs all clone, fetch, deepen (`share_history`) and push through `git_http`, metered as | |
| 450 | + | `git.info_refs`, `git.ls_refs`, `git.fetch` and `git.receive_pack`. Git an agent runs itself | |
| 451 | + | in its sandbox has the same remote and no other credential, so it is metered the same way. | |
| 452 | + | - `git_access` has no caller that is deployed: only `crates/sshd`, whose `/_internal/ssh/*` | |
| 453 | + | endpoints do not exist yet. When git over SSH ships, its bridge talks to the store directly | |
| 454 | + | and must report what it does (as backups do) or go through `git_http`. | |
| 455 | + | - The one sandbox that reads the store directly is a nightly backup (`backups.rs` | |
| 456 | + | `store.handout`): its runner reports the clone with `fetched_bytes`, metered as | |
| 457 | + | `internal.git.info_refs` and `internal.git.backup_fetch` (g1t's cost, never a workspace's). | |
| 458 | + | A clone that fails before reading anything, or a sandbox that dies without reporting, is not | |
| 459 | + | metered. | |
| 460 | + | ||
| 461 | + | Two things were wrong instead, both fixed: | |
| 462 | + | ||
| 463 | + | 1. **Counts left in memory.** An isolate wrote its counts only when a request found the last | |
| 464 | + | write 5 s old. Whatever was counted since stayed in memory until the next request on that | |
| 465 | + | isolate, and was lost if none came: when the isolate went idle, or a deploy replaced it. A | |
| 466 | + | clone is `info/refs`, `ls-refs` and `fetch` within a second, so its last request, the fetch | |
| 467 | + | (the one that is an operation), was the one most often stranded. Now a request that counts | |
| 468 | + | something before a write is due plans that write in its own `wait_until`, waiting until it is | |
| 469 | + | due (`meters::plan_flush`, `flush_after`): still one write per isolate every 5 s at most, and | |
| 470 | + | nothing on the request path. Only an isolate that dies outright loses its last few seconds. | |
| 471 | + | 2. **Working copies counted for nobody.** A pull request's working copy has the path | |
| 472 | + | `pulls/<pull id>`, so everything asked of it (an agent cloning it and pushing to it, checks | |
| 473 | + | and reviews cloning it, catching up, landing's fetch from it, making it with `fork`, removing | |
| 474 | + | it with `delete`) was counted for a workspace called `pulls`, which nobody is charged as. The | |
| 475 | + | counts now go to the workspace of the repository it came from, looked up when they are written | |
| 476 | + | (one query per write, kept 10 minutes), in `artifacts_meters.workspace` and `git_operations` | |
| 477 | + | alike. `artifacts_meters.repo` keeps the working copy's own name. The free-workspace limits | |
| 478 | + | still go by the path asked for, so requests to a working copy are counted but never slowed. | |
| 479 | + | ||
| 480 | + | What can still differ from Cloudflare's count: an isolate that dies with counts in memory; a | |
| 481 | + | store request that fails on g1t's side before it is metered; an `info/refs` GET retried after | |
| 482 | + | the store refused a kept credential (asked twice, metered once; not an operation by default); | |
| 483 | + | backups' unreported clones; and whatever Cloudflare counts that g1t does not ask (its `pull` may | |
| 484 | + | include ref listings, still open). Compare again with `scripts/ops/artifacts-usage.mjs` a day | |
| 485 | + | after this deploys. No migration: `operation_mapping` is unchanged. | |
| 486 | + | ||
| 439 | 487 | ### R2: running and reading `scripts/ops/fork-storage-test.mjs` | |
| 440 | 488 | ||
| 441 | 489 | ```sh |
| 21 | 21 | //! `GIT_OPERATIONS_FREE_HOURLY` (60) an hour, answered 429 with when to try | |
| 22 | 22 | //! again. Whether it is past its cap is decided from counts this isolate | |
| 23 | 23 | //! read a moment ago and has added to since, never by asking the database | |
| 24 | − | //! on the way (63 to 98 ms a request, measured). Pushes from agents' | |
| 25 | − | //! sandboxes go to the store directly and are counted as the store's | |
| 26 | − | //! meters see them, not here. | |
| 24 | + | //! on the way (63 to 98 ms a request, measured). | |
| 25 | + | //! | |
| 26 | + | //! Agents' sandboxes, checks, builds and workflow jobs clone, fetch and | |
| 27 | + | //! push through g1t's git endpoints (`https://g1t.sh/<path>.git`, with a | |
| 28 | + | //! run credential), never the store directly, so they are counted here | |
| 29 | + | //! like anyone's; so is git an agent runs itself in its sandbox. A pull | |
| 30 | + | //! request's working copy counts for the workspace of the repository it | |
| 31 | + | //! came from (meters.rs). The one sandbox that reads the store directly, | |
| 32 | + | //! a nightly backup, reports its clone, which is g1t's cost and never a | |
| 33 | + | //! workspace's (backups.rs). The limits here go by the path asked for, so | |
| 34 | + | //! a free workspace's requests to a working copy (`pulls/<pull id>`) are | |
| 35 | + | //! counted for it but never slowed down. | |
| 27 | 36 | ||
| 28 | 37 | use std::cell::RefCell; | |
| 29 | 38 | use std::collections::HashMap; |
| 2072 | 2072 | } | |
| 2073 | 2073 | ||
| 2074 | 2074 | /// Writes what this isolate metered once the answer has gone back, every | |
| 2075 | − | /// few seconds at most (meters.rs). | |
| 2075 | + | /// few seconds at most: now, or once it is due, waiting in this request's | |
| 2076 | + | /// `wait_until` so nothing counted is left for a request that may never | |
| 2077 | + | /// come (meters.rs). | |
| 2076 | 2078 | fn flush_later(env: &Env, ctx: &Context) { | |
| 2077 | − | if !meters::take_due() { | |
| 2079 | + | let Some(wait) = meters::plan_flush() else { | |
| 2078 | 2080 | return; | |
| 2079 | − | } | |
| 2081 | + | }; | |
| 2080 | 2082 | if let Ok(db) = env.d1("DB") { | |
| 2081 | − | ctx.wait_until(async move { meters::flush(&db).await }); | |
| 2083 | + | ctx.wait_until(async move { meters::flush_after(&db, wait).await }); | |
| 2082 | 2084 | } | |
| 2083 | 2085 | } | |
| 2084 | 2086 |
| 10 | 10 | //! and changed without a deploy. What a workspace is counted for goes to | |
| 11 | 11 | //! `git_operations` by the hour, as before, which billing reads. | |
| 12 | 12 | //! | |
| 13 | + | //! Sandboxes (agents, checks, builds, workflow jobs) use git like anyone | |
| 14 | + | //! else: through g1t's git endpoints with a run credential, so their | |
| 15 | + | //! clones, fetches and pushes are metered here as `git.*`, whoever runs | |
| 16 | + | //! them in the sandbox, the agent included. Only a nightly backup's clone | |
| 17 | + | //! goes to the store directly (backups.rs), and its sandbox reports it. | |
| 18 | + | //! What is asked of a pull request's working copy (`pulls--<pull id>`, | |
| 19 | + | //! where agents clone and push) is counted for the workspace of the | |
| 20 | + | //! repository it came from, looked up when the counts are written | |
| 21 | + | //! (`Pending::attributed`); before 2026-10-07 it was counted for a | |
| 22 | + | //! workspace called `pulls`, which nobody is charged as. | |
| 23 | + | //! | |
| 13 | 24 | //! The same place keeps how the store answered, by the minute | |
| 14 | 25 | //! (`store_health`), for the status page's "Git storage" part. | |
| 15 | 26 | //! | |
| 16 | 27 | //! Nothing here is on the request path. Each isolate adds up what it saw in | |
| 17 | 28 | //! memory, and writes it all in one batch once the answer has gone back | |
| 18 | − | //! (`flush`, from `ctx.wait_until`). A failed write puts the counts back | |
| 19 | − | //! for the next one. What an isolate holds when it is evicted is lost: a | |
| 20 | − | //! few seconds' worth at most, since every request ends with a flush. | |
| 29 | + | //! (`flush`, from `ctx.wait_until`), every few seconds at most. A request | |
| 30 | + | //! that counts something before the next write is due plans that write in | |
| 31 | + | //! its own `wait_until`, which waits until it is (`plan_flush`, | |
| 32 | + | //! `flush_after`): what was counted is never left for a later request on | |
| 33 | + | //! the same isolate, which may never come. Before 2026-10-07 it was, and a | |
| 34 | + | //! clone's last request (its fetch, the one that is an operation) was the | |
| 35 | + | //! one most often lost when the isolate then went idle or a deploy | |
| 36 | + | //! replaced it. A failed write puts the counts back for the next one. What | |
| 37 | + | //! an isolate holds when it dies outright is lost: a few seconds' worth at | |
| 38 | + | //! most. | |
| 21 | 39 | ||
| 22 | 40 | use std::cell::RefCell; | |
| 23 | 41 | use std::collections::HashMap; | |
| 126 | 144 | } | |
| 127 | 145 | } | |
| 128 | 146 | ||
| 147 | + | /// The pull requests whose working copies (`pulls--<pull id>`) are | |
| 148 | + | /// counted for nobody's workspace yet: see [`Pending::attributed`]. | |
| 149 | + | pub fn unattributed_pulls(&self) -> Vec<String> { | |
| 150 | + | let mut pulls: Vec<String> = self | |
| 151 | + | .usage | |
| 152 | + | .keys() | |
| 153 | + | .filter(|place| place.workspace == crate::PULLS_NAMESPACE) | |
| 154 | + | .filter_map(|place| working_copy(&place.repo).map(str::to_owned)) | |
| 155 | + | .collect(); | |
| 156 | + | pulls.sort(); | |
| 157 | + | pulls.dedup(); | |
| 158 | + | pulls | |
| 159 | + | } | |
| 160 | + | ||
| 161 | + | /// The same counts with each pull request's working copy counted for | |
| 162 | + | /// the workspace of the repository it came from, as `owners` (pull id | |
| 163 | + | /// to workspace) says. A working copy's path is `pulls/<pull id>`, so | |
| 164 | + | /// what is asked of it (an agent's clone of it and its pushes to it, a | |
| 165 | + | /// merge check, catching up, making and removing it) would otherwise be | |
| 166 | + | /// counted for a workspace called `pulls`, which nobody is charged as. | |
| 167 | + | /// One `owners` does not name stays as it was. | |
| 168 | + | pub fn attributed(&self, owners: &HashMap<String, String>) -> Pending { | |
| 169 | + | let mut out = Pending { usage: HashMap::with_capacity(self.usage.len()), health: self.health.clone() }; | |
| 170 | + | for (place, tally) in &self.usage { | |
| 171 | + | let mut place = place.clone(); | |
| 172 | + | if place.workspace == crate::PULLS_NAMESPACE | |
| 173 | + | && let Some(owner) = working_copy(&place.repo).and_then(|pull| owners.get(pull)) | |
| 174 | + | { | |
| 175 | + | place.workspace = owner.clone(); | |
| 176 | + | } | |
| 177 | + | out.add(place, tally.count, tally.bytes_in, tally.bytes_out); | |
| 178 | + | } | |
| 179 | + | out | |
| 180 | + | } | |
| 181 | + | ||
| 129 | 182 | /// What each workspace is counted for, by the hour, under `mapping`. | |
| 130 | 183 | pub fn billable(&self, mapping: &Mapping) -> HashMap<(String, String), f64> { | |
| 131 | 184 | let mut out: HashMap<(String, String), f64> = HashMap::new(); | |
| 232 | 285 | .unwrap_or_else(|| name.split_once("--").map(|(workspace, _)| workspace.to_owned()).unwrap_or_default()) | |
| 233 | 286 | } | |
| 234 | 287 | ||
| 288 | + | /// The pull id of a working copy's name in the store (`pulls--pul_7`). | |
| 289 | + | fn working_copy(name: &str) -> Option<&str> { | |
| 290 | + | name.strip_prefix(crate::PULLS_NAMESPACE)?.strip_prefix("--").filter(|pull| !pull.is_empty()) | |
| 291 | + | } | |
| 292 | + | ||
| 293 | + | /// How long the workspace a working copy is counted for is kept before it | |
| 294 | + | /// is read again (a workspace can be renamed). | |
| 295 | + | const PULL_OWNER_TTL_MS: u64 = 10 * 60 * 1000; | |
| 296 | + | ||
| 297 | + | thread_local! { | |
| 298 | + | /// The workspace each pull request's working copy is counted for, by | |
| 299 | + | /// pull id, as last read, and when. | |
| 300 | + | static PULL_OWNERS: RefCell<HashMap<String, (String, u64)>> = RefCell::new(HashMap::new()); | |
| 301 | + | } | |
| 302 | + | ||
| 303 | + | /// The workspace of the repository each of `pulls`' working copies came | |
| 304 | + | /// from: as kept for a while, else read in one query. Off the request | |
| 305 | + | /// path (from `flush`). | |
| 306 | + | async fn pull_owners(db: &D1Database, pulls: &[String]) -> Result<HashMap<String, String>> { | |
| 307 | + | let now = g1t_kit::now_ms(); | |
| 308 | + | let mut owners = HashMap::new(); | |
| 309 | + | let mut missing = Vec::new(); | |
| 310 | + | PULL_OWNERS.with(|kept| { | |
| 311 | + | let kept = kept.borrow(); | |
| 312 | + | for pull in pulls { | |
| 313 | + | match kept.get(pull).filter(|(_, at)| now.saturating_sub(*at) < PULL_OWNER_TTL_MS) { | |
| 314 | + | Some((owner, _)) => { | |
| 315 | + | owners.insert(pull.clone(), owner.clone()); | |
| 316 | + | } | |
| 317 | + | None => missing.push(pull.clone()), | |
| 318 | + | } | |
| 319 | + | } | |
| 320 | + | }); | |
| 321 | + | if missing.is_empty() { | |
| 322 | + | return Ok(owners); | |
| 323 | + | } | |
| 324 | + | #[derive(Deserialize)] | |
| 325 | + | struct Row { | |
| 326 | + | pull: String, | |
| 327 | + | workspace: String, | |
| 328 | + | } | |
| 329 | + | let rows = db | |
| 330 | + | .prepare( | |
| 331 | + | "SELECT f.name AS pull, s.namespace AS workspace FROM repos f JOIN repos s ON s.id = f.fork_of | |
| 332 | + | WHERE f.namespace = ?1 AND f.name IN (SELECT value FROM json_each(?2))", | |
| 333 | + | ) | |
| 334 | + | .bind(&[crate::PULLS_NAMESPACE.into(), serde_json::to_string(&missing)?.into()])? | |
| 335 | + | .all() | |
| 336 | + | .await? | |
| 337 | + | .results::<Row>()?; | |
| 338 | + | PULL_OWNERS.with(|kept| { | |
| 339 | + | let mut kept = kept.borrow_mut(); | |
| 340 | + | for row in &rows { | |
| 341 | + | kept.insert(row.pull.clone(), (row.workspace.clone(), now)); | |
| 342 | + | } | |
| 343 | + | }); | |
| 344 | + | owners.extend(rows.into_iter().map(|row| (row.pull, row.workspace))); | |
| 345 | + | Ok(owners) | |
| 346 | + | } | |
| 347 | + | ||
| 235 | 348 | /// Counts one `meter` for the repository stored under `key`, with the | |
| 236 | 349 | /// bytes it sent and received. | |
| 237 | 350 | pub fn record(meter: &str, key: &str, bytes_in: u64, bytes_out: u64) { | |
| 282 | 395 | /// not each need one. | |
| 283 | 396 | const FLUSH_EVERY_MS: u64 = 5_000; | |
| 284 | 397 | const FLUSH_AT_PLACES: usize = 200; | |
| 398 | + | /// A planned write not begun after this long is taken as never coming (its | |
| 399 | + | /// request's `wait_until` was cut short), so another is planned. | |
| 400 | + | const PLAN_STALE_MS: u64 = 30_000; | |
| 285 | 401 | ||
| 286 | 402 | thread_local! { | |
| 287 | 403 | static LAST_FLUSH: std::cell::Cell<u64> = const { std::cell::Cell::new(0) }; | |
| 404 | + | /// When the write now waiting in some request's `wait_until` was planned. | |
| 405 | + | static PLANNED: std::cell::Cell<Option<u64>> = const { std::cell::Cell::new(None) }; | |
| 288 | 406 | } | |
| 289 | 407 | ||
| 290 | − | /// Whether a write is due: something waits, and the last write was a | |
| 291 | − | /// while ago or much has piled up. | |
| 292 | − | pub fn flush_due(waiting: usize, last: u64, now: u64) -> bool { | |
| 293 | − | waiting > 0 && (now.saturating_sub(last) >= FLUSH_EVERY_MS || waiting >= FLUSH_AT_PLACES) | |
| 408 | + | /// How long the write a request should plan waits, in milliseconds, or | |
| 409 | + | /// `None` when it should plan none: nothing waits, or a write is planned | |
| 410 | + | /// already and has not gone stale. A pile-up is written at once, planned | |
| 411 | + | /// write or not. | |
| 412 | + | pub fn plan(waiting: usize, last: u64, planned: Option<u64>, now: u64) -> Option<u64> { | |
| 413 | + | if waiting == 0 { | |
| 414 | + | return None; | |
| 415 | + | } | |
| 416 | + | if waiting >= FLUSH_AT_PLACES { | |
| 417 | + | return Some(0); | |
| 418 | + | } | |
| 419 | + | if planned.is_some_and(|at| now.saturating_sub(at) < PLAN_STALE_MS) { | |
| 420 | + | return None; | |
| 421 | + | } | |
| 422 | + | Some(FLUSH_EVERY_MS.saturating_sub(now.saturating_sub(last))) | |
| 294 | 423 | } | |
| 295 | 424 | ||
| 296 | − | /// Whether this isolate should write what it counted now; if so, the | |
| 297 | − | /// write is taken as begun. | |
| 298 | − | pub fn take_due() -> bool { | |
| 425 | + | /// The write this request should plan, if any, as how long it waits (see | |
| 426 | + | /// [`plan`]); taken as planned. Never waits itself. | |
| 427 | + | pub fn plan_flush() -> Option<u64> { | |
| 299 | 428 | let now = g1t_kit::now_ms(); | |
| 300 | 429 | let waiting = PENDING.with(|pending| { | |
| 301 | 430 | let pending = pending.borrow(); | |
| 302 | 431 | pending.usage.len() + pending.health.len() | |
| 303 | 432 | }); | |
| 304 | − | let due = flush_due(waiting, LAST_FLUSH.with(std::cell::Cell::get), now); | |
| 305 | − | if due { | |
| 306 | − | LAST_FLUSH.with(|last| last.set(now)); | |
| 433 | + | let wait = plan(waiting, LAST_FLUSH.with(std::cell::Cell::get), PLANNED.with(std::cell::Cell::get), now)?; | |
| 434 | + | if wait > 0 { | |
| 435 | + | PLANNED.with(|planned| planned.set(Some(now))); | |
| 436 | + | } | |
| 437 | + | Some(wait) | |
| 438 | + | } | |
| 439 | + | ||
| 440 | + | /// A planned write: waits `wait_ms`, then writes everything counted by | |
| 441 | + | /// then. For `ctx.wait_until`, after the answer has gone back. | |
| 442 | + | pub async fn flush_after(db: &D1Database, wait_ms: u64) { | |
| 443 | + | if wait_ms > 0 { | |
| 444 | + | worker::Delay::from(std::time::Duration::from_millis(wait_ms)).await; | |
| 445 | + | PLANNED.with(|planned| planned.set(None)); | |
| 307 | 446 | } | |
| 308 | − | due | |
| 447 | + | LAST_FLUSH.with(|last| last.set(g1t_kit::now_ms())); | |
| 448 | + | flush(db).await; | |
| 309 | 449 | } | |
| 310 | 450 | ||
| 311 | 451 | /// The mapping as this isolate last read it, else the defaults. Never | |
| 372 | 512 | return; | |
| 373 | 513 | } | |
| 374 | 514 | let mapping = mapping(db).await; | |
| 375 | − | if let Err(error) = write(db, &taken, &mapping).await { | |
| 515 | + | // Pull requests' working copies count for their repositories' workspaces. | |
| 516 | + | let pulls = taken.unattributed_pulls(); | |
| 517 | + | let attributed = if pulls.is_empty() { | |
| 518 | + | None | |
| 519 | + | } else { | |
| 520 | + | match pull_owners(db, &pulls).await { | |
| 521 | + | Ok(owners) => Some(taken.attributed(&owners)), | |
| 522 | + | Err(error) => { | |
| 523 | + | worker::console_error!("working copies' workspaces not read, counted as they are: {error}"); | |
| 524 | + | None | |
| 525 | + | } | |
| 526 | + | } | |
| 527 | + | }; | |
| 528 | + | if let Err(error) = write(db, attributed.as_ref().unwrap_or(&taken), &mapping).await { | |
| 376 | 529 | worker::console_error!("git store meters not written, kept for the next try: {error}"); | |
| 377 | 530 | PENDING.with(|pending| pending.borrow_mut().merge(taken)); | |
| 378 | 531 | } | |
| 572 | 725 | ||
| 573 | 726 | #[test] | |
| 574 | 727 | fn counts_are_written_now_and_then_not_on_every_request() { | |
| 575 | − | assert!(!flush_due(0, 0, 100_000)); | |
| 576 | − | assert!(flush_due(1, 0, 100_000)); | |
| 577 | − | assert!(!flush_due(3, 100_000, 100_000 + FLUSH_EVERY_MS - 1)); | |
| 578 | − | assert!(flush_due(3, 100_000, 100_000 + FLUSH_EVERY_MS)); | |
| 579 | − | assert!(flush_due(FLUSH_AT_PLACES, 100_000, 100_001)); | |
| 728 | + | // At once only when the last write was a while ago or much has | |
| 729 | + | // piled up; otherwise once it is due. | |
| 730 | + | assert_eq!(plan(1, 0, None, 100_000), Some(0)); | |
| 731 | + | assert_eq!(plan(3, 100_000, None, 100_000 + FLUSH_EVERY_MS - 1), Some(1)); | |
| 732 | + | assert_eq!(plan(3, 100_000, None, 100_000 + FLUSH_EVERY_MS), Some(0)); | |
| 733 | + | assert_eq!(plan(FLUSH_AT_PLACES, 100_000, None, 100_001), Some(0)); | |
| 734 | + | } | |
| 735 | + | ||
| 736 | + | #[test] | |
| 737 | + | fn what_a_request_counts_is_written_even_if_no_request_follows() { | |
| 738 | + | // Nothing waiting: nothing planned. | |
| 739 | + | assert_eq!(plan(0, 0, None, 100_000), None); | |
| 740 | + | // Due: written at once. | |
| 741 | + | assert_eq!(plan(1, 0, None, 100_000), Some(0)); | |
| 742 | + | // Counted a second after the last write: the write is planned for | |
| 743 | + | // when it is due, not left for the next request. | |
| 744 | + | assert_eq!(plan(3, 100_000, None, 101_000), Some(FLUSH_EVERY_MS - 1_000)); | |
| 745 | + | // One planned already: the requests after it plan none... | |
| 746 | + | assert_eq!(plan(3, 100_000, Some(101_000), 102_000), None); | |
| 747 | + | // ...unless much has piled up, which is written at once, | |
| 748 | + | assert_eq!(plan(FLUSH_AT_PLACES, 100_000, Some(101_000), 102_000), Some(0)); | |
| 749 | + | // or the planned one never began (its request was cut short). | |
| 750 | + | assert_eq!(plan(3, 100_000, Some(101_000), 101_000 + PLAN_STALE_MS), Some(0)); | |
| 751 | + | // Never planned further off than the interval. | |
| 752 | + | assert!(plan(1, 100_000, None, 100_000).is_some_and(|wait| wait <= FLUSH_EVERY_MS)); | |
| 580 | 753 | } | |
| 581 | 754 | ||
| 582 | 755 | #[test] | |
| 617 | 790 | } | |
| 618 | 791 | ||
| 619 | 792 | #[test] | |
| 793 | + | fn a_working_copy_is_counted_for_its_repositorys_workspace() { | |
| 794 | + | let copy = |meter: &str, pull: &str| Place { | |
| 795 | + | hour: "2026-10-14T09".into(), | |
| 796 | + | store: "g1t".into(), | |
| 797 | + | repo: format!("pulls--{pull}"), | |
| 798 | + | workspace: "pulls".into(), | |
| 799 | + | meter: meter.into(), | |
| 800 | + | }; | |
| 801 | + | let mut pending = Pending::default(); | |
| 802 | + | // A pull request's working copy is made, an agent's sandbox clones | |
| 803 | + | // it and pushes to it, and checks clone it again; another pull | |
| 804 | + | // request's repository is not found. | |
| 805 | + | pending.meter(copy("binding.fork", "pul_7"), 0, 0); | |
| 806 | + | pending.meter(copy("git.fetch", "pul_7"), 100, 9_000); | |
| 807 | + | pending.meter(copy("git.receive_pack", "pul_7"), 4_000, 50); | |
| 808 | + | pending.meter(copy("git.fetch", "pul_7"), 100, 9_000); | |
| 809 | + | pending.meter(copy("git.fetch", "pul_8"), 0, 0); | |
| 810 | + | pending.meter(place("2026-10-14T09", "acme", "git.fetch"), 0, 0); | |
| 811 | + | pending.health("g1t", "2026-10-14T09:01", Outcome::Ok, 5); | |
| 812 | + | assert_eq!(pending.unattributed_pulls(), ["pul_7", "pul_8"]); | |
| 813 | + | // As they are: counted for a workspace called `pulls`. | |
| 814 | + | let before = pending.billable(&Mapping::defaults()); | |
| 815 | + | assert_eq!(before[&("pulls".to_owned(), "2026-10-14T09".to_owned())], 5.0); | |
| 816 | + | ||
| 817 | + | let owners = HashMap::from([("pul_7".to_owned(), "acme".to_owned())]); | |
| 818 | + | let attributed = pending.attributed(&owners); | |
| 819 | + | let billable = attributed.billable(&Mapping::defaults()); | |
| 820 | + | assert_eq!(billable[&("acme".to_owned(), "2026-10-14T09".to_owned())], 5.0); | |
| 821 | + | assert_eq!(billable[&("pulls".to_owned(), "2026-10-14T09".to_owned())], 1.0); | |
| 822 | + | // The meters keep the working copy's own name, with its workspace. | |
| 823 | + | let mut counted = copy("git.fetch", "pul_7"); | |
| 824 | + | counted.workspace = "acme".into(); | |
| 825 | + | assert_eq!(attributed.usage[&counted], Tally { count: 2, bytes_in: 200, bytes_out: 18_000 }); | |
| 826 | + | assert_eq!(attributed.health.len(), 1); | |
| 827 | + | assert_eq!(attributed.unattributed_pulls(), ["pul_8"]); | |
| 828 | + | // Nothing else is a working copy. | |
| 829 | + | assert_eq!(working_copy("pulls--pul_7"), Some("pul_7")); | |
| 830 | + | assert_eq!(working_copy("pulls--"), None); | |
| 831 | + | assert_eq!(working_copy("pullsx--pul_7"), None); | |
| 832 | + | assert_eq!(working_copy("acme--pulls"), None); | |
| 833 | + | } | |
| 834 | + | ||
| 835 | + | #[test] | |
| 620 | 836 | fn health_counts_failures_and_rejections_apart() { | |
| 621 | 837 | let mut pending = Pending::default(); | |
| 622 | 838 | pending.health("g1t", "2026-10-14T09:01", Outcome::Ok, 40); |
| 30 | 30 | /// minutes, so each one used has at least ten minutes left. | |
| 31 | 31 | const INTERNAL_TTL_SECONDS: u32 = 3_600; | |
| 32 | 32 | const INTERNAL_REUSE_MS: u64 = 50 * 60 * 1000; | |
| 33 | − | /// Credentials handed out (`git_access`, to sandboxes): five minutes, used | |
| 34 | − | /// for three, so whoever gets one has at least two. | |
| 33 | + | /// Credentials handed out: to a nightly backup's sandbox (backups.rs), and | |
| 34 | + | /// by `git_access`, which nothing deployed asks yet (git over SSH will). | |
| 35 | + | /// Other sandboxes never get one: they use g1t's git endpoints. Five | |
| 36 | + | /// minutes, used for three, so whoever gets one has at least two. | |
| 35 | 37 | const HANDOUT_TTL_SECONDS: u32 = 300; | |
| 36 | 38 | const HANDOUT_REUSE_MS: u64 = 180_000; | |
| 37 | 39 | /// How long a handed-out credential stays valid, in milliseconds. |