Merge branch 'worktree-agent-a57ff9fecefa1eaf7'
7 files+321−340/7 viewed
| 186 | 186 | ||
| 187 | 187 | ### Git operations | |
| 188 | 188 | ||
| 189 | − | Each clone, fetch and push is a git operation. Every workspace has 50,000 | |
| 189 | + | Each clone, fetch and push is a git operation, your agents' included: | |
| 190 | + | their sandboxes use the same git endpoints you do, and a pull request's | |
| 191 | + | working copy counts for its repository's workspace. Every workspace has 50,000 | |
| 190 | 192 | a month included. Past that, a workspace on the g1t plan pays $0.18 per | |
| 191 | 193 | 1,000, and a free workspace is never charged: past 50,000 in a month, its | |
| 192 | 194 | git requests past 60 in an hour are answered `429` with when to try again, |
| 337 | 337 | what counts follows what Cloudflare confirms it bills; this page changes | |
| 338 | 338 | with it. | |
| 339 | 339 | ||
| 340 | + | Your agents' git counts the same as yours. Agent runs, checks, reviews, | |
| 341 | + | builds and workflow jobs clone, fetch and push through the same git | |
| 342 | + | endpoints you use, and so does any git command an agent runs itself. Each | |
| 343 | + | counts for the workspace whose repository it is. A pull request's working | |
| 344 | + | copy, where an agent clones and pushes its changes, counts for the | |
| 345 | + | workspace of the repository the pull request is in. g1t's own nightly | |
| 346 | + | backups are never counted for your workspace. | |
| 347 | + | ||
| 340 | 348 | | | Each month (UTC) | | |
| 341 | 349 | | --- | --- | | |
| 342 | 350 | | Free, on every workspace | 50,000 | |
| 235 | 235 | | M12 | Errors | Typed `ArtifactsError`, `rateLimited` events | `ALREADY_EXISTS`/`NOT_FOUND` tolerated; everything else returns 500; no retry, no breaker | Brief Artifacts errors become user-visible failures | R5 | | |
| 236 | 236 | | M13 | Durability | Synchronous replication, asynchronous snapshots; no SLA; no export | No copy outside Artifacts | Beta incident or account issue with no recovery path | R11 | | |
| 237 | 237 | | M14 | Data location | Jurisdiction per namespace, fixed | `PLAN.md` promises residency per workspace; only `g1t` exists | EU customers cannot be offered residency | R7 | | |
| 238 | − | | M15 | Direct credentials | Tokens are bearer, repo-scoped | `git_access` hands out raw write tokens; pushes with them skip branch protection and push protection (`refs_open` only stops caching) | Policy bypass if a token leaks out of a sandbox | Keep TTL minimal on this path | | |
| 238 | + | | M15 | Direct credentials | Tokens are bearer, repo-scoped | `git_access` hands out raw write tokens (no caller is deployed today: sandboxes use g1t's git endpoints, and only backups get a read token); pushes with them skip branch protection and push protection (`refs_open` only stops caching) | Policy bypass if a token leaks out of a sandbox | Keep TTL minimal on this path | | |
| 239 | 239 | | M16 | Hot repository | 2,000 git requests per 10 s per repository; DO soft limit 1,000 req/s | Agents clone the same repository many times per pull request | Not near the limit today; a monorepo with many agents could be | R6 | | |
| 240 | 240 | ||
| 241 | 241 | ## 5. Capacity model: 3,000 workspaces | |
| 363 | 363 | ||
| 364 | 364 | | # | Status | What | | |
| 365 | 365 | | --- | --- | --- | | |
| 366 | − | | R1 | Built; Cloudflare's answer still needed | Every interaction with the store is metered raw (`meters.rs` → `artifacts_meters`, per day, namespace, repository, workspace and meter, with bytes where known): client git (`git.info_refs`, `git.ls_refs`, `git.fetch`, `git.receive_pack`), g1t's own git (`internal.git.*`: landing, catch-up, mirrors, branch listings, fork retirement), every binding call (`binding.get`, `binding.create_token`, `binding.log`, `binding.read_tree`, …, each retry included), and answers g1t served from its own cache (`cache.*`, never operations). Which meters are operations is data: `operation_mapping` (`cost_operations` for g1t's bill, `billable_operations` for workspaces), read every 5 minutes, changed with `set_operation_mapping` without a deploy. Default: `git.fetch`, `git.receive_pack`, `internal.git.fetch`, `internal.git.receive_pack`, `binding.create`, `binding.fork`, `binding.delete` = 1, everything else 0. `git_operations` (what billing reads) is filled from the meters × `billable_operations`, by the hour. RPCs: `artifacts_usage { from, to, workspace?, by_repo? }` (raw meters and the mapping, for the reconciler), `operation_mapping`, `set_operation_mapping { meter, cost_operations, billable_operations, note? }`. Script: `scripts/ops/artifacts-usage.mjs`. | | |
| 366 | + | | R1 | Built; Cloudflare's answer still needed | Every interaction with the store is metered raw (`meters.rs` → `artifacts_meters`, per day, namespace, repository, workspace and meter, with bytes where known): client git (`git.info_refs`, `git.ls_refs`, `git.fetch`, `git.receive_pack`), g1t's own git (`internal.git.*`: landing, catch-up, mirrors, branch listings, fork retirement), every binding call (`binding.get`, `binding.create_token`, `binding.log`, `binding.read_tree`, …, each retry included), and answers g1t served from its own cache (`cache.*`, never operations). Which meters are operations is data: `operation_mapping` (`cost_operations` for g1t's bill, `billable_operations` for workspaces), read every 5 minutes, changed with `set_operation_mapping` without a deploy. Default: `git.fetch`, `git.receive_pack`, `internal.git.fetch`, `internal.git.receive_pack`, `binding.create`, `binding.fork`, `binding.delete` = 1, everything else 0. `git_operations` (what billing reads) is filled from the meters × `billable_operations`, by the hour. RPCs: `artifacts_usage { from, to, workspace?, by_repo? }` (raw meters and the mapping, for the reconciler), `operation_mapping`, `set_operation_mapping { meter, cost_operations, billable_operations, note? }`. Script: `scripts/ops/artifacts-usage.mjs`. Sandboxes' git goes through `git_http` and is metered there; a pull request's working copy counts for its repository's workspace; a request that counts something plans the write that follows in its own `wait_until`, so counts are not stranded in an idle isolate (2026-10-07, see "2026-10-07: where the gap came from"). | | |
| 367 | 367 | | R13 | Built | Nothing on the request path writes D1 for counting. Meters add up per isolate and are written in one batch from `ctx.wait_until` after every request (and at the end of the cron and queue handlers); a failed write is kept for the next. The free-workspace slow-down decides from counts the isolate read back after its last write plus what it added since (`git_ops::standing`, at most 10 minutes old) and billing's plan answer kept 5 minutes. The 63–98 ms `kept` step's D1 upsert is gone. A workspace whose counts this isolate never read is not slowed: nothing slows anyone on a guess. | | |
| 368 | 368 | | R2 | Built; the fork storage test is yours to run | `pull.merged` and `pull.closed` set the fork's `retire_after` (`FORK_RETENTION_DAYS`, 1 day in production; 7 when unset); `pull.reopened` clears it, or makes the fork again. The hourly sweep (`23 * * * *`, 25 a run) keeps the fork's head in its repository as `refs/pull/<pull id>/head` (only missing objects travel; an empty pack when merged), records `retired_at` and `retired_head`, then deletes the fork from the store (a failed delete puts the row back). Reads of a retired fork (the pull request's changes, divergence, tree, blob, log, branches) are answered from the repository with the fork's branch mapped to the kept head (`forks.rs` `Viewed`). Anything that writes or uses git on it (git over HTTPS, `git_access`, catch-up, land, `delete_branch`) makes it again first (`revive`: fork, then move its branch to the head) and schedules it to go again. Work never emits `pull.reopened` today; the handler is ready for it. Script: `scripts/ops/fork-storage-test.mjs`. | | |
| 369 | 369 | | R3 | Built | Credentials g1t uses itself: TTL 3,600 s, reused for 50 minutes (isolate and KV, key `cred2:<key>:<scope>:internal`). `git_access` hands out its own: TTL 300 s, reused 180 s (`…:handout`); `refs_open` still uses 300 s. Every internal path (land, catch-up, mirrors, branch listing, commits, deleting a branch) now reuses kept credentials instead of minting each time. The remote is worked out as `https://<account>.artifacts.cloudflare.net/git/<namespace>/<name>.git` (the documented format, `api/git-protocol`), learned per namespace from the first `info()` an isolate makes, which runs alongside `createToken` and so costs no time; after that a mint is `get` and `createToken`. Optional `ARTIFACTS_REMOTE_BASE` skips even the first `info()`. | | |
| 396 | 396 | `cost_operations` for those meters to 1 (`set_operation_mapping`), and decide whether | |
| 397 | 397 | `billable_operations` follows (cost pass-through says yes). | |
| 398 | 398 | - Every ratio well under 1: Cloudflare counts per clone or fetch session, not per request. | |
| 399 | − | Ratios over 1: something reaches Artifacts that g1t does not meter, such as sandboxes pushing | |
| 400 | − | directly with handed-out credentials. | |
| 399 | + | Ratios over 1: something reaches Artifacts that g1t does not meter, or metered counts were | |
| 400 | + | lost before they were written. Sandboxes are not it: every sandbox but a backup's clones, | |
| 401 | + | fetches and pushes through g1t's git endpoints (see "2026-10-07: where the gap came from"). | |
| 401 | 402 | - Only days after the meters were deployed compare; before that only `git_operations` exists. | |
| 402 | 403 | - Binding calls do appear: Cloudflare's events include `read` and `token_create` actions (and | |
| 403 | 404 | `namespace_*`) besides the five documented ones. If Cloudflare says they are billed, map the | |
| 436 | 437 | - 476 client errors on 2026-10-06 are unexplained; the fetch fix below accounts for some (every | |
| 437 | 438 | failed negotiation was one). | |
| 438 | 439 | ||
| 440 | + | **2026-10-07: where the gap came from.** Cloudflare counted 581 operations (pull 535, push 39, | |
| 441 | + | create 3, fork 4) against g1t's 458 (`git.fetch` 417, `git.receive_pack` 33, ...). The suspicion | |
| 442 | + | was that agents' sandboxes clone and push straight to the store with `git_access` credentials. | |
| 443 | + | They do not. Read from the code: | |
| 444 | + | ||
| 445 | + | - Every sandbox's remote is `https://g1t.sh/<path>.git` (`services/runner/src/index.ts`, | |
| 446 | + | `bump.ts`; Actions' checkout from `services/actions/src/plan.rs`), with a run credential passed | |
| 447 | + | per command (`crates/runner/src/main.rs` `auth_option`, `clone.rs`). Agent runs, answers, | |
| 448 | + | checks, reviews, plans, updates, the merge queue, merge checks, bumps, deploys and workflow | |
| 449 | + | jobs all clone, fetch, deepen (`share_history`) and push through `git_http`, metered as | |
| 450 | + | `git.info_refs`, `git.ls_refs`, `git.fetch` and `git.receive_pack`. Git an agent runs itself | |
| 451 | + | in its sandbox has the same remote and no other credential, so it is metered the same way. | |
| 452 | + | - `git_access` has no caller that is deployed: only `crates/sshd`, whose `/_internal/ssh/*` | |
| 453 | + | endpoints do not exist yet. When git over SSH ships, its bridge talks to the store directly | |
| 454 | + | and must report what it does (as backups do) or go through `git_http`. | |
| 455 | + | - The one sandbox that reads the store directly is a nightly backup (`backups.rs` | |
| 456 | + | `store.handout`): its runner reports the clone with `fetched_bytes`, metered as | |
| 457 | + | `internal.git.info_refs` and `internal.git.backup_fetch` (g1t's cost, never a workspace's). | |
| 458 | + | A clone that fails before reading anything, or a sandbox that dies without reporting, is not | |
| 459 | + | metered. | |
| 460 | + | ||
| 461 | + | Two things were wrong instead, both fixed: | |
| 462 | + | ||
| 463 | + | 1. **Counts left in memory.** An isolate wrote its counts only when a request found the last | |
| 464 | + | write 5 s old. Whatever was counted since stayed in memory until the next request on that | |
| 465 | + | isolate, and was lost if none came: when the isolate went idle, or a deploy replaced it. A | |
| 466 | + | clone is `info/refs`, `ls-refs` and `fetch` within a second, so its last request, the fetch | |
| 467 | + | (the one that is an operation), was the one most often stranded. Now a request that counts | |
| 468 | + | something before a write is due plans that write in its own `wait_until`, waiting until it is | |
| 469 | + | due (`meters::plan_flush`, `flush_after`): still one write per isolate every 5 s at most, and | |
| 470 | + | nothing on the request path. Only an isolate that dies outright loses its last few seconds. | |
| 471 | + | 2. **Working copies counted for nobody.** A pull request's working copy has the path | |
| 472 | + | `pulls/<pull id>`, so everything asked of it (an agent cloning it and pushing to it, checks | |
| 473 | + | and reviews cloning it, catching up, landing's fetch from it, making it with `fork`, removing | |
| 474 | + | it with `delete`) was counted for a workspace called `pulls`, which nobody is charged as. The | |
| 475 | + | counts now go to the workspace of the repository it came from, looked up when they are written | |
| 476 | + | (one query per write, kept 10 minutes), in `artifacts_meters.workspace` and `git_operations` | |
| 477 | + | alike. `artifacts_meters.repo` keeps the working copy's own name. The free-workspace limits | |
| 478 | + | still go by the path asked for, so requests to a working copy are counted but never slowed. | |
| 479 | + | ||
| 480 | + | What can still differ from Cloudflare's count: an isolate that dies with counts in memory; a | |
| 481 | + | store request that fails on g1t's side before it is metered; an `info/refs` GET retried after | |
| 482 | + | the store refused a kept credential (asked twice, metered once; not an operation by default); | |
| 483 | + | backups' unreported clones; and whatever Cloudflare counts that g1t does not ask (its `pull` may | |
| 484 | + | include ref listings, still open). Compare again with `scripts/ops/artifacts-usage.mjs` a day | |
| 485 | + | after this deploys. No migration: `operation_mapping` is unchanged. | |
| 486 | + | ||
| 439 | 487 | ### R2: running and reading `scripts/ops/fork-storage-test.mjs` | |
| 440 | 488 | ||
| 441 | 489 | ```sh |
| 21 | 21 | //! `GIT_OPERATIONS_FREE_HOURLY` (60) an hour, answered 429 with when to try | |
| 22 | 22 | //! again. Whether it is past its cap is decided from counts this isolate | |
| 23 | 23 | //! read a moment ago and has added to since, never by asking the database | |
| 24 | − | //! on the way (63 to 98 ms a request, measured). Pushes from agents' | |
| 25 | − | //! sandboxes go to the store directly and are counted as the store's | |
| 26 | − | //! meters see them, not here. | |
| 24 | + | //! on the way (63 to 98 ms a request, measured). | |
| 25 | + | //! | |
| 26 | + | //! Agents' sandboxes, checks, builds and workflow jobs clone, fetch and | |
| 27 | + | //! push through g1t's git endpoints (`https://g1t.sh/<path>.git`, with a | |
| 28 | + | //! run credential), never the store directly, so they are counted here | |
| 29 | + | //! like anyone's; so is git an agent runs itself in its sandbox. A pull | |
| 30 | + | //! request's working copy counts for the workspace of the repository it | |
| 31 | + | //! came from (meters.rs). The one sandbox that reads the store directly, | |
| 32 | + | //! a nightly backup, reports its clone, which is g1t's cost and never a | |
| 33 | + | //! workspace's (backups.rs). The limits here go by the path asked for, so | |
| 34 | + | //! a free workspace's requests to a working copy (`pulls/<pull id>`) are | |
| 35 | + | //! counted for it but never slowed down. | |
| 27 | 36 | ||
| 28 | 37 | use std::cell::RefCell; | |
| 29 | 38 | use std::collections::HashMap; |
| 2072 | 2072 | } | |
| 2073 | 2073 | ||
| 2074 | 2074 | /// Writes what this isolate metered once the answer has gone back, every | |
| 2075 | − | /// few seconds at most (meters.rs). | |
| 2075 | + | /// few seconds at most: now, or once it is due, waiting in this request's | |
| 2076 | + | /// `wait_until` so nothing counted is left for a request that may never | |
| 2077 | + | /// come (meters.rs). | |
| 2076 | 2078 | fn flush_later(env: &Env, ctx: &Context) { | |
| 2077 | − | if !meters::take_due() { | |
| 2079 | + | let Some(wait) = meters::plan_flush() else { | |
| 2078 | 2080 | return; | |
| 2079 | − | } | |
| 2081 | + | }; | |
| 2080 | 2082 | if let Ok(db) = env.d1("DB") { | |
| 2081 | − | ctx.wait_until(async move { meters::flush(&db).await }); | |
| 2083 | + | ctx.wait_until(async move { meters::flush_after(&db, wait).await }); | |
| 2082 | 2084 | } | |
| 2083 | 2085 | } | |
| 2084 | 2086 |
| 10 | 10 | //! and changed without a deploy. What a workspace is counted for goes to | |
| 11 | 11 | //! `git_operations` by the hour, as before, which billing reads. | |
| 12 | 12 | //! | |
| 13 | + | //! Sandboxes (agents, checks, builds, workflow jobs) use git like anyone | |
| 14 | + | //! else: through g1t's git endpoints with a run credential, so their | |
| 15 | + | //! clones, fetches and pushes are metered here as `git.*`, whoever runs | |
| 16 | + | //! them in the sandbox, the agent included. Only a nightly backup's clone | |
| 17 | + | //! goes to the store directly (backups.rs), and its sandbox reports it. | |
| 18 | + | //! What is asked of a pull request's working copy (`pulls--<pull id>`, | |
| 19 | + | //! where agents clone and push) is counted for the workspace of the | |
| 20 | + | //! repository it came from, looked up when the counts are written | |
| 21 | + | //! (`Pending::attributed`); before 2026-10-07 it was counted for a | |
| 22 | + | //! workspace called `pulls`, which nobody is charged as. | |
| 23 | + | //! | |
| 13 | 24 | //! The same place keeps how the store answered, by the minute | |
| 14 | 25 | //! (`store_health`), for the status page's "Git storage" part. | |
| 15 | 26 | //! | |
| 16 | 27 | //! Nothing here is on the request path. Each isolate adds up what it saw in | |
| 17 | 28 | //! memory, and writes it all in one batch once the answer has gone back | |
| 18 | − | //! (`flush`, from `ctx.wait_until`). A failed write puts the counts back | |
| 19 | − | //! for the next one. What an isolate holds when it is evicted is lost: a | |
| 20 | − | //! few seconds' worth at most, since every request ends with a flush. | |
| 29 | + | //! (`flush`, from `ctx.wait_until`), every few seconds at most. A request | |
| 30 | + | //! that counts something before the next write is due plans that write in | |
| 31 | + | //! its own `wait_until`, which waits until it is (`plan_flush`, | |
| 32 | + | //! `flush_after`): what was counted is never left for a later request on | |
| 33 | + | //! the same isolate, which may never come. Before 2026-10-07 it was, and a | |
| 34 | + | //! clone's last request (its fetch, the one that is an operation) was the | |
| 35 | + | //! one most often lost when the isolate then went idle or a deploy | |
| 36 | + | //! replaced it. A failed write puts the counts back for the next one. What | |
| 37 | + | //! an isolate holds when it dies outright is lost: a few seconds' worth at | |
| 38 | + | //! most. | |
| 21 | 39 | ||
| 22 | 40 | use std::cell::RefCell; | |
| 23 | 41 | use std::collections::HashMap; | |
| 126 | 144 | } | |
| 127 | 145 | } | |
| 128 | 146 | ||
| 147 | + | /// The pull requests whose working copies (`pulls--<pull id>`) are | |
| 148 | + | /// counted for nobody's workspace yet: see [`Pending::attributed`]. | |
| 149 | + | pub fn unattributed_pulls(&self) -> Vec<String> { | |
| 150 | + | let mut pulls: Vec<String> = self | |
| 151 | + | .usage | |
| 152 | + | .keys() | |
| 153 | + | .filter(|place| place.workspace == crate::PULLS_NAMESPACE) | |
| 154 | + | .filter_map(|place| working_copy(&place.repo).map(str::to_owned)) | |
| 155 | + | .collect(); | |
| 156 | + | pulls.sort(); | |
| 157 | + | pulls.dedup(); | |
| 158 | + | pulls | |
| 159 | + | } | |
| 160 | + | ||
| 161 | + | /// The same counts with each pull request's working copy counted for | |
| 162 | + | /// the workspace of the repository it came from, as `owners` (pull id | |
| 163 | + | /// to workspace) says. A working copy's path is `pulls/<pull id>`, so | |
| 164 | + | /// what is asked of it (an agent's clone of it and its pushes to it, a | |
| 165 | + | /// merge check, catching up, making and removing it) would otherwise be | |
| 166 | + | /// counted for a workspace called `pulls`, which nobody is charged as. | |
| 167 | + | /// One `owners` does not name stays as it was. | |
| 168 | + | pub fn attributed(&self, owners: &HashMap<String, String>) -> Pending { | |
| 169 | + | let mut out = Pending { usage: HashMap::with_capacity(self.usage.len()), health: self.health.clone() }; | |
| 170 | + | for (place, tally) in &self.usage { | |
| 171 | + | let mut place = place.clone(); | |
| 172 | + | if place.workspace == crate::PULLS_NAMESPACE | |
| 173 | + | && let Some(owner) = working_copy(&place.repo).and_then(|pull| owners.get(pull)) | |
| 174 | + | { | |
| 175 | + | place.workspace = owner.clone(); | |
| 176 | + | } | |
| 177 | + | out.add(place, tally.count, tally.bytes_in, tally.bytes_out); | |
| 178 | + | } | |
| 179 | + | out | |
| 180 | + | } | |
| 181 | + | ||
| 129 | 182 | /// What each workspace is counted for, by the hour, under `mapping`. | |
| 130 | 183 | pub fn billable(&self, mapping: &Mapping) -> HashMap<(String, String), f64> { | |
| 131 | 184 | let mut out: HashMap<(String, String), f64> = HashMap::new(); | |
| 232 | 285 | .unwrap_or_else(|| name.split_once("--").map(|(workspace, _)| workspace.to_owned()).unwrap_or_default()) | |
| 233 | 286 | } | |
| 234 | 287 | ||
| 288 | + | /// The pull id of a working copy's name in the store (`pulls--pul_7`). | |
| 289 | + | fn working_copy(name: &str) -> Option<&str> { | |
| 290 | + | name.strip_prefix(crate::PULLS_NAMESPACE)?.strip_prefix("--").filter(|pull| !pull.is_empty()) | |
| 291 | + | } | |
| 292 | + | ||
| 293 | + | /// How long the workspace a working copy is counted for is kept before it | |
| 294 | + | /// is read again (a workspace can be renamed). | |
| 295 | + | const PULL_OWNER_TTL_MS: u64 = 10 * 60 * 1000; | |
| 296 | + | ||
| 297 | + | thread_local! { | |
| 298 | + | /// The workspace each pull request's working copy is counted for, by | |
| 299 | + | /// pull id, as last read, and when. | |
| 300 | + | static PULL_OWNERS: RefCell<HashMap<String, (String, u64)>> = RefCell::new(HashMap::new()); | |
| 301 | + | } | |
| 302 | + | ||
| 303 | + | /// The workspace of the repository each of `pulls`' working copies came | |
| 304 | + | /// from: as kept for a while, else read in one query. Off the request | |
| 305 | + | /// path (from `flush`). | |
| 306 | + | async fn pull_owners(db: &D1Database, pulls: &[String]) -> Result<HashMap<String, String>> { | |
| 307 | + | let now = g1t_kit::now_ms(); | |
| 308 | + | let mut owners = HashMap::new(); | |
| 309 | + | let mut missing = Vec::new(); | |
| 310 | + | PULL_OWNERS.with(|kept| { | |
| 311 | + | let kept = kept.borrow(); | |
| 312 | + | for pull in pulls { | |
| 313 | + | match kept.get(pull).filter(|(_, at)| now.saturating_sub(*at) < PULL_OWNER_TTL_MS) { | |
| 314 | + | Some((owner, _)) => { | |
| 315 | + | owners.insert(pull.clone(), owner.clone()); | |
| 316 | + | } | |
| 317 | + | None => missing.push(pull.clone()), | |
| 318 | + | } | |
| 319 | + | } | |
| 320 | + | }); | |
| 321 | + | if missing.is_empty() { | |
| 322 | + | return Ok(owners); | |
| 323 | + | } | |
| 324 | + | #[derive(Deserialize)] | |
| 325 | + | struct Row { | |
| 326 | + | pull: String, | |
| 327 | + | workspace: String, | |
| 328 | + | } | |
| 329 | + | let rows = db | |
| 330 | + | .prepare( | |
| 331 | + | "SELECT f.name AS pull, s.namespace AS workspace FROM repos f JOIN repos s ON s.id = f.fork_of | |
| 332 | + | WHERE f.namespace = ?1 AND f.name IN (SELECT value FROM json_each(?2))", | |
| 333 | + | ) | |
| 334 | + | .bind(&[crate::PULLS_NAMESPACE.into(), serde_json::to_string(&missing)?.into()])? | |
| 335 | + | .all() | |
| 336 | + | .await? | |
| 337 | + | .results::<Row>()?; | |
| 338 | + | PULL_OWNERS.with(|kept| { | |
| 339 | + | let mut kept = kept.borrow_mut(); | |
| 340 | + | for row in &rows { | |
| 341 | + | kept.insert(row.pull.clone(), (row.workspace.clone(), now)); | |
| 342 | + | } | |
| 343 | + | }); | |
| 344 | + | owners.extend(rows.into_iter().map(|row| (row.pull, row.workspace))); | |
| 345 | + | Ok(owners) | |
| 346 | + | } | |
| 347 | + | ||
| 235 | 348 | /// Counts one `meter` for the repository stored under `key`, with the | |
| 236 | 349 | /// bytes it sent and received. | |
| 237 | 350 | pub fn record(meter: &str, key: &str, bytes_in: u64, bytes_out: u64) { | |
| 282 | 395 | /// not each need one. | |
| 283 | 396 | const FLUSH_EVERY_MS: u64 = 5_000; | |
| 284 | 397 | const FLUSH_AT_PLACES: usize = 200; | |
| 398 | + | /// A planned write not begun after this long is taken as never coming (its | |
| 399 | + | /// request's `wait_until` was cut short), so another is planned. | |
| 400 | + | const PLAN_STALE_MS: u64 = 30_000; | |
| 285 | 401 | ||
| 286 | 402 | thread_local! { | |
| 287 | 403 | static LAST_FLUSH: std::cell::Cell<u64> = const { std::cell::Cell::new(0) }; | |
| 404 | + | /// When the write now waiting in some request's `wait_until` was planned. | |
| 405 | + | static PLANNED: std::cell::Cell<Option<u64>> = const { std::cell::Cell::new(None) }; | |
| 288 | 406 | } | |
| 289 | 407 | ||
| 290 | − | /// Whether a write is due: something waits, and the last write was a | |
| 291 | − | /// while ago or much has piled up. | |
| 292 | − | pub fn flush_due(waiting: usize, last: u64, now: u64) -> bool { | |
| 293 | − | waiting > 0 && (now.saturating_sub(last) >= FLUSH_EVERY_MS || waiting >= FLUSH_AT_PLACES) | |
| 408 | + | /// How long the write a request should plan waits, in milliseconds, or | |
| 409 | + | /// `None` when it should plan none: nothing waits, or a write is planned | |
| 410 | + | /// already and has not gone stale. A pile-up is written at once, planned | |
| 411 | + | /// write or not. | |
| 412 | + | pub fn plan(waiting: usize, last: u64, planned: Option<u64>, now: u64) -> Option<u64> { | |
| 413 | + | if waiting == 0 { | |
| 414 | + | return None; | |
| 415 | + | } | |
| 416 | + | if waiting >= FLUSH_AT_PLACES { | |
| 417 | + | return Some(0); | |
| 418 | + | } | |
| 419 | + | if planned.is_some_and(|at| now.saturating_sub(at) < PLAN_STALE_MS) { | |
| 420 | + | return None; | |
| 421 | + | } | |
| 422 | + | Some(FLUSH_EVERY_MS.saturating_sub(now.saturating_sub(last))) | |
| 294 | 423 | } | |
| 295 | 424 | ||
| 296 | − | /// Whether this isolate should write what it counted now; if so, the | |
| 297 | − | /// write is taken as begun. | |
| 298 | − | pub fn take_due() -> bool { | |
| 425 | + | /// The write this request should plan, if any, as how long it waits (see | |
| 426 | + | /// [`plan`]); taken as planned. Never waits itself. | |
| 427 | + | pub fn plan_flush() -> Option<u64> { | |
| 299 | 428 | let now = g1t_kit::now_ms(); | |
| 300 | 429 | let waiting = PENDING.with(|pending| { | |
| 301 | 430 | let pending = pending.borrow(); | |
| 302 | 431 | pending.usage.len() + pending.health.len() | |
| 303 | 432 | }); | |
| 304 | − | let due = flush_due(waiting, LAST_FLUSH.with(std::cell::Cell::get), now); | |
| 305 | − | if due { | |
| 306 | − | LAST_FLUSH.with(|last| last.set(now)); | |
| 433 | + | let wait = plan(waiting, LAST_FLUSH.with(std::cell::Cell::get), PLANNED.with(std::cell::Cell::get), now)?; | |
| 434 | + | if wait > 0 { | |
| 435 | + | PLANNED.with(|planned| planned.set(Some(now))); | |
| 436 | + | } | |
| 437 | + | Some(wait) | |
| 438 | + | } | |
| 439 | + | ||
| 440 | + | /// A planned write: waits `wait_ms`, then writes everything counted by | |
| 441 | + | /// then. For `ctx.wait_until`, after the answer has gone back. | |
| 442 | + | pub async fn flush_after(db: &D1Database, wait_ms: u64) { | |
| 443 | + | if wait_ms > 0 { | |
| 444 | + | worker::Delay::from(std::time::Duration::from_millis(wait_ms)).await; | |
| 445 | + | PLANNED.with(|planned| planned.set(None)); | |
| 307 | 446 | } | |
| 308 | − | due | |
| 447 | + | LAST_FLUSH.with(|last| last.set(g1t_kit::now_ms())); | |
| 448 | + | flush(db).await; | |
| 309 | 449 | } | |
| 310 | 450 | ||
| 311 | 451 | /// The mapping as this isolate last read it, else the defaults. Never | |
| 372 | 512 | return; | |
| 373 | 513 | } | |
| 374 | 514 | let mapping = mapping(db).await; | |
| 375 | − | if let Err(error) = write(db, &taken, &mapping).await { | |
| 515 | + | // Pull requests' working copies count for their repositories' workspaces. | |
| 516 | + | let pulls = taken.unattributed_pulls(); | |
| 517 | + | let attributed = if pulls.is_empty() { | |
| 518 | + | None | |
| 519 | + | } else { | |
| 520 | + | match pull_owners(db, &pulls).await { | |
| 521 | + | Ok(owners) => Some(taken.attributed(&owners)), | |
| 522 | + | Err(error) => { | |
| 523 | + | worker::console_error!("working copies' workspaces not read, counted as they are: {error}"); | |
| 524 | + | None | |
| 525 | + | } | |
| 526 | + | } | |
| 527 | + | }; | |
| 528 | + | if let Err(error) = write(db, attributed.as_ref().unwrap_or(&taken), &mapping).await { | |
| 376 | 529 | worker::console_error!("git store meters not written, kept for the next try: {error}"); | |
| 377 | 530 | PENDING.with(|pending| pending.borrow_mut().merge(taken)); | |
| 378 | 531 | } | |
| 572 | 725 | ||
| 573 | 726 | #[test] | |
| 574 | 727 | fn counts_are_written_now_and_then_not_on_every_request() { | |
| 575 | − | assert!(!flush_due(0, 0, 100_000)); | |
| 576 | − | assert!(flush_due(1, 0, 100_000)); | |
| 577 | − | assert!(!flush_due(3, 100_000, 100_000 + FLUSH_EVERY_MS - 1)); | |
| 578 | − | assert!(flush_due(3, 100_000, 100_000 + FLUSH_EVERY_MS)); | |
| 579 | − | assert!(flush_due(FLUSH_AT_PLACES, 100_000, 100_001)); | |
| 728 | + | // At once only when the last write was a while ago or much has | |
| 729 | + | // piled up; otherwise once it is due. | |
| 730 | + | assert_eq!(plan(1, 0, None, 100_000), Some(0)); | |
| 731 | + | assert_eq!(plan(3, 100_000, None, 100_000 + FLUSH_EVERY_MS - 1), Some(1)); | |
| 732 | + | assert_eq!(plan(3, 100_000, None, 100_000 + FLUSH_EVERY_MS), Some(0)); | |
| 733 | + | assert_eq!(plan(FLUSH_AT_PLACES, 100_000, None, 100_001), Some(0)); | |
| 734 | + | } | |
| 735 | + | ||
| 736 | + | #[test] | |
| 737 | + | fn what_a_request_counts_is_written_even_if_no_request_follows() { | |
| 738 | + | // Nothing waiting: nothing planned. | |
| 739 | + | assert_eq!(plan(0, 0, None, 100_000), None); | |
| 740 | + | // Due: written at once. | |
| 741 | + | assert_eq!(plan(1, 0, None, 100_000), Some(0)); | |
| 742 | + | // Counted a second after the last write: the write is planned for | |
| 743 | + | // when it is due, not left for the next request. | |
| 744 | + | assert_eq!(plan(3, 100_000, None, 101_000), Some(FLUSH_EVERY_MS - 1_000)); | |
| 745 | + | // One planned already: the requests after it plan none... | |
| 746 | + | assert_eq!(plan(3, 100_000, Some(101_000), 102_000), None); | |
| 747 | + | // ...unless much has piled up, which is written at once, | |
| 748 | + | assert_eq!(plan(FLUSH_AT_PLACES, 100_000, Some(101_000), 102_000), Some(0)); | |
| 749 | + | // or the planned one never began (its request was cut short). | |
| 750 | + | assert_eq!(plan(3, 100_000, Some(101_000), 101_000 + PLAN_STALE_MS), Some(0)); | |
| 751 | + | // Never planned further off than the interval. | |
| 752 | + | assert!(plan(1, 100_000, None, 100_000).is_some_and(|wait| wait <= FLUSH_EVERY_MS)); | |
| 580 | 753 | } | |
| 581 | 754 | ||
| 582 | 755 | #[test] | |
| 617 | 790 | } | |
| 618 | 791 | ||
| 619 | 792 | #[test] | |
| 793 | + | fn a_working_copy_is_counted_for_its_repositorys_workspace() { | |
| 794 | + | let copy = |meter: &str, pull: &str| Place { | |
| 795 | + | hour: "2026-10-14T09".into(), | |
| 796 | + | store: "g1t".into(), | |
| 797 | + | repo: format!("pulls--{pull}"), | |
| 798 | + | workspace: "pulls".into(), | |
| 799 | + | meter: meter.into(), | |
| 800 | + | }; | |
| 801 | + | let mut pending = Pending::default(); | |
| 802 | + | // A pull request's working copy is made, an agent's sandbox clones | |
| 803 | + | // it and pushes to it, and checks clone it again; another pull | |
| 804 | + | // request's repository is not found. | |
| 805 | + | pending.meter(copy("binding.fork", "pul_7"), 0, 0); | |
| 806 | + | pending.meter(copy("git.fetch", "pul_7"), 100, 9_000); | |
| 807 | + | pending.meter(copy("git.receive_pack", "pul_7"), 4_000, 50); | |
| 808 | + | pending.meter(copy("git.fetch", "pul_7"), 100, 9_000); | |
| 809 | + | pending.meter(copy("git.fetch", "pul_8"), 0, 0); | |
| 810 | + | pending.meter(place("2026-10-14T09", "acme", "git.fetch"), 0, 0); | |
| 811 | + | pending.health("g1t", "2026-10-14T09:01", Outcome::Ok, 5); | |
| 812 | + | assert_eq!(pending.unattributed_pulls(), ["pul_7", "pul_8"]); | |
| 813 | + | // As they are: counted for a workspace called `pulls`. | |
| 814 | + | let before = pending.billable(&Mapping::defaults()); | |
| 815 | + | assert_eq!(before[&("pulls".to_owned(), "2026-10-14T09".to_owned())], 5.0); | |
| 816 | + | ||
| 817 | + | let owners = HashMap::from([("pul_7".to_owned(), "acme".to_owned())]); | |
| 818 | + | let attributed = pending.attributed(&owners); | |
| 819 | + | let billable = attributed.billable(&Mapping::defaults()); | |
| 820 | + | assert_eq!(billable[&("acme".to_owned(), "2026-10-14T09".to_owned())], 5.0); | |
| 821 | + | assert_eq!(billable[&("pulls".to_owned(), "2026-10-14T09".to_owned())], 1.0); | |
| 822 | + | // The meters keep the working copy's own name, with its workspace. | |
| 823 | + | let mut counted = copy("git.fetch", "pul_7"); | |
| 824 | + | counted.workspace = "acme".into(); | |
| 825 | + | assert_eq!(attributed.usage[&counted], Tally { count: 2, bytes_in: 200, bytes_out: 18_000 }); | |
| 826 | + | assert_eq!(attributed.health.len(), 1); | |
| 827 | + | assert_eq!(attributed.unattributed_pulls(), ["pul_8"]); | |
| 828 | + | // Nothing else is a working copy. | |
| 829 | + | assert_eq!(working_copy("pulls--pul_7"), Some("pul_7")); | |
| 830 | + | assert_eq!(working_copy("pulls--"), None); | |
| 831 | + | assert_eq!(working_copy("pullsx--pul_7"), None); | |
| 832 | + | assert_eq!(working_copy("acme--pulls"), None); | |
| 833 | + | } | |
| 834 | + | ||
| 835 | + | #[test] | |
| 620 | 836 | fn health_counts_failures_and_rejections_apart() { | |
| 621 | 837 | let mut pending = Pending::default(); | |
| 622 | 838 | pending.health("g1t", "2026-10-14T09:01", Outcome::Ok, 40); |
| 30 | 30 | /// minutes, so each one used has at least ten minutes left. | |
| 31 | 31 | const INTERNAL_TTL_SECONDS: u32 = 3_600; | |
| 32 | 32 | const INTERNAL_REUSE_MS: u64 = 50 * 60 * 1000; | |
| 33 | − | /// Credentials handed out (`git_access`, to sandboxes): five minutes, used | |
| 34 | − | /// for three, so whoever gets one has at least two. | |
| 33 | + | /// Credentials handed out: to a nightly backup's sandbox (backups.rs), and | |
| 34 | + | /// by `git_access`, which nothing deployed asks yet (git over SSH will). | |
| 35 | + | /// Other sandboxes never get one: they use g1t's git endpoints. Five | |
| 36 | + | /// minutes, used for three, so whoever gets one has at least two. | |
| 35 | 37 | const HANDOUT_TTL_SECONDS: u32 = 300; | |
| 36 | 38 | const HANDOUT_REUSE_MS: u64 = 180_000; | |
| 37 | 39 | /// How long a handed-out credential stays valid, in milliseconds. |