Merge memory from docs: only docs on how to work here, whole sentences, no near-duplicates, stale doc suggestions cleared
14 files+900−750/14 viewed
| 149 | 149 | ||
| 150 | 150 | Memory also fills itself. At the end of every run that changes code, the | |
| 151 | 151 | agent is asked what it learned; a person's correction in a review, a merged | |
| 152 | − | pull request's decision, and what a project's `AGENTS.md`, README and | |
| 153 | − | manifests say are captured too. These arrive as **candidates**, which no | |
| 152 | + | pull request's decision, and what a project's `AGENTS.md`, README, | |
| 153 | + | CONTRIBUTING, docs on how to work in it and manifests say are captured | |
| 154 | + | too ([which docs](/guides/context-hub/#which-docs-are-read-for-memory)). | |
| 155 | + | These arrive as **candidates**, which no | |
| 154 | 156 | agent is given until they are kept: at once when two independent sources | |
| 155 | 157 | say the same thing or a project's `AGENTS.md` or manifests state it, | |
| 156 | 158 | otherwise by a member in the **Review** list on **Agents → Memory** or the | |
| ⋯ | |||
| 190 | 192 | - **edit** one that has drifted, or change its kind; | |
| 191 | 193 | - **forget** one that no longer holds; | |
| 192 | 194 | - **keep**, **edit** or **dismiss** a candidate waiting for review. A | |
| 193 | − | dismissed candidate is never suggested again in the same words. | |
| 195 | + | dismissed candidate is never suggested again, in the same words or near enough to them. | |
| 194 | 196 | ||
| 195 | 197 | Each shows who added it, when, and when it was last given to an agent. A | |
| 196 | 198 | memory that has not been given to an agent for a long time is a good one to | |
| 36 | 36 | | `README`, `AGENTS.md` (or `CLAUDE.md`), `CONTRIBUTING`, `docs/*.md`, `runbooks/*.md` | **Docs**, split into sections for search | | |
| 37 | 37 | | `.g1t/workflows/*`, `.github/workflows/*` | Whether the project's tests run in checks | | |
| 38 | 38 | | `owners:` in `.g1t/project.yml`, and `CODEOWNERS` | **Owners** | | |
| 39 | + | | `memory:` in `.g1t/project.yml` | Which docs are [read for memory](#which-docs-are-read-for-memory) | | |
| 39 | 40 | ||
| 40 | 41 | and joins them with what g1t already knows: | |
| 41 | 42 | ||
| ⋯ | |||
| 76 | 77 | | **Agent runs** | At the end of every run that changes code, the agent is asked what it learned that the next agent would need, with what showed it | fact, convention, decision or gotcha | | |
| 77 | 78 | | **Reviews** | A person's request for changes, or a comment that corrects the agent ("we use the shared client instead"), on a pull request | convention, quoting the comment | | |
| 78 | 79 | | **Merges** | A merged pull request's title, why (the first paragraph of its description) and the files it changed | decision | | |
| 79 | − | | **Docs and manifests** | Bullets in `AGENTS.md`; commands under a README's setup and testing sections; conventions sections; the package manager and test commands from manifests | fact, convention or gotcha | | |
| 80 | + | | **Docs and manifests** | Bullets in `AGENTS.md`; commands and bullets under the setup, testing and conventions sections of docs on how to work in the project; the package manager and test commands from manifests | fact, convention or gotcha | | |
| 80 | 81 | ||
| 81 | 82 | What is captured arrives as a **candidate**. No agent is given a candidate | |
| 82 | 83 | until it is **kept**: | |
| ⋯ | |||
| 87 | 88 | - **by a person**, in the Review queue. | |
| 88 | 89 | ||
| 89 | 90 | The same thing said again in different case, punctuation or spacing counts | |
| 90 | − | as the same memory, seen once more. A memory seen again from the same | |
| 91 | + | as the same memory, seen once more. So does the same thing in slightly | |
| 92 | + | different words: a sentence that grew (a list with one more item), or one | |
| 93 | + | that holds all of another's words. A memory seen again from the same | |
| 91 | 94 | source (the same run, the same file) is not counted twice. | |
| 92 | 95 | ||
| 96 | + | ### Which docs are read for memory | |
| 97 | + | ||
| 98 | + | Memory is for how to work in a project, so only docs that say how are | |
| 99 | + | read for it: | |
| 100 | + | ||
| 101 | + | | Read for memory | Not read for memory | | |
| 102 | + | | --- | --- | | |
| 103 | + | | `README`, `AGENTS.md` (or `CLAUDE.md`), `CONTRIBUTING`, `runbooks/*.md` | Plans and roadmaps (`PLAN.md`, `roadmap.md`) | | |
| 104 | + | | A doc in `docs/` whose name says how to work: `DEPLOYING.md`, `testing.md`, `architecture.md`, `setup.md`, `conventions.md`, `getting-started.md` | Changelogs, release notes, history, incidents, demos, research, notes, feedback | | |
| 105 | + | ||
| 106 | + | Within those docs: | |
| 107 | + | ||
| 108 | + | - Only the **setup, testing and conventions** sections are read (in | |
| 109 | + | `AGENTS.md`, every section). A section about plans, a roadmap, asks or | |
| 110 | + | feedback is skipped with everything under it, wherever it is. | |
| 111 | + | - A doc that reads as a plan or a report (most of its headings are plans, | |
| 112 | + | or its bullets are labelled "What we tried", "Ask" and the like) gives | |
| 113 | + | nothing, whatever its name. | |
| 114 | + | - A line that says what someone wants rather than how things are ("g1t | |
| 115 | + | will", "should", "TODO") is left out. In `AGENTS.md`, "should" states a | |
| 116 | + | rule, and is kept. | |
| 117 | + | - A bullet is taken whole, with the lines it wraps onto. A long one is cut | |
| 118 | + | only between sentences, to 300 characters; one whose first sentence is | |
| 119 | + | longer is left out rather than cut. | |
| 120 | + | - A bullet that tells you to do or never do something is a convention or | |
| 121 | + | a gotcha; anything else is a fact. Outside `AGENTS.md`, a bullet waits | |
| 122 | + | for review at 50 to 60% sure. | |
| 123 | + | ||
| 124 | + | To read another doc for memory, or never read one, say so in the | |
| 125 | + | project's `.g1t/project.yml`, by path or a whole folder: | |
| 126 | + | ||
| 127 | + | ```yaml | |
| 128 | + | memory: | |
| 129 | + | docs: | |
| 130 | + | - docs/PERFORMANCE.md | |
| 131 | + | skip: | |
| 132 | + | - README.md | |
| 133 | + | - docs/legacy/* | |
| 134 | + | ``` | |
| 135 | + | ||
| 136 | + | `skip` wins over `docs`. A change to these lists applies to each doc the | |
| 137 | + | next time it changes, or for every doc at once with **Rebuild**. | |
| 138 | + | ||
| 139 | + | Once g1t has read every file of a project as these rules read it, a | |
| 140 | + | candidate from its docs that is still waiting and that its docs no longer | |
| 141 | + | suggest (the doc changed, or the rules leave the line out) is removed from | |
| 142 | + | the queue. A candidate another source saw too, and every kept or dismissed | |
| 143 | + | memory, stays. A waiting candidate its doc now says in other words takes | |
| 144 | + | the new words. | |
| 145 | + | ||
| 93 | 146 | ### Review | |
| 94 | 147 | ||
| 95 | 148 | The **Memory** tab of the Context page lists every candidate in the | |
| ⋯ | |||
| 99 | 152 | ||
| 100 | 153 | - **Keep** gives it to every agent from their next run on. | |
| 101 | 154 | - **Edit** rewords it, or changes its kind, and keeps it. | |
| 102 | − | - **Dismiss** throws it away, and the same wording is never suggested | |
| 103 | − | again. | |
| 155 | + | - **Dismiss** throws it away, and the same thing, in the same words or | |
| 156 | + | near enough to them, is never suggested again. | |
| 104 | 157 | ||
| 105 | 158 | ### Never a secret | |
| 106 | 159 | ||
| 1 | 1 | import assert from "node:assert/strict"; | |
| 2 | 2 | import { test } from "node:test"; | |
| 3 | 3 | ||
| 4 | − | import { distinctFacts } from "./memory-facts.ts"; | |
| 4 | + | import { distinctFacts, sameFact } from "./memory-facts.ts"; | |
| 5 | 5 | ||
| 6 | 6 | test("a fact remembered twice is shown once, the first kept", () => { | |
| 7 | 7 | const facts = [ | |
| ⋯ | |||
| 13 | 13 | assert.deepEqual(distinctFacts(facts).map((fact) => fact.id), ["a", "b"]); | |
| 14 | 14 | assert.deepEqual(distinctFacts([]), []); | |
| 15 | 15 | }); | |
| 16 | + | ||
| 17 | + | test("near-duplicates collapse: a list that grew, a sentence cut short", () => { | |
| 18 | + | const facts = [ | |
| 19 | + | { id: "a", text: "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing); `cargo test` runs its tests." }, | |
| 20 | + | { id: "b", text: "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing, services/work); `cargo test` runs its tests." }, | |
| 21 | + | { id: "c", text: "Roles. Viewer, commenter, planner and approver map onto" }, | |
| 22 | + | { id: "d", text: "Roles. Viewer, commenter, planner and approver map onto the five repository roles." }, | |
| 23 | + | { id: "e", text: "Use npm to install." }, | |
| 24 | + | { id: "f", text: "Use pnpm to install." }, | |
| 25 | + | ]; | |
| 26 | + | assert.deepEqual(distinctFacts(facts).map((fact) => fact.id), ["a", "c", "e", "f"]); | |
| 27 | + | assert.ok(!sameFact("Run cargo test in the crate you changed.", "Run npm test in the app you changed.")); | |
| 28 | + | assert.ok(!sameFact("use pnpm", "use pnpm in the web app and npm in the docs, which predates it")); | |
| 29 | + | }); | |
| 1 | + | /** How alike two facts' words must be (shared over all, as sets) to be one fact. */ | |
| 2 | + | const SAME_WORDS = 0.85; | |
| 3 | + | /** The fewest distinct words a fact needs before it can be found inside another. */ | |
| 4 | + | const CONTAINED_MIN_WORDS = 6; | |
| 5 | + | ||
| 6 | + | /** A fact's words, lowercase, one space apart: case, punctuation and spacing aside. */ | |
| 7 | + | function folded(text: string): string { | |
| 8 | + | return text | |
| 9 | + | .toLowerCase() | |
| 10 | + | .split(/[^\p{L}\p{N}]+/u) | |
| 11 | + | .filter(Boolean) | |
| 12 | + | .join(" "); | |
| 13 | + | } | |
| 14 | + | ||
| 15 | + | /** | |
| 16 | + | * Whether two facts say the same thing, as the work service decides it | |
| 17 | + | * (`same_memory`): the same words; one's words, in order, inside the | |
| 18 | + | * other's; all of one's words (at least six) among the other's; or most of | |
| 19 | + | * their words shared. | |
| 20 | + | */ | |
| 21 | + | export function sameFact(a: string, b: string): boolean { | |
| 22 | + | const [x, y] = [folded(a), folded(b)]; | |
| 23 | + | if (!x || !y) return false; | |
| 24 | + | if (x === y) return true; | |
| 25 | + | const [short, long] = x.length <= y.length ? [x, y] : [y, x]; | |
| 26 | + | if (short.split(" ").length >= 4 && ` ${long} `.includes(` ${short} `)) return true; | |
| 27 | + | const shortWords = new Set(short.split(" ")); | |
| 28 | + | const longWords = new Set(long.split(" ")); | |
| 29 | + | let shared = 0; | |
| 30 | + | for (const word of shortWords) if (longWords.has(word)) shared++; | |
| 31 | + | if (shortWords.size >= CONTAINED_MIN_WORDS && shared === shortWords.size) return true; | |
| 32 | + | const all = new Set([...shortWords, ...longWords]).size; | |
| 33 | + | return all > 0 && shared / all >= SAME_WORDS; | |
| 34 | + | } | |
| 35 | + | ||
| 1 | 36 | /** | |
| 2 | − | * The same fact remembered twice (two runs that learned it, or one saved | |
| 3 | − | * again) is shown once: the first of them in the order given, so a pinned | |
| 4 | − | * one or the newest wins. Facts match when their words do, whatever the | |
| 5 | − | * case, spacing or a closing full stop. | |
| 37 | + | * The same fact remembered twice (two runs that learned it, one saved | |
| 38 | + | * again, or a sentence that grew since) is shown once: the first of them in | |
| 39 | + | * the order given, so a pinned one or the newest wins. Facts match when | |
| 40 | + | * their words do, whatever the case, spacing or punctuation, or when they | |
| 41 | + | * are near enough to be one (see `sameFact`). | |
| 6 | 42 | */ | |
| 7 | 43 | export function distinctFacts<T extends { text: string }>(facts: T[]): T[] { | |
| 8 | − | const seen = new Set<string>(); | |
| 9 | − | return facts.filter((fact) => { | |
| 10 | − | const key = fact.text.trim().toLowerCase().replace(/\s+/g, " ").replace(/[.!]+$/, ""); | |
| 11 | − | if (seen.has(key)) return false; | |
| 12 | − | seen.add(key); | |
| 13 | − | return true; | |
| 14 | − | }); | |
| 44 | + | const shown: T[] = []; | |
| 45 | + | for (const fact of facts) { | |
| 46 | + | if (!shown.some((other) => sameFact(other.text, fact.text))) shown.push(fact); | |
| 47 | + | } | |
| 48 | + | return shown; | |
| 15 | 49 | } |
| 75 | 75 | .join(" ") | |
| 76 | 76 | } | |
| 77 | 77 | ||
| 78 | + | /// How alike two memories' words must be (shared over all, as sets) to be | |
| 79 | + | /// one memory. | |
| 80 | + | pub const SAME_WORDS: f64 = 0.85; | |
| 81 | + | /// The fewest distinct words a memory needs before it can be found inside | |
| 82 | + | /// another: shorter ones are too general to be the same thing. | |
| 83 | + | const CONTAINED_MIN_WORDS: usize = 6; | |
| 84 | + | ||
| 85 | + | /// Whether two memories say the same thing: the same words once folded | |
| 86 | + | /// ([`fingerprint`]); one's words, in order, inside the other's; all of a | |
| 87 | + | /// memory's words (at least six of them) among the other's; or most of | |
| 88 | + | /// their words shared (Jaccard at or above [`SAME_WORDS`]). "g1t is a | |
| 89 | + | /// Cargo workspace (apps/*, crates/*)" and the same sentence listing every | |
| 90 | + | /// crate are one memory. | |
| 91 | + | pub fn same_memory(a: &str, b: &str) -> bool { | |
| 92 | + | let (a, b) = (fingerprint(a), fingerprint(b)); | |
| 93 | + | if a.is_empty() || b.is_empty() { | |
| 94 | + | return false; | |
| 95 | + | } | |
| 96 | + | if a == b { | |
| 97 | + | return true; | |
| 98 | + | } | |
| 99 | + | let (short, long) = if a.len() <= b.len() { (&a, &b) } else { (&b, &a) }; | |
| 100 | + | let short_words: std::collections::HashSet<&str> = short.split(' ').collect(); | |
| 101 | + | let long_words: std::collections::HashSet<&str> = long.split(' ').collect(); | |
| 102 | + | // Contiguous: the shorter one's words, in order, inside the longer one's. | |
| 103 | + | if short.split(' ').count() >= 4 && format!(" {long} ").contains(&format!(" {short} ")) { | |
| 104 | + | return true; | |
| 105 | + | } | |
| 106 | + | if short_words.len() >= CONTAINED_MIN_WORDS && short_words.is_subset(&long_words) { | |
| 107 | + | return true; | |
| 108 | + | } | |
| 109 | + | let shared = short_words.intersection(&long_words).count(); | |
| 110 | + | let all = short_words.union(&long_words).count(); | |
| 111 | + | all > 0 && shared as f64 / all as f64 >= SAME_WORDS | |
| 112 | + | } | |
| 113 | + | ||
| 78 | 114 | /// One thing learned, as a source reports it. | |
| 79 | 115 | #[derive(Clone, Debug, Serialize, Deserialize)] | |
| 80 | 116 | #[serde(rename_all = "camelCase")] | |
| ⋯ | |||
| 229 | 265 | pub limit: Option<u32>, | |
| 230 | 266 | } | |
| 231 | 267 | ||
| 268 | + | /// `prune_doc_candidates`: after the context service has read every file | |
| 269 | + | /// of a project, removes the project's candidates that came only from its | |
| 270 | + | /// docs and manifests and are still waiting, unless one of `texts` (what | |
| 271 | + | /// they suggest now) is the same memory ([`same_memory`]). Kept and | |
| 272 | + | /// dismissed memory is never touched. Services only. Returns `Pruned`. | |
| 273 | + | #[derive(Debug, Serialize, Deserialize)] | |
| 274 | + | #[serde(rename_all = "camelCase")] | |
| 275 | + | pub struct PruneDocCandidatesArgs { | |
| 276 | + | pub workspace: String, | |
| 277 | + | pub repo_id: String, | |
| 278 | + | #[serde(default)] | |
| 279 | + | pub texts: Vec<String>, | |
| 280 | + | } | |
| 281 | + | ||
| 282 | + | #[derive(Clone, Debug, Default, PartialEq, Eq, Serialize, Deserialize)] | |
| 283 | + | pub struct Pruned { | |
| 284 | + | /// Candidates removed. | |
| 285 | + | pub removed: u32, | |
| 286 | + | } | |
| 287 | + | ||
| 232 | 288 | /// Memories, newest first, for listing a review queue. | |
| 233 | 289 | pub type Candidates = Vec<Memory>; | |
| 234 | 290 | ||
| ⋯ | |||
| 272 | 328 | assert_eq!(fingerprint("use pnpm never NPM"), fingerprint("Use pnpm; never npm.")); | |
| 273 | 329 | assert_ne!(fingerprint("use pnpm"), fingerprint("use npm")); | |
| 274 | 330 | } | |
| 331 | + | ||
| 332 | + | #[test] | |
| 333 | + | fn near_duplicates_are_one_memory() { | |
| 334 | + | assert!(same_memory("Use pnpm, never npm.", "use pnpm never NPM")); | |
| 335 | + | // A crate list that grew, and the same fact without the list. | |
| 336 | + | let before = "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing); `cargo test` runs its tests."; | |
| 337 | + | let after = "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing, services/work); `cargo test` runs its tests."; | |
| 338 | + | assert!(same_memory(before, after)); | |
| 339 | + | assert!(same_memory(before, "g1t is a Cargo workspace (apps/*, crates/*, services/*); `cargo test` runs its tests.")); | |
| 340 | + | // One cut short, inside the whole of it. | |
| 341 | + | assert!(same_memory( | |
| 342 | + | "Roles. Viewer, commenter, planner and approver map onto", | |
| 343 | + | "Roles. Viewer, commenter, planner and approver map onto the five repository roles." | |
| 344 | + | )); | |
| 345 | + | // Different facts that share words are not. | |
| 346 | + | assert!(!same_memory("Use pnpm to install.", "Use npm to install.")); | |
| 347 | + | assert!(!same_memory("Run cargo test in the crate you changed.", "Run npm test in the app you changed.")); | |
| 348 | + | assert!(!same_memory("use pnpm", "use pnpm in the web app and npm in the docs, which predates it")); | |
| 349 | + | assert!(!same_memory("", "anything")); | |
| 350 | + | } | |
| 275 | 351 | } | |
| 237 | 237 | ): Promise<Result<Memory>>; | |
| 238 | 238 | /** For services: candidates from docs and backfills. */ | |
| 239 | 239 | captureMemories(workspace: string, items: CaptureItem[], by?: string): Promise<Captured>; | |
| 240 | + | /** | |
| 241 | + | * For services: removes a project's candidates that came only from its | |
| 242 | + | * docs and are still waiting, unless one of `texts` (what its docs suggest | |
| 243 | + | * now) says the same thing. Kept and dismissed memory is never touched. | |
| 244 | + | */ | |
| 245 | + | pruneDocCandidates(workspace: string, repoId: string, texts: string[]): Promise<{ removed: number }>; | |
| 240 | 246 | /** For services: memories by id, in any status. */ | |
| 241 | 247 | memoriesById(workspace: string, ids: string[]): Promise<Memory[]>; | |
| 242 | 248 | /** For services: kept memories with every word of `query`; the caller has checked the viewer may read them. */ | |
| ⋯ | |||
| 252 | 258 | listCandidates: (viewer, workspace, repo) => call("list_candidates", { viewer, workspace, repo: repo ?? null }), | |
| 253 | 259 | reviewMemory: (actor, workspace, id, decision, change = {}) => call("review_memory", { actor, workspace, id, decision, ...change }), | |
| 254 | 260 | captureMemories: (workspace, items, by) => call("capture_memories", { workspace, items, by: by ?? null }), | |
| 261 | + | pruneDocCandidates: (workspace, repoId, texts) => call("prune_doc_candidates", { workspace, repoId, texts }), | |
| 255 | 262 | memoriesById: (workspace, ids) => call("memories_by_id", { workspace, ids }), | |
| 256 | 263 | searchMemories: (workspace, query, options = {}) => | |
| 257 | 264 | call("search_memories", { workspace, query, repoIds: options.repoIds ?? null, limit: options.limit ?? 20 }), | |
| 9 | 9 | import type { EntityKind, RelationKind } from "@g1t/contracts"; | |
| 10 | 10 | ||
| 11 | 11 | import type { FileFacts, Hint } from "./extract"; | |
| 12 | + | import { harvestable } from "./harvest.ts"; | |
| 12 | 13 | ||
| 13 | 14 | export type ProjectInput = { | |
| 14 | 15 | id: string; | |
| ⋯ | |||
| 198 | 199 | if (integration.repo?.toLowerCase() === repoPath) relate(me, "uses", { kind: "integration", key: integration.id }); | |
| 199 | 200 | } | |
| 200 | 201 | ||
| 201 | − | const hints = files.flatMap((file) => file.facts.hints.map((hint) => ({ ...hint, path: file.path }))); | |
| 202 | + | // A doc is a source of memory only when it says how to work here (see | |
| 203 | + | // ./harvest), or project.yml says it is. | |
| 204 | + | const memory = files.find((file) => file.facts.memory)?.facts.memory ?? null; | |
| 205 | + | const hints = files | |
| 206 | + | .filter((file) => !file.facts.doc || harvestable(file.path, file.facts.doc.role, memory)) | |
| 207 | + | .flatMap((file) => file.facts.hints.map((hint) => ({ ...hint, path: file.path }))); | |
| 202 | 208 | return { entities, relations, hints, tests }; | |
| 203 | 209 | } | |
| 204 | 210 | ||
| 53 | 53 | assert.deepEqual(crate.packages[0].dependencies, ["serde", "worker.workspace", "pretty"]); | |
| 54 | 54 | const workspace = extract("Cargo.toml", `[workspace]\nmembers = [\n "crates/a",\n "services/b",\n]\n`, ctx); | |
| 55 | 55 | assert.deepEqual(workspace.packages[0].members, ["crates/a", "services/b"]); | |
| 56 | − | assert.match(workspace.hints[0].text, /Cargo workspace \(crates\/a, services\/b\)/); | |
| 56 | + | // Where crates live, not each one, so adding a crate says nothing new. | |
| 57 | + | assert.match(workspace.hints[0].text, /Cargo workspace \(crates\/\*, services\/\*\)/); | |
| 58 | + | const grown = extract("Cargo.toml", `[workspace]\nmembers = ["crates/a", "crates/c", "services/b"]\n`, ctx); | |
| 59 | + | assert.equal(grown.hints[0].text, workspace.hints[0].text); | |
| 57 | 60 | }); | |
| 58 | 61 | ||
| 59 | 62 | test("go.mod and pyproject.toml", () => { | |
| ⋯ | |||
| 88 | 91 | assert.equal(agents.doc?.role, "agents"); | |
| 89 | 92 | const kinds = agents.hints.map((hint) => [hint.kind, hint.text, hint.confidence]); | |
| 90 | 93 | assert.deepEqual(kinds, [ | |
| 94 | + | ["fact", "In web, for working here: `npm run db:reset`.", 0.9], | |
| 91 | 95 | ["convention", "Components live in src/components/ui; import from there.", 0.9], | |
| 92 | 96 | ["gotcha", "Never edit generated.rs by hand.", 0.9], | |
| 93 | − | ["fact", "In web, for working here: `npm run db:reset`.", 0.9], | |
| 94 | 97 | ]); | |
| 95 | 98 | const readme = extract("README.md", "# Web\n\nThe storefront.\n\n## Testing\n\n```\n$ TZ=UTC npm test\nnpm test\n```\n\n## Features\n\n- Lots of features that are really great\n", ctx); | |
| 96 | 99 | assert.equal(readme.doc?.title, "Web"); | |
| 9 | 9 | ||
| 10 | 10 | import { parse as parseYaml } from "yaml"; | |
| 11 | 11 | ||
| 12 | + | import { aspirational, bulletsOf, classify, bulletSection, MAX_DOC_HINT_CHARS, planLike, planSection, plain, SETUP_HEADING, wholeSentences, type MemoryDocsConfig } from "./harvest.ts"; | |
| 13 | + | ||
| 12 | 14 | export type Ecosystem = "npm" | "cargo" | "go" | "pypi"; | |
| 13 | 15 | ||
| 14 | 16 | export type PackageFact = { | |
| ⋯ | |||
| 61 | 63 | /** Usernames this file names as owners. */ | |
| 62 | 64 | owners: string[]; | |
| 63 | 65 | hints: Hint[]; | |
| 66 | + | /** For `.g1t/project.yml`: which docs memory is suggested from. */ | |
| 67 | + | memory?: MemoryDocsConfig | null; | |
| 68 | + | /** The `EXTRACT_VERSION` that read it; facts from an older one are read again. */ | |
| 69 | + | version?: number; | |
| 64 | 70 | }; | |
| 65 | 71 | ||
| 66 | 72 | /** What extraction knows besides the file itself. */ | |
| ⋯ | |||
| 71 | 77 | siblings: string[]; | |
| 72 | 78 | }; | |
| 73 | 79 | ||
| 80 | + | /** Bumped when what is read from a file changes, so stored facts are read again. */ | |
| 81 | + | export const EXTRACT_VERSION = 2; | |
| 74 | 82 | export const MAX_DEPENDENCIES = 60; | |
| 75 | 83 | export const CHUNK_CHARS = 1500; | |
| 76 | 84 | export const MAX_CHUNKS = 12; | |
| 77 | 85 | const MAX_ROUTES = 40; | |
| 78 | 86 | const MAX_HINTS_PER_FILE = 12; | |
| 79 | − | const MAX_HINT_CHARS = 300; | |
| 87 | + | const MAX_HINT_CHARS = MAX_DOC_HINT_CHARS; | |
| 80 | 88 | ||
| 81 | 89 | const EMPTY: FileFacts = { languages: [], packages: [], apis: [], doc: null, tests: false, owners: [], hints: [] }; | |
| 82 | 90 | ||
| ⋯ | |||
| 279 | 287 | const name = str(toml.package?.name); | |
| 280 | 288 | const members = Array.isArray(toml.workspace?.members) ? (toml.workspace.members as string[]) : []; | |
| 281 | 289 | const deps = [...keys(toml.dependencies), ...keys(toml["dev-dependencies"]), ...keys(toml["workspace.dependencies"])]; | |
| 290 | + | // Where its crates live (apps/*, crates/*), not each one: the list | |
| 291 | + | // changes with every crate added, and the fact does not. | |
| 292 | + | const places = [...new Set(members.map((member) => (member.includes("/") ? `${member.slice(0, member.indexOf("/"))}/*` : member)))]; | |
| 282 | 293 | const hints: Hint[] = [ | |
| 283 | 294 | { | |
| 284 | 295 | kind: "fact", | |
| 285 | 296 | text: clip( | |
| 286 | 297 | members.length | |
| 287 | − | ? `${ctx.project} is a Cargo workspace (${members.join(", ")}); \`cargo test\` runs its tests.` | |
| 298 | + | ? `${ctx.project} is a Cargo workspace (${places.join(", ")}); \`cargo test\` runs its tests.` | |
| 288 | 299 | : `${ctx.project} is written in Rust; \`cargo test\` runs its tests.`, | |
| 289 | 300 | ), | |
| 290 | 301 | confidence: 0.9, | |
| ⋯ | |||
| 431 | 442 | } | |
| 432 | 443 | ||
| 433 | 444 | const COMMAND = /^\s*(?:\$\s*)?((?:npm|pnpm|yarn|bun|npx|cargo|go|make|pytest|python3?|poetry|uv|docker|wrangler|just|mix|bundle|rake|gradle|\.\/gradlew|mvn|dotnet)\b[^\n]{0,160})$/; | |
| 434 | − | const SETUP_HEADING = /\b(develop|development|getting started|setup|set up|install|build|test|testing|contribut|running|run locally|local)\b/i; | |
| 435 | − | const CONVENTION_HEADING = /\b(convention|guideline|rules|style|standards|gotcha|pitfall|caveat|known issue|do not|don't)\b/i; | |
| 436 | − | const GOTCHA = /\b(never|don't|do not|must not|careful|beware|gotcha|warning|avoid)\b/i; | |
| 437 | 445 | ||
| 438 | − | /** What a doc says worth remembering: commands in its setup sections, and its conventions. */ | |
| 446 | + | /** | |
| 447 | + | * What a doc says worth remembering: the commands in its setup sections, | |
| 448 | + | * and the bullets of its conventions and setup sections, each whole (see | |
| 449 | + | * `./harvest`). A doc that reads as a plan, a report or feedback says | |
| 450 | + | * nothing; nor does a plan section in any doc, or a line that says what | |
| 451 | + | * someone wants rather than how things are. | |
| 452 | + | */ | |
| 439 | 453 | function docHints(path: string, text: string, role: DocFact["role"], ctx: ExtractContext): Hint[] { | |
| 454 | + | const agents = role === "agents"; | |
| 455 | + | if (!agents && planLike(text)) return []; | |
| 440 | 456 | const hints: Hint[] = []; | |
| 441 | − | const agents = role === "agents"; | |
| 457 | + | const where = (heading: string) => `${path}${heading ? ` (${heading})` : ""}`; | |
| 458 | + | ||
| 459 | + | // Commands in code blocks under a setup heading (any, in AGENTS.md). | |
| 442 | 460 | let heading = ""; | |
| 461 | + | let skipping = 0; | |
| 443 | 462 | let fenced = false; | |
| 444 | 463 | for (const raw of text.split(/\r?\n/)) { | |
| 445 | 464 | const line = raw.trimEnd(); | |
| ⋯ | |||
| 447 | 466 | fenced = !fenced; | |
| 448 | 467 | continue; | |
| 449 | 468 | } | |
| 450 | − | const h = /^#{1,6}\s+(.*)$/.exec(line); | |
| 469 | + | const h = /^(#{1,6})\s+(.*)$/.exec(line); | |
| 451 | 470 | if (h && !fenced) { | |
| 452 | − | heading = h[1].trim(); | |
| 471 | + | const level = h[1].length; | |
| 472 | + | if (skipping && level <= skipping) skipping = 0; | |
| 473 | + | heading = h[2].trim(); | |
| 474 | + | if (!skipping && planSection(heading)) skipping = level; | |
| 453 | 475 | continue; | |
| 454 | 476 | } | |
| 455 | − | if (fenced) { | |
| 456 | − | const command = COMMAND.exec(line)?.[1]; | |
| 457 | − | if (command && (agents || SETUP_HEADING.test(heading))) { | |
| 458 | − | const what = heading ? heading.replace(/[:.]$/, "").toLowerCase() : "work on it"; | |
| 459 | − | hints.push({ | |
| 460 | − | kind: "fact", | |
| 461 | − | text: clip(`In ${ctx.project}, for ${what}: \`${command.trim()}\`.`), | |
| 462 | − | confidence: agents ? 0.9 : 0.7, | |
| 463 | − | evidence: `${path}${heading ? ` (${heading})` : ""}`, | |
| 464 | − | }); | |
| 465 | − | } | |
| 466 | − | continue; | |
| 477 | + | if (!fenced || skipping) continue; | |
| 478 | + | const command = COMMAND.exec(line)?.[1]; | |
| 479 | + | if (command && (agents || SETUP_HEADING.test(heading))) { | |
| 480 | + | const what = heading ? plain(heading).replace(/[:.]$/, "").toLowerCase() : "work on it"; | |
| 481 | + | const said = wholeSentences(`In ${ctx.project}, for ${what}: \`${command.trim()}\`.`, MAX_HINT_CHARS); | |
| 482 | + | if (said) hints.push({ kind: "fact", text: said, confidence: agents ? 0.9 : 0.7, evidence: where(heading) }); | |
| 467 | 483 | } | |
| 468 | − | const bullet = /^\s*[-*+]\s+(.*)$/.exec(line)?.[1]?.trim(); | |
| 469 | − | if (!bullet || bullet.length < 20 || bullet.length > MAX_HINT_CHARS) continue; | |
| 484 | + | } | |
| 485 | + | ||
| 486 | + | // Bullets, whole, from the sections that say how things are done. | |
| 487 | + | for (const bullet of bulletsOf(text)) { | |
| 488 | + | if (!bulletSection(bullet.heading, role)) continue; | |
| 470 | 489 | // A bullet that is only a link is a table of contents. | |
| 471 | − | if (/^\[[^\]]*\]\([^)]*\)\.?$/.test(bullet)) continue; | |
| 472 | − | if (agents || CONVENTION_HEADING.test(heading)) { | |
| 473 | − | hints.push({ | |
| 474 | − | kind: GOTCHA.test(bullet) ? "gotcha" : "convention", | |
| 475 | − | text: clip(bullet.replace(/\*\*/g, "")), | |
| 476 | − | confidence: agents ? 0.9 : 0.7, | |
| 477 | − | evidence: `${path}${heading ? ` (${heading})` : ""}`, | |
| 478 | − | }); | |
| 479 | − | } | |
| 490 | + | if (/^\[[^\]]*\]\([^)]*\)\.?$/.test(bullet.text)) continue; | |
| 491 | + | if (aspirational(bullet.text, role)) continue; | |
| 492 | + | const words = plain(bullet.text); | |
| 493 | + | // Too short to mean anything, or a lead-in to a list of its own. | |
| 494 | + | if (words.length < 20 || /:$/.test(words)) continue; | |
| 495 | + | const said = wholeSentences(words, MAX_HINT_CHARS); | |
| 496 | + | if (!said) continue; | |
| 497 | + | hints.push({ ...classify(said, bullet.heading, role), text: said, evidence: where(bullet.heading) }); | |
| 480 | 498 | } | |
| 481 | − | // The same command under two headings is one hint. | |
| 499 | + | // The same line under two headings is one hint. | |
| 482 | 500 | const seen = new Set<string>(); | |
| 483 | − | return hints.filter((hint) => !seen.has(hint.text) && seen.add(hint.text)).slice(0, MAX_HINTS_PER_FILE); | |
| 501 | + | return hints | |
| 502 | + | .filter((hint) => { | |
| 503 | + | const key = hint.text.toLowerCase().replace(/[^\p{L}\p{N}]+/gu, " ").trim(); | |
| 504 | + | return !seen.has(key) && seen.add(key); | |
| 505 | + | }) | |
| 506 | + | .slice(0, MAX_HINTS_PER_FILE); | |
| 484 | 507 | } | |
| 485 | 508 | ||
| 486 | 509 | function doc(path: string, text: string, ctx: ExtractContext): FileFacts { | |
| ⋯ | |||
| 504 | 527 | }; | |
| 505 | 528 | } | |
| 506 | 529 | ||
| 530 | + | /** Strings, from a YAML value that is a string or a list of them. */ | |
| 531 | + | function strings(value: unknown): string[] { | |
| 532 | + | const list = Array.isArray(value) ? value : typeof value === "string" ? [value] : []; | |
| 533 | + | return list.filter((item): item is string => typeof item === "string" && item.trim() !== "").map((item) => item.trim()).slice(0, 50); | |
| 534 | + | } | |
| 535 | + | ||
| 507 | 536 | function owners(path: string, text: string): FileFacts { | |
| 508 | 537 | const names = new Set<string>(); | |
| 538 | + | let memory: MemoryDocsConfig | null = null; | |
| 509 | 539 | if (path.endsWith("project.yml")) { | |
| 510 | − | const parsed = parseYaml(text) as { owners?: unknown } | null; | |
| 540 | + | const parsed = parseYaml(text) as { owners?: unknown; memory?: { docs?: unknown; skip?: unknown } } | null; | |
| 541 | + | if (parsed?.memory && typeof parsed.memory === "object") memory = { docs: strings(parsed.memory.docs), skip: strings(parsed.memory.skip) }; | |
| 511 | 542 | const list = Array.isArray(parsed?.owners) ? parsed.owners : typeof parsed?.owners === "string" ? [parsed.owners] : []; | |
| 512 | 543 | for (const owner of list) if (typeof owner === "string") names.add(owner.replace(/^@/, "").trim()); | |
| 513 | 544 | } else { | |
| ⋯ | |||
| 516 | 547 | for (const match of line.matchAll(/@([A-Za-z0-9][A-Za-z0-9_.-]*)(?=\s|$)/g)) names.add(match[1]); | |
| 517 | 548 | } | |
| 518 | 549 | } | |
| 519 | − | return { ...EMPTY, owners: [...names].filter(Boolean).slice(0, 20) }; | |
| 550 | + | return { ...EMPTY, owners: [...names].filter(Boolean).slice(0, 20), ...(memory ? { memory } : {}) }; | |
| 520 | 551 | } | |
| 521 | 552 | ||
| 522 | 553 | const TESTS = /\b(test|tests|pytest|vitest|jest|mocha|cargo test|go test|npm test|playwright|cypress)\b/i; | |
| 523 | 554 | ||
| 524 | 555 | /** The facts in one file. A file that cannot be parsed says nothing. */ | |
| 525 | 556 | export function extract(path: string, text: string, ctx: ExtractContext): FileFacts { | |
| 557 | + | return { ...read(path, text, ctx), version: EXTRACT_VERSION }; | |
| 558 | + | } | |
| 559 | + | ||
| 560 | + | function read(path: string, text: string, ctx: ExtractContext): FileFacts { | |
| 526 | 561 | const name = base(path); | |
| 527 | 562 | try { | |
| 528 | 563 | if (path.startsWith(".g1t/workflows/") || path.startsWith(".github/workflows/")) { | |
| 1 | + | import { test } from "node:test"; | |
| 2 | + | import assert from "node:assert/strict"; | |
| 3 | + | ||
| 4 | + | import { assemble } from "./assemble.ts"; | |
| 5 | + | import { extract } from "./extract.ts"; | |
| 6 | + | import { aspirational, bulletsOf, classify, harvestable, planLike, sentencesOf, wholeSentences } from "./harvest.ts"; | |
| 7 | + | ||
| 8 | + | const ctx = { project: "g1t", siblings: [] }; | |
| 9 | + | ||
| 10 | + | test("memory comes from docs on how to work here, never plans, feedback or history", () => { | |
| 11 | + | for (const path of ["README.md", "AGENTS.md", "CLAUDE.md", "CONTRIBUTING.md"]) { | |
| 12 | + | const role = path.startsWith("README") ? "readme" : path.startsWith("CONTRIBUTING") ? "contributing" : "agents"; | |
| 13 | + | assert.ok(harvestable(path, role), path); | |
| 14 | + | } | |
| 15 | + | for (const path of ["docs/DEPLOYING.md", "docs/testing.md", "docs/architecture.md", "docs/SELF_HOSTING.md", "runbooks/restore.md", "docs/getting-started.md"]) { | |
| 16 | + | assert.ok(harvestable(path, "doc"), path); | |
| 17 | + | } | |
| 18 | + | for (const path of ["docs/PLAN.md", "docs/CLOUDFLARE_FEEDBACK.md", "docs/CHANGELOG.md", "docs/roadmap.md", "docs/DEMO.md", "docs/INCIDENTS.md", "docs/PERFORMANCE.md", "docs/release-notes.md"]) { | |
| 19 | + | assert.ok(!harvestable(path, "doc"), path); | |
| 20 | + | } | |
| 21 | + | // project.yml adds a doc, or leaves one out, and leaving out wins. | |
| 22 | + | assert.ok(harvestable("docs/PERFORMANCE.md", "doc", { docs: ["docs/PERFORMANCE.md"], skip: [] })); | |
| 23 | + | assert.ok(harvestable("docs/notes/x.md", "doc", { docs: ["docs/notes/*"], skip: [] })); | |
| 24 | + | assert.ok(!harvestable("README.md", "readme", { docs: [], skip: ["README.md"] })); | |
| 25 | + | assert.ok(!harvestable("docs/testing.md", "doc", { docs: ["docs/testing.md"], skip: ["docs/*"] })); | |
| 26 | + | }); | |
| 27 | + | ||
| 28 | + | const FEEDBACK = `# Building a platform: a field report | |
| 29 | + | ||
| 30 | + | ### A1. What is a billable operation? (blocking) | |
| 31 | + | ||
| 32 | + | - **What we tried.** Price from cost: pass storage through to each workspace with a modest | |
| 33 | + | uniform overhead, so a bill tracks what it really costs us. | |
| 34 | + | - **What we hit.** Pricing defines operations only loosely. | |
| 35 | + | - **What it means.** Our capacity model gives a 15x spread. | |
| 36 | + | - **Ask.** A table: each binding method and each endpoint, billable or not. | |
| 37 | + | ||
| 38 | + | ### Conventions we would like | |
| 39 | + | ||
| 40 | + | - **Ask.** Metrics with the same event names as the invoice, per repository. | |
| 41 | + | `; | |
| 42 | + | ||
| 43 | + | const PLAN = `# Plan | |
| 44 | + | ||
| 45 | + | ## For people who do not write code | |
| 46 | + | ||
| 47 | + | - **Roles.** A person can plan and approve work without ever cloning a | |
| 48 | + | repo. The roles that were sketched here map onto the five repository | |
| 49 | + | roles that are built. | |
| 50 | + | ||
| 51 | + | ### Rules every automation obeys | |
| 52 | + | ||
| 53 | + | - **Deduplication.** The same Sentry issue firing 500 times maps to one | |
| 54 | + | issue. | |
| 55 | + | - **Loop protection.** Work started by an automation cannot retrigger the | |
| 56 | + | same automation without a person in between. | |
| 57 | + | ||
| 58 | + | ## Roadmap | |
| 59 | + | ||
| 60 | + | ## Later | |
| 61 | + | ||
| 62 | + | - g1t will learn from docs. | |
| 63 | + | `; | |
| 64 | + | ||
| 65 | + | test("a doc that reads as feedback or a plan says nothing, whatever its name", () => { | |
| 66 | + | assert.ok(planLike(FEEDBACK)); | |
| 67 | + | assert.deepEqual(extract("CONTRIBUTING.md", FEEDBACK, ctx).hints, []); | |
| 68 | + | assert.ok(planLike(PLAN)); | |
| 69 | + | assert.deepEqual(extract("docs/testing.md", PLAN, ctx).hints, []); | |
| 70 | + | const files = [ | |
| 71 | + | { path: "docs/PLAN.md", facts: extract("docs/PLAN.md", "# Working on g1t\n\n## Conventions\n\n- Run the whole suite before you push, every time.\n", ctx) }, | |
| 72 | + | { path: "docs/CLOUDFLARE_FEEDBACK.md", facts: extract("docs/CLOUDFLARE_FEEDBACK.md", FEEDBACK, ctx) }, | |
| 73 | + | ]; | |
| 74 | + | const project = { id: "p", workspace: "w", slug: "g1t", name: "g1t", description: null, private: false, repoId: "r", repo: { namespace: "w", name: "g1t" }, rootDir: "", defaultBranch: "main" }; | |
| 75 | + | const around = { owners: [], dependsOn: [], deploy: null, integrations: [] }; | |
| 76 | + | assert.deepEqual(assemble(project, files, around).hints, [], "docs/PLAN.md is not a source of memory"); | |
| 77 | + | // Unless project.yml says it is. | |
| 78 | + | const withConfig = [...files, { path: ".g1t/project.yml", facts: extract(".g1t/project.yml", "memory:\n docs:\n - docs/PLAN.md\n", ctx) }]; | |
| 79 | + | assert.deepEqual(assemble(project, withConfig, around).hints.map((hint) => hint.path), ["docs/PLAN.md"]); | |
| 80 | + | }); | |
| 81 | + | ||
| 82 | + | test("a bullet is taken whole, with the lines it wraps onto", () => { | |
| 83 | + | const bullets = bulletsOf( | |
| 84 | + | "## Conventions\n\n- **Loop protection.** Work started by an automation cannot retrigger the\n same automation without a person in between.\n- Short one.\n\nA paragraph.\n\n```\n- not a bullet\n```\n", | |
| 85 | + | ); | |
| 86 | + | assert.deepEqual(bullets, [ | |
| 87 | + | { text: "**Loop protection.** Work started by an automation cannot retrigger the same automation without a person in between.", heading: "Conventions" }, | |
| 88 | + | { text: "Short one.", heading: "Conventions" }, | |
| 89 | + | ]); | |
| 90 | + | // Bullets in a plan section, and its subsections, are left out. | |
| 91 | + | assert.deepEqual(bulletsOf("## Roadmap\n\n### Next\n\n- One.\n\n## Testing\n\n- Two.\n").map((b) => b.text), ["Two."]); | |
| 92 | + | }); | |
| 93 | + | ||
| 94 | + | test("never cut inside a sentence: whole sentences up to the limit, or nothing", () => { | |
| 95 | + | assert.deepEqual(sentencesOf("Run `npm test` first. It reads README.md, e.g. the setup. Done!"), [ | |
| 96 | + | "Run `npm test` first.", | |
| 97 | + | "It reads README.md, e.g. the setup.", | |
| 98 | + | "Done!", | |
| 99 | + | ]); | |
| 100 | + | assert.equal(wholeSentences("Short and whole", 40), "Short and whole"); | |
| 101 | + | assert.equal(wholeSentences("First sentence is here. Second sentence pushes it past the limit.", 30), "First sentence is here."); | |
| 102 | + | assert.equal(wholeSentences("One very long sentence that never ends before the limit is reached at all", 30), null); | |
| 103 | + | }); | |
| 104 | + | ||
| 105 | + | test("plans and wishes are not memory; kinds are earned, not assumed", () => { | |
| 106 | + | assert.ok(aspirational("g1t will learn from docs.", "doc")); | |
| 107 | + | assert.ok(aspirational("Tests should run on every push.", "contributing")); | |
| 108 | + | assert.ok(aspirational("**Ask.** A usage endpoint.", "doc")); | |
| 109 | + | assert.ok(aspirational("**What we tried.** Pass storage through.", "doc")); | |
| 110 | + | assert.ok(!aspirational("Run `cargo test` in the crate you changed.", "contributing")); | |
| 111 | + | // In AGENTS.md "should" states a rule. | |
| 112 | + | assert.ok(!aspirational("You should run the formatter before committing.", "agents")); | |
| 113 | + | assert.ok(aspirational("TODO: document the release flow.", "agents")); | |
| 114 | + | ||
| 115 | + | assert.deepEqual(classify("Never edit generated files by hand.", "Conventions", "contributing"), { kind: "gotcha", confidence: 0.6 }); | |
| 116 | + | assert.deepEqual(classify("Use the shared client for every call.", "Conventions", "contributing"), { kind: "convention", confidence: 0.6 }); | |
| 117 | + | assert.deepEqual(classify("The API lives in apps/api and serves api.g1t.sh.", "Conventions", "readme"), { kind: "fact", confidence: 0.5 }); | |
| 118 | + | }); | |
| 119 | + | ||
| 120 | + | test("CONTRIBUTING's before-you-push bullets are suggested whole; nothing is cut", () => { | |
| 121 | + | const facts = extract( | |
| 122 | + | "CONTRIBUTING.md", | |
| 123 | + | "# Contributing\n\n## Before you push\n\n- `cargo test` in the crate or service you changed.\n- `npx tsc -b --force` in `apps/web` (the incremental build misses changes\n in `packages/contracts`).\n- Look at what you changed in a browser. Screenshots catch what type\n checks do not.\n\n## Deploying\n\nPushes to `main` deploy themselves.\n", | |
| 124 | + | ctx, | |
| 125 | + | ); | |
| 126 | + | assert.deepEqual( | |
| 127 | + | facts.hints.map((hint) => [hint.kind, hint.text, hint.confidence]), | |
| 128 | + | [ | |
| 129 | + | ["fact", "`cargo test` in the crate or service you changed.", 0.5], | |
| 130 | + | ["fact", "`npx tsc -b --force` in `apps/web` (the incremental build misses changes in `packages/contracts`).", 0.5], | |
| 131 | + | ["convention", "Look at what you changed in a browser. Screenshots catch what type checks do not.", 0.6], | |
| 132 | + | ], | |
| 133 | + | ); | |
| 134 | + | for (const hint of facts.hints) assert.match(hint.text, /[.!?)`]$/); | |
| 135 | + | }); |
| 1 | + | /** | |
| 2 | + | * Which docs, and which of their lines, are worth suggesting as memory. | |
| 3 | + | * | |
| 4 | + | * Memory is for how to work in a project: how to build, test and run it, | |
| 5 | + | * its conventions and its traps. A project's docs also hold plans, | |
| 6 | + | * roadmaps, feedback to vendors and release notes, which say what someone | |
| 7 | + | * wants, not how things are; none of that is suggested. A line is taken | |
| 8 | + | * whole (a bullet with its continuation lines), cut only between sentences | |
| 9 | + | * to fit, and left out when even its first sentence does not fit. Pure. | |
| 10 | + | */ | |
| 11 | + | ||
| 12 | + | /** README, AGENTS.md (or CLAUDE.md), CONTRIBUTING, or another doc. */ | |
| 13 | + | export type DocRole = "readme" | "agents" | "contributing" | "doc"; | |
| 14 | + | ||
| 15 | + | /** What a project's `.g1t/project.yml` says about memory from its docs. */ | |
| 16 | + | export type MemoryDocsConfig = { | |
| 17 | + | /** Docs to read for memory besides the ones g1t picks, by path or `dir/*`. */ | |
| 18 | + | docs: string[]; | |
| 19 | + | /** Docs never read for memory, by path or `dir/*`. Wins over `docs`. */ | |
| 20 | + | skip: string[]; | |
| 21 | + | }; | |
| 22 | + | ||
| 23 | + | /** The longest memory suggested from a doc, in characters. */ | |
| 24 | + | export const MAX_DOC_HINT_CHARS = 300; | |
| 25 | + | ||
| 26 | + | /** Doc names that say how to work on a project. */ | |
| 27 | + | const WORKING_NAME = | |
| 28 | + | /(build|test|setup|set-up|develop|contribut|hacking|architecture|convention|style|guideline|standard|deploy|install|getting[-_ ]?started|local|onboard|troubleshoot|debug|runbook|self[-_ ]?host|coding|structure|layout|workflow|agents|claude)/i; | |
| 29 | + | /** Doc names that hold plans, feedback, history or news. */ | |
| 30 | + | const PLAN_NAME = | |
| 31 | + | /(plan|roadmap|feedback|changelog|changes|history|news|release[-_ ]?notes|rfc|proposal|idea|todo|backlog|vision|strategy|research|demo|incident|postmortem|post-mortem|retro|meeting|minutes|notes|draft|wish|faq|announce|blog|pitch)/i; | |
| 32 | + | ||
| 33 | + | /** Headings of sections that hold plans, asks or history. */ | |
| 34 | + | const PLAN_HEADING = | |
| 35 | + | /\b(roadmap|plans?|planned|planning|milestones?|phase \d+|future|later|next steps|what'?s next|what we'?d love|wish ?list|asks?|feedback|changelog|release notes|what'?s new|open questions|non-goals|goals|ideas|proposals?|backlog|todo|executive summary|what we (hit|tried|built|want))\b/i; | |
| 36 | + | /** Headings of sections that say how things are done here. */ | |
| 37 | + | const CONVENTION_HEADING = | |
| 38 | + | /\b(conventions?|guidelines?|code style|coding style|style guide|standards|gotchas?|pitfalls?|caveats?|known issues|troubleshooting|before you (push|commit|open)|house rules|how we work|working (here|on|in)|dos and don'?ts|rules for (contributors|agents|code))\b/i; | |
| 39 | + | /** Headings of sections that say how to set up, build or test. */ | |
| 40 | + | export const SETUP_HEADING = | |
| 41 | + | /\b(develop|development|getting started|setup|set up|install|installing|build|building|test|testing|contribut|running|run locally|local|before you push|deploying)\b/i; | |
| 42 | + | ||
| 43 | + | /** A bullet's bold label that marks it as a plan, an ask or a report. */ | |
| 44 | + | const PLAN_LABEL = | |
| 45 | + | /^(ask|asks|what we (tried|hit|built|want|need|'d love)|what it means|proposal|idea|plan|next|later|future|todo|status|why it matters)\b/i; | |
| 46 | + | /** Words that make a line a plan or a wish rather than how things are. */ | |
| 47 | + | const ASPIRATION = | |
| 48 | + | /\b(will|won'?t|would|should|could|might|shall|plans? to|planned|we'?d|we'?ll|eventually|someday|one day|in the future|not yet|coming soon|todo|tbd|wip|roadmap)\b|^ask\b/i; | |
| 49 | + | /** The modal words an AGENTS.md uses to state a rule; aspirational elsewhere. */ | |
| 50 | + | const RULE_MODAL = /\b(will|won'?t|would|should|could|might|shall)\b/gi; | |
| 51 | + | ||
| 52 | + | /** Words that mark a line as a trap. */ | |
| 53 | + | const GOTCHA = | |
| 54 | + | /^(never|don'?t|do not|avoid|beware|careful|warning|watch out)\b|\b(must not|must never|gotcha|beware|careful|fails? (unless|if|when|until|without)|breaks? (if|when|unless)|otherwise)\b/i; | |
| 55 | + | /** Words that mark a line as a rule of how things are done. */ | |
| 56 | + | const CONVENTION = | |
| 57 | + | /^(use|run|prefer|keep|put|name|write|import|add|look|always|never|don'?t|do not|avoid|make sure|remember|check|update|commit|open|follow|place|call|read|test|document)\b|\b(we use|must|prefer|instead of|rather than|always|never|convention|is required|are required)\b/i; | |
| 58 | + | ||
| 59 | + | /** Whether `path` matches one of `patterns`: an exact path, or `dir/*`. */ | |
| 60 | + | function matches(path: string, patterns: string[]): boolean { | |
| 61 | + | const lower = path.toLowerCase(); | |
| 62 | + | return patterns.some((raw) => { | |
| 63 | + | const pattern = raw.trim().replace(/^\.?\//, "").toLowerCase(); | |
| 64 | + | if (!pattern) return false; | |
| 65 | + | if (pattern.endsWith("/*")) return lower.startsWith(pattern.slice(0, -1)) && !lower.slice(pattern.length - 1).includes("/"); | |
| 66 | + | if (pattern.endsWith("/")) return lower.startsWith(pattern); | |
| 67 | + | return lower === pattern; | |
| 68 | + | }); | |
| 69 | + | } | |
| 70 | + | ||
| 71 | + | /** | |
| 72 | + | * Whether memory is suggested from the doc at `path`: a README, AGENTS.md, | |
| 73 | + | * CLAUDE.md or CONTRIBUTING always, a runbook, and another doc when its | |
| 74 | + | * name says how to work on the project (`docs/DEPLOYING.md`, not | |
| 75 | + | * `docs/PLAN.md` or `docs/CHANGELOG.md`). `.g1t/project.yml` can add docs | |
| 76 | + | * or leave them out. | |
| 77 | + | */ | |
| 78 | + | export function harvestable(path: string, role: DocRole, config?: MemoryDocsConfig | null): boolean { | |
| 79 | + | if (config && matches(path, config.skip)) return false; | |
| 80 | + | if (config && matches(path, config.docs)) return true; | |
| 81 | + | if (role !== "doc") return true; | |
| 82 | + | const name = path.slice(path.lastIndexOf("/") + 1).replace(/\.(md|markdown|txt)$/i, ""); | |
| 83 | + | if (PLAN_NAME.test(name)) return false; | |
| 84 | + | if (/(^|\/)runbooks\//i.test(path)) return true; | |
| 85 | + | return WORKING_NAME.test(name); | |
| 86 | + | } | |
| 87 | + | ||
| 88 | + | /** | |
| 89 | + | * Whether a doc reads as a plan, a report or feedback rather than how to | |
| 90 | + | * work: many of its headings are plans or asks, or many of its bullets are | |
| 91 | + | * labelled as asks and findings. | |
| 92 | + | */ | |
| 93 | + | export function planLike(text: string): boolean { | |
| 94 | + | let headings = 0; | |
| 95 | + | let planHeadings = 0; | |
| 96 | + | let labels = 0; | |
| 97 | + | let aspirations = 0; | |
| 98 | + | let lines = 0; | |
| 99 | + | for (const line of text.split(/\r?\n/)) { | |
| 100 | + | const heading = /^#{1,6}\s+(.*)$/.exec(line)?.[1]; | |
| 101 | + | if (heading) { | |
| 102 | + | headings++; | |
| 103 | + | if (PLAN_HEADING.test(heading)) planHeadings++; | |
| 104 | + | continue; | |
| 105 | + | } | |
| 106 | + | const label = /^\s*[-*+]\s+\*\*([^*]+)\*\*/.exec(line)?.[1]; | |
| 107 | + | if (label && PLAN_LABEL.test(label.trim())) labels++; | |
| 108 | + | if (line.trim()) { | |
| 109 | + | lines++; | |
| 110 | + | if (/\b(we will|we'?ll|we plan|we'?d love|will be|is planned|are planned|not yet built|coming soon)\b/i.test(line)) aspirations++; | |
| 111 | + | } | |
| 112 | + | } | |
| 113 | + | return labels >= 3 || (headings >= 3 && planHeadings * 3 >= headings) || (lines >= 20 && aspirations * 8 >= lines); | |
| 114 | + | } | |
| 115 | + | ||
| 116 | + | /** Markdown inline marks taken out: bold, italics, links to their text, images. */ | |
| 117 | + | export function plain(text: string): string { | |
| 118 | + | return text | |
| 119 | + | .replace(/!\[[^\]]*\]\([^)]*\)/g, "") | |
| 120 | + | .replace(/\[([^\]]*)\]\([^)]*\)/g, "$1") | |
| 121 | + | .replace(/\*\*([^*]+)\*\*/g, "$1") | |
| 122 | + | .replace(/__([^_]+)__/g, "$1") | |
| 123 | + | .replace(/(^|\s)\*([^*\s][^*]*)\*(?=\s|[.,;:!?)]|$)/g, "$1$2") | |
| 124 | + | .replace(/\s+/g, " ") | |
| 125 | + | .trim(); | |
| 126 | + | } | |
| 127 | + | ||
| 128 | + | const ABBREVIATION = /(\b(e\.g|i\.e|etc|vs|cf|approx|no|fig)\.|\b[A-Z]\.)$/i; | |
| 129 | + | ||
| 130 | + | /** A line's sentences, never split inside `code` or after an abbreviation. */ | |
| 131 | + | export function sentencesOf(text: string): string[] { | |
| 132 | + | const out: string[] = []; | |
| 133 | + | let current = ""; | |
| 134 | + | let code = false; | |
| 135 | + | for (let i = 0; i < text.length; i++) { | |
| 136 | + | const c = text[i]; | |
| 137 | + | current += c; | |
| 138 | + | if (c === "`") code = !code; | |
| 139 | + | if (code || !/[.!?]/.test(c)) continue; | |
| 140 | + | const next = text[i + 1]; | |
| 141 | + | // A full stop inside a word (README.md, v1.2) or before more punctuation. | |
| 142 | + | if (next !== undefined && !/\s/.test(next)) continue; | |
| 143 | + | if (ABBREVIATION.test(current.trimEnd())) continue; | |
| 144 | + | out.push(current.trim()); | |
| 145 | + | current = ""; | |
| 146 | + | } | |
| 147 | + | if (current.trim()) out.push(current.trim()); | |
| 148 | + | return out; | |
| 149 | + | } | |
| 150 | + | ||
| 151 | + | /** | |
| 152 | + | * `text` whole if it fits in `max` characters; otherwise as many of its | |
| 153 | + | * first sentences as fit; null when not even the first one does. Never cut | |
| 154 | + | * inside a sentence. | |
| 155 | + | */ | |
| 156 | + | export function wholeSentences(text: string, max = MAX_DOC_HINT_CHARS): string | null { | |
| 157 | + | const line = text.replace(/\s+/g, " ").trim(); | |
| 158 | + | if (line.length <= max) return line; | |
| 159 | + | let out = ""; | |
| 160 | + | for (const sentence of sentencesOf(line)) { | |
| 161 | + | const next = out ? `${out} ${sentence}` : sentence; | |
| 162 | + | if (next.length > max) break; | |
| 163 | + | out = next; | |
| 164 | + | } | |
| 165 | + | return out && /[.!?]$/.test(out) ? out : null; | |
| 166 | + | } | |
| 167 | + | ||
| 168 | + | /** Whether a line says what someone wants or plans rather than how things are. */ | |
| 169 | + | export function aspirational(text: string, role: DocRole): boolean { | |
| 170 | + | const label = /^\*\*([^*]+)\*\*/.exec(text.trim())?.[1]?.trim(); | |
| 171 | + | if (label && PLAN_LABEL.test(label)) return true; | |
| 172 | + | const words = plain(text); | |
| 173 | + | if (/^(ask|what we (tried|hit|built))\b/i.test(words)) return true; | |
| 174 | + | // AGENTS.md states its rules with "should" and "will": those are rules. | |
| 175 | + | const checked = role === "agents" ? words.replace(RULE_MODAL, "") : words; | |
| 176 | + | return ASPIRATION.test(checked); | |
| 177 | + | } | |
| 178 | + | ||
| 179 | + | /** The kind of memory a line from a doc is, and how sure g1t is of it. */ | |
| 180 | + | export function classify(text: string, heading: string, role: DocRole): { kind: "fact" | "convention" | "gotcha"; confidence: number } { | |
| 181 | + | const agents = role === "agents"; | |
| 182 | + | if (GOTCHA.test(text)) return { kind: "gotcha", confidence: agents ? 0.9 : 0.6 }; | |
| 183 | + | if (CONVENTION.test(text) || (agents && CONVENTION_HEADING.test(heading))) return { kind: "convention", confidence: agents ? 0.9 : 0.6 }; | |
| 184 | + | return { kind: "fact", confidence: agents ? 0.85 : 0.5 }; | |
| 185 | + | } | |
| 186 | + | ||
| 187 | + | /** Headings of sections that report or explain rather than say how. */ | |
| 188 | + | const REPORT_HEADING = /\b(speed|performance|benchmarks?|timings?|costs?|numbers|results|findings|lessons|why|background|motivation|history|incidents?|features|status)\b/i; | |
| 189 | + | ||
| 190 | + | /** Whether a section's bullets are worth reading for memory. */ | |
| 191 | + | export function bulletSection(heading: string, role: DocRole): boolean { | |
| 192 | + | if (PLAN_HEADING.test(heading) || (role !== "agents" && REPORT_HEADING.test(heading))) return false; | |
| 193 | + | return role === "agents" || CONVENTION_HEADING.test(heading) || SETUP_HEADING.test(heading); | |
| 194 | + | } | |
| 195 | + | ||
| 196 | + | /** Whether a section is a plan, an ask or history, with everything under it. */ | |
| 197 | + | export function planSection(heading: string): boolean { | |
| 198 | + | return PLAN_HEADING.test(heading); | |
| 199 | + | } | |
| 200 | + | ||
| 201 | + | export type Bullet = { text: string; heading: string }; | |
| 202 | + | ||
| 203 | + | /** | |
| 204 | + | * A doc's bullets, each whole with its continuation lines, with the heading | |
| 205 | + | * it is under. Bullets in code blocks and in plan sections (and their | |
| 206 | + | * subsections) are left out. | |
| 207 | + | */ | |
| 208 | + | export function bulletsOf(markdown: string): Bullet[] { | |
| 209 | + | const out: Bullet[] = []; | |
| 210 | + | let heading = ""; | |
| 211 | + | // The level of the plan section being skipped, if any. | |
| 212 | + | let skipping = 0; | |
| 213 | + | let fenced = false; | |
| 214 | + | let current: { text: string; indent: number } | null = null; | |
| 215 | + | const flush = () => { | |
| 216 | + | if (current && !skipping) out.push({ text: current.text.replace(/\s+/g, " ").trim(), heading }); | |
| 217 | + | current = null; | |
| 218 | + | }; | |
| 219 | + | for (const raw of markdown.split(/\r?\n/)) { | |
| 220 | + | if (/^\s*(```|~~~)/.test(raw)) { | |
| 221 | + | flush(); | |
| 222 | + | fenced = !fenced; | |
| 223 | + | continue; | |
| 224 | + | } | |
| 225 | + | if (fenced) continue; | |
| 226 | + | const h = /^(#{1,6})\s+(.*)$/.exec(raw); | |
| 227 | + | if (h) { | |
| 228 | + | flush(); | |
| 229 | + | const level = h[1].length; | |
| 230 | + | if (skipping && level <= skipping) skipping = 0; | |
| 231 | + | heading = h[2].trim(); | |
| 232 | + | if (!skipping && planSection(heading)) skipping = level; | |
| 233 | + | continue; | |
| 234 | + | } | |
| 235 | + | const bullet = /^(\s*)(?:[-*+]|\d+[.)])\s+(.*)$/.exec(raw); | |
| 236 | + | if (bullet) { | |
| 237 | + | flush(); | |
| 238 | + | current = { text: bullet[2], indent: bullet[1].length }; | |
| 239 | + | continue; | |
| 240 | + | } | |
| 241 | + | if (current && raw.trim() && /^\s+/.test(raw) && raw.search(/\S/) > current.indent) { | |
| 242 | + | current.text += ` ${raw.trim()}`; | |
| 243 | + | continue; | |
| 244 | + | } | |
| 245 | + | // A blank line or a paragraph ends the bullet. A bullet's lazy | |
| 246 | + | // continuation (not indented) is still part of it. | |
| 247 | + | if (current && raw.trim() && !/^\s*[|<>]/.test(raw)) { | |
| 248 | + | current.text += ` ${raw.trim()}`; | |
| 249 | + | continue; | |
| 250 | + | } | |
| 251 | + | flush(); | |
| 252 | + | } | |
| 253 | + | flush(); | |
| 254 | + | return out; | |
| 255 | + | } |
| 68 | 68 | } from "@g1t/contracts"; | |
| 69 | 69 | ||
| 70 | 70 | import { assemble, authorsOf, integrationEntities, type EntityDraft, type FileRecord, type ProjectInput, type Surroundings } from "./assemble"; | |
| 71 | − | import { extract, interesting, type FileFacts } from "./extract"; | |
| 71 | + | import { EXTRACT_VERSION, extract, interesting, type FileFacts } from "./extract"; | |
| 72 | 72 | import { composeRunContext, type ContextNote, type ProjectContext } from "./runcontext"; | |
| 73 | 73 | import { evaluate } from "./scorecards"; | |
| 74 | 74 | import { allowedKinds, countVisible, indexFilter, memoryReadable, merge, projectReadable, readable, runMemoryReadable, type IndexMeta, type Reader } from "./visibility"; | |
| ⋯ | |||
| 271 | 271 | } | |
| 272 | 272 | } | |
| 273 | 273 | ||
| 274 | − | type ScanStats = { entities: number; candidates: number; kept: number; indexed: number }; | |
| 274 | + | type ScanStats = { entities: number; candidates: number; kept: number; indexed: number; pruned?: number }; | |
| 275 | 275 | ||
| 276 | 276 | class Context { | |
| 277 | 277 | constructor(private readonly env: Env) {} | |
| ⋯ | |||
| 454 | 454 | // ---- Building the catalog -------------------------------------------------- | |
| 455 | 455 | ||
| 456 | 456 | /** The files of a project worth reading, with their blobs, found in a few tree reads. */ | |
| 457 | − | private async candidates(project: Project, actor: User, ref: string): Promise<{ files: { path: string; hash: string }[]; siblings: string[]; head: string | null } | null> { | |
| 457 | + | private async candidates(project: Project, actor: User, ref: string): Promise<{ files: { path: string; hash: string }[]; siblings: string[]; head: string | null; partial: boolean } | null> { | |
| 458 | 458 | if (project.source.kind !== "hosted") return null; | |
| 459 | 459 | const repos = reposClient(this.env.REPOS); | |
| 460 | 460 | const { repo, rootDir: root } = project.source; | |
| ⋯ | |||
| 472 | 472 | }; | |
| 473 | 473 | blobs(top.value.entries, ""); | |
| 474 | 474 | const dirs = new Set(top.value.entries.filter((entry) => entry.kind === "tree").map((entry) => entry.name)); | |
| 475 | + | // A folder that could not be listed leaves the list partial. | |
| 476 | + | let partial = false; | |
| 475 | 477 | const look = async (dir: string) => { | |
| 476 | 478 | const found = await repos.tree(repo, actor, ref, at(dir)).catch(() => null); | |
| 479 | + | if (!found?.ok) partial = true; | |
| 477 | 480 | return found?.ok ? found.value.entries : []; | |
| 478 | 481 | }; | |
| 479 | 482 | for (const dir of ["docs", "doc", "runbooks"]) if (dirs.has(dir)) blobs(await look(dir), dir); | |
| ⋯ | |||
| 483 | 486 | blobs(inside, dir); | |
| 484 | 487 | if (inside.some((entry) => entry.name === "workflows" && entry.kind === "tree")) blobs(await look(`${dir}/workflows`), `${dir}/workflows`); | |
| 485 | 488 | } | |
| 486 | − | return { files, siblings, head: top.value.head?.hash ?? null }; | |
| 489 | + | return { files, siblings, head: top.value.head?.hash ?? null, partial }; | |
| 487 | 490 | } | |
| 488 | 491 | ||
| 489 | 492 | /** | |
| ⋯ | |||
| 511 | 514 | const writes: D1PreparedStatement[] = []; | |
| 512 | 515 | let reads = 0; | |
| 513 | 516 | let workflowReads = 0; | |
| 517 | + | // Whether every file is known as this version of extract reads it: only | |
| 518 | + | // then are the doc candidates it no longer suggests let go. | |
| 519 | + | let complete = !found.partial; | |
| 514 | 520 | const repos = reposClient(this.env.REPOS); | |
| 515 | 521 | for (const file of found.files) { | |
| 516 | 522 | const before = stored.get(file.path); | |
| 523 | + | const kept = before ? (JSON.parse(before.facts) as FileFacts) : null; | |
| 517 | 524 | const workflow = file.path.includes("/workflows/"); | |
| 518 | − | const fresh = before && before.hash === file.hash && !force; | |
| 525 | + | // Facts an older extract read are read again, as a changed file is. | |
| 526 | + | const fresh = before && before.hash === file.hash && kept?.version === EXTRACT_VERSION && !force; | |
| 519 | 527 | const canRead = reads < MAX_READS && (!workflow || workflowReads < MAX_WORKFLOW_READS); | |
| 520 | 528 | if (fresh || !canRead) { | |
| 521 | − | if (before) files.push({ path: file.path, facts: JSON.parse(before.facts) as FileFacts }); | |
| 529 | + | if (!fresh) complete = false; | |
| 530 | + | if (kept) files.push({ path: file.path, facts: kept }); | |
| 522 | 531 | continue; | |
| 523 | 532 | } | |
| 524 | 533 | reads++; | |
| 525 | 534 | if (workflow) workflowReads++; | |
| 526 | 535 | const blob = await repos.blob(repo, cache.actor, found.head ?? ref, [rootDir, file.path].filter(Boolean).join("/")).catch(() => null); | |
| 527 | − | if (!blob?.ok || blob.value.text == null) continue; | |
| 536 | + | if (!blob?.ok || blob.value.text == null) { | |
| 537 | + | complete = false; | |
| 538 | + | continue; | |
| 539 | + | } | |
| 528 | 540 | const facts = extract(file.path, blob.value.text, { project: project.name, siblings: found.siblings }); | |
| 529 | 541 | files.push({ path: file.path, facts }); | |
| 530 | 542 | changed.push(file.path); | |
| ⋯ | |||
| 724 | 736 | stats.candidates += captured?.added ?? 0; | |
| 725 | 737 | stats.kept += captured?.kept ?? 0; | |
| 726 | 738 | } | |
| 739 | + | // Candidates from this project's docs that are still waiting but that | |
| 740 | + | // its docs, as read now, no longer suggest: a line since changed, or one | |
| 741 | + | // the rules for what is worth remembering leave out. Kept and dismissed | |
| 742 | + | // memory is never touched. | |
| 743 | + | if (complete) { | |
| 744 | + | const pruned = await memoryReviewClient(this.env.WORK) | |
| 745 | + | .pruneDocCandidates(workspace, repoId, built.hints.map((hint) => hint.text)) | |
| 746 | + | .catch(() => null); | |
| 747 | + | stats.pruned = pruned?.removed ?? 0; | |
| 748 | + | } | |
| 727 | 749 | return stats; | |
| 728 | 750 | } | |
| 729 | 751 | ||
| 1 | 1 | { | |
| 2 | 2 | "extends": "../../tsconfig.base.json", | |
| 3 | + | "compilerOptions": { | |
| 4 | + | // Node runs the tests on these files as they are, which needs the | |
| 5 | + | // extension on relative imports. | |
| 6 | + | "allowImportingTsExtensions": true | |
| 7 | + | }, | |
| 3 | 8 | "include": [ | |
| 4 | 9 | "src/**/*", | |
| 5 | 10 | "worker-configuration.d.ts" |
| 11 | 11 | //! decision candidate. No model is asked: the pull request's own words | |
| 12 | 12 | //! are the decision, and a candidate waits for review anyway. | |
| 13 | 13 | //! - **Docs.** The context service reads a project's README, AGENTS.md, | |
| 14 | − | //! docs and manifests and sends what they say (`capture_memories`). | |
| 14 | + | //! CONTRIBUTING, docs on how to work in it, and manifests, and sends | |
| 15 | + | //! what they say (`capture_memories`). Once it has read every file, the | |
| 16 | + | //! waiting candidates its docs no longer suggest are let go | |
| 17 | + | //! (`prune_doc_candidates`). | |
| 15 | 18 | //! | |
| 16 | 19 | //! The same thing said again, by an independent source, is one memory seen | |
| 17 | − | //! twice, which keeps it (`promotes`). A dismissed memory stays dismissed, | |
| 18 | − | //! so the same wording is never suggested again. Nothing that looks like a | |
| 19 | − | //! secret is stored, in the text or in its evidence. | |
| 20 | + | //! twice, which keeps it (`promotes`). The same thing in other words (a | |
| 21 | + | //! sentence that grew, one cut short) is the memory already there | |
| 22 | + | //! (`same_memory`). A dismissed memory stays dismissed, so the same thing | |
| 23 | + | //! is never suggested again. Nothing that looks like a secret is stored, in | |
| 24 | + | //! the text or in its evidence. | |
| 20 | 25 | ||
| 21 | 26 | use std::collections::HashMap; | |
| 22 | 27 | ||
| ⋯ | |||
| 76 | 81 | "the right way", | |
| 77 | 82 | ]; | |
| 78 | 83 | ||
| 79 | − | #[derive(Deserialize)] | |
| 84 | + | #[derive(Clone, Deserialize)] | |
| 80 | 85 | struct Existing { | |
| 81 | 86 | id: String, | |
| 82 | 87 | status: String, | |
| 83 | 88 | sources: String, | |
| 84 | 89 | confidence: Option<f64>, | |
| 90 | + | text: String, | |
| 91 | + | } | |
| 92 | + | ||
| 93 | + | /// The most memories of one scope compared with a new one for being the | |
| 94 | + | /// same thing in other words: every kept and waiting one (at most 500) and | |
| 95 | + | /// the most recently dismissed. | |
| 96 | + | const MAX_COMPARED: u32 = 2000; | |
| 97 | + | ||
| 98 | + | /// The memory in `among` that says the same thing as `text`, if any. | |
| 99 | + | fn same_as<'a>(among: &'a [Existing], text: &str) -> Option<&'a Existing> { | |
| 100 | + | among.iter().find(|existing| same_memory(&existing.text, text)) | |
| 85 | 101 | } | |
| 86 | 102 | ||
| 103 | + | /// Whether `existing` is a waiting candidate that the source `reference` | |
| 104 | + | /// said before and now says in other words (`print`). | |
| 105 | + | fn reworded_by(existing: &Existing, reference: &str, print: &str) -> bool { | |
| 106 | + | let sources: Vec<String> = serde_json::from_str(&existing.sources).unwrap_or_default(); | |
| 107 | + | existing.status == MemoryStatus::Candidate.as_str() | |
| 108 | + | && fingerprint(&existing.text) != print | |
| 109 | + | && sources.iter().any(|source| source == reference) | |
| 110 | + | } | |
| 111 | + | ||
| 112 | + | /// Whether a candidate came from docs and manifests alone. | |
| 113 | + | fn only_from_docs(sources: &str) -> bool { | |
| 114 | + | let list: Vec<String> = serde_json::from_str(sources).unwrap_or_default(); | |
| 115 | + | !list.is_empty() && list.iter().all(|source| source.starts_with("doc:")) | |
| 116 | + | } | |
| 117 | + | ||
| 118 | + | /// Of a project's waiting doc candidates, the ids its docs no longer | |
| 119 | + | /// suggest: none of `texts` is the same memory. | |
| 120 | + | fn unsuggested(candidates: &[Existing], texts: &[String]) -> Vec<String> { | |
| 121 | + | candidates | |
| 122 | + | .iter() | |
| 123 | + | .filter(|candidate| only_from_docs(&candidate.sources)) | |
| 124 | + | .filter(|candidate| !texts.iter().any(|text| same_memory(&candidate.text, text))) | |
| 125 | + | .map(|candidate| candidate.id.clone()) | |
| 126 | + | .collect() | |
| 127 | + | } | |
| 128 | + | ||
| 87 | 129 | #[derive(Deserialize)] | |
| 88 | 130 | struct RunTicket { | |
| 89 | 131 | id: String, | |
| ⋯ | |||
| 281 | 323 | let workspace = workspace.to_lowercase(); | |
| 282 | 324 | let mut done = Captured::default(); | |
| 283 | 325 | let mut paths = HashMap::new(); | |
| 326 | + | // Each scope's memories, read once a capture, for near-duplicates. | |
| 327 | + | let mut scopes: HashMap<(&'static str, String), Vec<Existing>> = HashMap::new(); | |
| 284 | 328 | for item in items.iter().take(MAX_CAPTURE) { | |
| 285 | 329 | let Ok((text, evidence)) = checked(item) else { | |
| 286 | 330 | done.refused += 1; | |
| ⋯ | |||
| 307 | 351 | } | |
| 308 | 352 | let print = fingerprint(&text); | |
| 309 | 353 | let now = rfc3339(now_ms()); | |
| 310 | − | let existing = self | |
| 354 | + | let mut existing = self | |
| 311 | 355 | .db | |
| 312 | 356 | .prepare( | |
| 313 | − | "SELECT id, status, sources, confidence FROM memories | |
| 357 | + | "SELECT id, status, sources, confidence, text FROM memories | |
| 314 | 358 | WHERE scope = ? AND scope_key = ? AND (fingerprint = ? OR text = ?) LIMIT 1", | |
| 315 | 359 | ) | |
| 316 | 360 | .bind(&[item.scope.as_str().into(), key.as_str().into(), print.as_str().into(), text.as_str().into()])? | |
| 317 | 361 | .first::<Existing>(None) | |
| 318 | 362 | .await?; | |
| 363 | + | // The same thing in other words: a sentence that grew, or one | |
| 364 | + | // cut short, is the memory already there, kept, waiting or | |
| 365 | + | // dismissed. | |
| 366 | + | let scope_key = (item.scope.as_str(), key.clone()); | |
| 367 | + | if existing.is_none() { | |
| 368 | + | if !scopes.contains_key(&scope_key) { | |
| 369 | + | let rows = self | |
| 370 | + | .db | |
| 371 | + | .prepare( | |
| 372 | + | "SELECT id, status, sources, confidence, text FROM memories | |
| 373 | + | WHERE scope = ? AND scope_key = ? | |
| 374 | + | ORDER BY status = 'dismissed', updated_at DESC LIMIT ?", | |
| 375 | + | ) | |
| 376 | + | .bind(&[item.scope.as_str().into(), key.as_str().into(), MAX_COMPARED.into()])? | |
| 377 | + | .all() | |
| 378 | + | .await? | |
| 379 | + | .results::<Existing>()?; | |
| 380 | + | scopes.insert(scope_key.clone(), rows); | |
| 381 | + | } | |
| 382 | + | existing = scopes.get(&scope_key).and_then(|among| same_as(among, &text)).cloned(); | |
| 383 | + | } | |
| 319 | 384 | if let Some(existing) = existing { | |
| 320 | 385 | done.merged += 1; | |
| 321 | 386 | if existing.status == MemoryStatus::Dismissed.as_str() { | |
| 322 | 387 | continue; | |
| 323 | 388 | } | |
| 324 | 389 | let sources = with_source(&existing.sources, &item.reference); | |
| 390 | + | // A waiting candidate its own source now says in other words | |
| 391 | + | // (a doc's line read whole, or edited) takes the new words, | |
| 392 | + | // kind and confidence. Kept memory keeps its words. | |
| 393 | + | let reworded = reworded_by(&existing, &item.reference, &print); | |
| 325 | 394 | let confidence = match (existing.confidence, item.confidence) { | |
| 395 | + | _ if reworded => item.confidence, | |
| 326 | 396 | (Some(a), Some(b)) => Some(a.max(b)), | |
| 327 | 397 | (a, b) => a.or(b), | |
| 328 | 398 | }; | |
| ⋯ | |||
| 332 | 402 | done.kept += 1; | |
| 333 | 403 | } | |
| 334 | 404 | let status = if keep { MemoryStatus::Kept.as_str() } else { existing.status.as_str() }; | |
| 405 | + | let (new_text, new_print) = | |
| 406 | + | if reworded { (text.clone(), print.clone()) } else { (existing.text.clone(), fingerprint(&existing.text)) }; | |
| 335 | 407 | self.db | |
| 336 | 408 | .prepare( | |
| 337 | − | "UPDATE memories SET sources = ?, confidence = ?, status = ?, fingerprint = ?, updated_at = ? | |
| 409 | + | "UPDATE memories SET sources = ?, confidence = ?, status = ?, text = ?, fingerprint = ?, | |
| 410 | + | kind = COALESCE(?, kind), evidence = COALESCE(?, evidence), updated_at = ? | |
| 338 | 411 | WHERE id = ?", | |
| 339 | 412 | ) | |
| 340 | 413 | .bind(&[ | |
| 341 | 414 | serde_json::to_string(&sources)?.into(), | |
| 342 | 415 | confidence.map_or(JsValue::NULL, JsValue::from), | |
| 343 | 416 | status.into(), | |
| 344 | − | print.as_str().into(), | |
| 417 | + | new_text.as_str().into(), | |
| 418 | + | new_print.as_str().into(), | |
| 419 | + | if reworded { item.kind.as_str().into() } else { JsValue::NULL }, | |
| 420 | + | if reworded { evidence.clone().map_or(JsValue::NULL, JsValue::from) } else { JsValue::NULL }, | |
| 345 | 421 | now.as_str().into(), | |
| 346 | 422 | existing.id.as_str().into(), | |
| 347 | 423 | ])? | |
| ⋯ | |||
| 403 | 479 | if keep { | |
| 404 | 480 | done.kept += 1; | |
| 405 | 481 | } | |
| 482 | + | // The same thing twice in one capture is one memory. | |
| 483 | + | if let Some(among) = scopes.get_mut(&scope_key) { | |
| 484 | + | among.push(Existing { | |
| 485 | + | id: id.clone(), | |
| 486 | + | status: status.as_str().to_owned(), | |
| 487 | + | sources: serde_json::to_string(&[&item.reference])?, | |
| 488 | + | confidence: item.confidence, | |
| 489 | + | text: text.clone(), | |
| 490 | + | }); | |
| 491 | + | } | |
| 406 | 492 | self.memory_changed(&id, &workspace, status.as_str(), repo.as_deref()).await; | |
| 407 | 493 | } | |
| 408 | 494 | Ok(done) | |
| 409 | 495 | } | |
| 410 | 496 | ||
| 497 | + | /// Removes a project's waiting candidates that came from its docs alone | |
| 498 | + | /// and that its docs, as read now, no longer suggest. Kept and dismissed | |
| 499 | + | /// memory is never touched, nor a candidate another source saw too. | |
| 500 | + | pub(crate) async fn prune_doc_candidates(&self, a: PruneDocCandidatesArgs) -> Result<Pruned> { | |
| 501 | + | let workspace = a.workspace.to_lowercase(); | |
| 502 | + | let candidates = self | |
| 503 | + | .db | |
| 504 | + | .prepare( | |
| 505 | + | "SELECT id, status, sources, confidence, text FROM memories | |
| 506 | + | WHERE workspace = ? AND scope = 'project' AND scope_key = ? AND status = 'candidate' | |
| 507 | + | AND source_kind = 'doc' | |
| 508 | + | LIMIT ?", | |
| 509 | + | ) | |
| 510 | + | .bind(&[workspace.as_str().into(), a.repo_id.as_str().into(), MAX_COMPARED.into()])? | |
| 511 | + | .all() | |
| 512 | + | .await? | |
| 513 | + | .results::<Existing>()?; | |
| 514 | + | let gone = unsuggested(&candidates, &a.texts); | |
| 515 | + | if gone.is_empty() { | |
| 516 | + | return Ok(Pruned::default()); | |
| 517 | + | } | |
| 518 | + | let removed = self | |
| 519 | + | .db | |
| 520 | + | .prepare( | |
| 521 | + | "DELETE FROM memories WHERE workspace = ? AND scope_key = ? AND status = 'candidate' | |
| 522 | + | AND id IN (SELECT value FROM json_each(?)) RETURNING id", | |
| 523 | + | ) | |
| 524 | + | .bind(&[workspace.as_str().into(), a.repo_id.as_str().into(), serde_json::to_string(&gone)?.into()])? | |
| 525 | + | .all() | |
| 526 | + | .await? | |
| 527 | + | .results::<serde_json::Value>()? | |
| 528 | + | .len(); | |
| 529 | + | worker::console_log!("pruned {removed} doc candidates of {} in {workspace}", a.repo_id); | |
| 530 | + | Ok(Pruned { removed: removed as u32 }) | |
| 531 | + | } | |
| 532 | + | ||
| 411 | 533 | pub(crate) async fn capture_memories(&self, a: CaptureMemoriesArgs) -> Result<Captured> { | |
| 412 | 534 | self.capture(&a.workspace, &a.items, a.by.as_deref().unwrap_or("g1t")).await | |
| 413 | 535 | } | |
| ⋯ | |||
| 744 | 866 | } | |
| 745 | 867 | ||
| 746 | 868 | /// The methods this module serves at `POST /rpc/<method>`. | |
| 747 | − | pub(crate) const METHODS: [&str; 7] = [ | |
| 869 | + | pub(crate) const METHODS: [&str; 8] = [ | |
| 748 | 870 | "capture_memories", | |
| 871 | + | "prune_doc_candidates", | |
| 749 | 872 | "report_learned", | |
| 750 | 873 | "list_candidates", | |
| 751 | 874 | "review_memory", | |
| ⋯ | |||
| 758 | 881 | use g1t_kit::{args, reply}; | |
| 759 | 882 | match method { | |
| 760 | 883 | "capture_memories" => reply(&work.capture_memories(args(body)?).await?), | |
| 884 | + | "prune_doc_candidates" => reply(&work.prune_doc_candidates(args(body)?).await?), | |
| 761 | 885 | "report_learned" => reply(&work.report_learned(args(body)?).await?), | |
| 762 | 886 | "list_candidates" => reply(&work.list_candidates(args(body)?).await?), | |
| 763 | 887 | "review_memory" => reply(&work.review_memory(args(body)?).await?), | |
| ⋯ | |||
| 834 | 958 | assert_eq!(with_source("not json", "doc:x"), ["doc:x"]); | |
| 835 | 959 | } | |
| 836 | 960 | ||
| 961 | + | fn existing(id: &str, text: &str, sources: &str) -> Existing { | |
| 962 | + | Existing { id: id.into(), status: "candidate".into(), sources: sources.into(), confidence: Some(0.7), text: text.into() } | |
| 963 | + | } | |
| 964 | + | ||
| 965 | + | #[test] | |
| 966 | + | fn the_same_thing_in_other_words_is_found() { | |
| 967 | + | let among = [ | |
| 968 | + | existing("a", "Run cargo test in the crate you changed.", "[]"), | |
| 969 | + | existing("b", "g1t is a Cargo workspace (apps/api, crates/*, services/actions); `cargo test` runs its tests.", "[]"), | |
| 970 | + | ]; | |
| 971 | + | assert_eq!(same_as(&among, "run cargo test in the crate you changed").map(|e| e.id.as_str()), Some("a")); | |
| 972 | + | assert_eq!( | |
| 973 | + | same_as(&among, "g1t is a Cargo workspace (apps/*, crates/*, services/*); `cargo test` runs its tests.").map(|e| e.id.as_str()), | |
| 974 | + | Some("b") | |
| 975 | + | ); | |
| 976 | + | assert!(same_as(&among, "Use pnpm, never npm.").is_none()); | |
| 977 | + | } | |
| 978 | + | ||
| 979 | + | #[test] | |
| 980 | + | fn a_candidate_cut_short_takes_its_doc_s_whole_words() { | |
| 981 | + | let cut = existing("a", "Work started by an automation cannot retrigger the", "[\"doc:r:CONTRIBUTING.md\"]"); | |
| 982 | + | let whole = fingerprint("Work started by an automation cannot retrigger the same automation."); | |
| 983 | + | assert!(reworded_by(&cut, "doc:r:CONTRIBUTING.md", &whole)); | |
| 984 | + | // Another source saying it is a sighting, not new words. | |
| 985 | + | assert!(!reworded_by(&cut, "run:x", &whole)); | |
| 986 | + | // Kept memory keeps the words someone kept. | |
| 987 | + | let kept = Existing { status: "kept".into(), ..cut.clone() }; | |
| 988 | + | assert!(!reworded_by(&kept, "doc:r:CONTRIBUTING.md", &whole)); | |
| 989 | + | // The same words are not new ones. | |
| 990 | + | assert!(!reworded_by(&cut, "doc:r:CONTRIBUTING.md", &fingerprint(&cut.text))); | |
| 991 | + | } | |
| 992 | + | ||
| 993 | + | #[test] | |
| 994 | + | fn waiting_doc_candidates_the_docs_no_longer_suggest_are_let_go() { | |
| 995 | + | let candidates = [ | |
| 996 | + | // Cut mid-sentence by the old reader; the docs now say it whole. | |
| 997 | + | existing("cut", "Run the whole suite before you push, every", "[\"doc:r:CONTRIBUTING.md\"]"), | |
| 998 | + | // From a plan, which is no longer read for memory. | |
| 999 | + | existing("plan", "Loop protection. Work started by an automation cannot retrigger the", "[\"doc:r:docs/PLAN.md\"]"), | |
| 1000 | + | // Still suggested. | |
| 1001 | + | existing("still", "`cargo test` in the crate or service you changed.", "[\"doc:r:CONTRIBUTING.md\"]"), | |
| 1002 | + | // Seen by a run too: not the docs' alone to take back. | |
| 1003 | + | existing("run", "Deduplication. The same Sentry issue firing 500 times maps to one", "[\"doc:r:docs/PLAN.md\",\"run:x\"]"), | |
| 1004 | + | ]; | |
| 1005 | + | let texts = [ | |
| 1006 | + | "Run the whole suite before you push, every time, on every branch.".to_owned(), | |
| 1007 | + | "`cargo test` in the crate or service you changed.".to_owned(), | |
| 1008 | + | ]; | |
| 1009 | + | assert_eq!(unsuggested(&candidates, &texts), ["plan"]); | |
| 1010 | + | assert!(only_from_docs("[\"doc:a\",\"doc:b\"]")); | |
| 1011 | + | assert!(!only_from_docs("[]")); | |
| 1012 | + | assert!(!only_from_docs("[\"doc:a\",\"review:b\"]")); | |
| 1013 | + | } | |
| 1014 | + | ||
| 837 | 1015 | fn item(text: &str, evidence: Option<&str>) -> CaptureItem { | |
| 838 | 1016 | CaptureItem { | |
| 839 | 1017 | scope: MemoryScope::Project, | |