Skip to content

Commit

Merge memory from docs: only docs on how to work here, whole sentences, no near-duplicates, stale doc suggestions cleared

syntaqxcommitted Parentsd9ea97d1cba074Browse files
14 files+900−750/14 viewed
+5−3
149149
150150 Memory also fills itself. At the end of every run that changes code, the
151151 agent is asked what it learned; a person's correction in a review, a merged
152−pull request's decision, and what a project's `AGENTS.md`, README and
153−manifests say are captured too. These arrive as **candidates**, which no
152+pull request's decision, and what a project's `AGENTS.md`, README,
153+CONTRIBUTING, docs on how to work in it and manifests say are captured
154+too ([which docs](/guides/context-hub/#which-docs-are-read-for-memory)).
155+These arrive as **candidates**, which no
154156 agent is given until they are kept: at once when two independent sources
155157 say the same thing or a project's `AGENTS.md` or manifests state it,
156158 otherwise by a member in the **Review** list on **Agents → Memory** or the
190192 - **edit** one that has drifted, or change its kind;
191193 - **forget** one that no longer holds;
192194 - **keep**, **edit** or **dismiss** a candidate waiting for review. A
193− dismissed candidate is never suggested again in the same words.
195+ dismissed candidate is never suggested again, in the same words or near enough to them.
194196
195197 Each shows who added it, when, and when it was last given to an agent. A
196198 memory that has not been given to an agent for a long time is a good one to
+57−4
3636 | `README`, `AGENTS.md` (or `CLAUDE.md`), `CONTRIBUTING`, `docs/*.md`, `runbooks/*.md` | **Docs**, split into sections for search |
3737 | `.g1t/workflows/*`, `.github/workflows/*` | Whether the project's tests run in checks |
3838 | `owners:` in `.g1t/project.yml`, and `CODEOWNERS` | **Owners** |
39+| `memory:` in `.g1t/project.yml` | Which docs are [read for memory](#which-docs-are-read-for-memory) |
3940
4041 and joins them with what g1t already knows:
4142
7677 | **Agent runs** | At the end of every run that changes code, the agent is asked what it learned that the next agent would need, with what showed it | fact, convention, decision or gotcha |
7778 | **Reviews** | A person's request for changes, or a comment that corrects the agent ("we use the shared client instead"), on a pull request | convention, quoting the comment |
7879 | **Merges** | A merged pull request's title, why (the first paragraph of its description) and the files it changed | decision |
79−| **Docs and manifests** | Bullets in `AGENTS.md`; commands under a README's setup and testing sections; conventions sections; the package manager and test commands from manifests | fact, convention or gotcha |
80+| **Docs and manifests** | Bullets in `AGENTS.md`; commands and bullets under the setup, testing and conventions sections of docs on how to work in the project; the package manager and test commands from manifests | fact, convention or gotcha |
8081
8182 What is captured arrives as a **candidate**. No agent is given a candidate
8283 until it is **kept**:
8788 - **by a person**, in the Review queue.
8889
8990 The same thing said again in different case, punctuation or spacing counts
90−as the same memory, seen once more. A memory seen again from the same
91+as the same memory, seen once more. So does the same thing in slightly
92+different words: a sentence that grew (a list with one more item), or one
93+that holds all of another's words. A memory seen again from the same
9194 source (the same run, the same file) is not counted twice.
9295
96+### Which docs are read for memory
97+
98+Memory is for how to work in a project, so only docs that say how are
99+read for it:
100+
101+| Read for memory | Not read for memory |
102+| --- | --- |
103+| `README`, `AGENTS.md` (or `CLAUDE.md`), `CONTRIBUTING`, `runbooks/*.md` | Plans and roadmaps (`PLAN.md`, `roadmap.md`) |
104+| A doc in `docs/` whose name says how to work: `DEPLOYING.md`, `testing.md`, `architecture.md`, `setup.md`, `conventions.md`, `getting-started.md` | Changelogs, release notes, history, incidents, demos, research, notes, feedback |
105+
106+Within those docs:
107+
108+- Only the **setup, testing and conventions** sections are read (in
109+ `AGENTS.md`, every section). A section about plans, a roadmap, asks or
110+ feedback is skipped with everything under it, wherever it is.
111+- A doc that reads as a plan or a report (most of its headings are plans,
112+ or its bullets are labelled "What we tried", "Ask" and the like) gives
113+ nothing, whatever its name.
114+- A line that says what someone wants rather than how things are ("g1t
115+ will", "should", "TODO") is left out. In `AGENTS.md`, "should" states a
116+ rule, and is kept.
117+- A bullet is taken whole, with the lines it wraps onto. A long one is cut
118+ only between sentences, to 300 characters; one whose first sentence is
119+ longer is left out rather than cut.
120+- A bullet that tells you to do or never do something is a convention or
121+ a gotcha; anything else is a fact. Outside `AGENTS.md`, a bullet waits
122+ for review at 50 to 60% sure.
123+
124+To read another doc for memory, or never read one, say so in the
125+project's `.g1t/project.yml`, by path or a whole folder:
126+
127+```yaml
128+memory:
129+ docs:
130+ - docs/PERFORMANCE.md
131+ skip:
132+ - README.md
133+ - docs/legacy/*
134+```
135+
136+`skip` wins over `docs`. A change to these lists applies to each doc the
137+next time it changes, or for every doc at once with **Rebuild**.
138+
139+Once g1t has read every file of a project as these rules read it, a
140+candidate from its docs that is still waiting and that its docs no longer
141+suggest (the doc changed, or the rules leave the line out) is removed from
142+the queue. A candidate another source saw too, and every kept or dismissed
143+memory, stays. A waiting candidate its doc now says in other words takes
144+the new words.
145+
93146 ### Review
94147
95148 The **Memory** tab of the Context page lists every candidate in the
99152
100153 - **Keep** gives it to every agent from their next run on.
101154 - **Edit** rewords it, or changes its kind, and keeps it.
102−- **Dismiss** throws it away, and the same wording is never suggested
103− again.
155+- **Dismiss** throws it away, and the same thing, in the same words or
156+ near enough to them, is never suggested again.
104157
105158 ### Never a secret
106159
+15−1
11 import assert from "node:assert/strict";
22 import { test } from "node:test";
33
4−import { distinctFacts } from "./memory-facts.ts";
4+import { distinctFacts, sameFact } from "./memory-facts.ts";
55
66 test("a fact remembered twice is shown once, the first kept", () => {
77 const facts = [
1313 assert.deepEqual(distinctFacts(facts).map((fact) => fact.id), ["a", "b"]);
1414 assert.deepEqual(distinctFacts([]), []);
1515 });
16+
17+test("near-duplicates collapse: a list that grew, a sentence cut short", () => {
18+ const facts = [
19+ { id: "a", text: "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing); `cargo test` runs its tests." },
20+ { id: "b", text: "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing, services/work); `cargo test` runs its tests." },
21+ { id: "c", text: "Roles. Viewer, commenter, planner and approver map onto" },
22+ { id: "d", text: "Roles. Viewer, commenter, planner and approver map onto the five repository roles." },
23+ { id: "e", text: "Use npm to install." },
24+ { id: "f", text: "Use pnpm to install." },
25+ ];
26+ assert.deepEqual(distinctFacts(facts).map((fact) => fact.id), ["a", "c", "e", "f"]);
27+ assert.ok(!sameFact("Run cargo test in the crate you changed.", "Run npm test in the app you changed."));
28+ assert.ok(!sameFact("use pnpm", "use pnpm in the web app and npm in the docs, which predates it"));
29+});
+45−11
1+/** How alike two facts' words must be (shared over all, as sets) to be one fact. */
2+const SAME_WORDS = 0.85;
3+/** The fewest distinct words a fact needs before it can be found inside another. */
4+const CONTAINED_MIN_WORDS = 6;
5+
6+/** A fact's words, lowercase, one space apart: case, punctuation and spacing aside. */
7+function folded(text: string): string {
8+ return text
9+ .toLowerCase()
10+ .split(/[^\p{L}\p{N}]+/u)
11+ .filter(Boolean)
12+ .join(" ");
13+}
14+
15+/**
16+ * Whether two facts say the same thing, as the work service decides it
17+ * (`same_memory`): the same words; one's words, in order, inside the
18+ * other's; all of one's words (at least six) among the other's; or most of
19+ * their words shared.
20+ */
21+export function sameFact(a: string, b: string): boolean {
22+ const [x, y] = [folded(a), folded(b)];
23+ if (!x || !y) return false;
24+ if (x === y) return true;
25+ const [short, long] = x.length <= y.length ? [x, y] : [y, x];
26+ if (short.split(" ").length >= 4 && ` ${long} `.includes(` ${short} `)) return true;
27+ const shortWords = new Set(short.split(" "));
28+ const longWords = new Set(long.split(" "));
29+ let shared = 0;
30+ for (const word of shortWords) if (longWords.has(word)) shared++;
31+ if (shortWords.size >= CONTAINED_MIN_WORDS && shared === shortWords.size) return true;
32+ const all = new Set([...shortWords, ...longWords]).size;
33+ return all > 0 && shared / all >= SAME_WORDS;
34+}
35+
136 /**
2− * The same fact remembered twice (two runs that learned it, or one saved
3− * again) is shown once: the first of them in the order given, so a pinned
4− * one or the newest wins. Facts match when their words do, whatever the
5− * case, spacing or a closing full stop.
37+ * The same fact remembered twice (two runs that learned it, one saved
38+ * again, or a sentence that grew since) is shown once: the first of them in
39+ * the order given, so a pinned one or the newest wins. Facts match when
40+ * their words do, whatever the case, spacing or punctuation, or when they
41+ * are near enough to be one (see `sameFact`).
642 */
743 export function distinctFacts<T extends { text: string }>(facts: T[]): T[] {
8− const seen = new Set<string>();
9− return facts.filter((fact) => {
10− const key = fact.text.trim().toLowerCase().replace(/\s+/g, " ").replace(/[.!]+$/, "");
11− if (seen.has(key)) return false;
12− seen.add(key);
13− return true;
14− });
44+ const shown: T[] = [];
45+ for (const fact of facts) {
46+ if (!shown.some((other) => sameFact(other.text, fact.text))) shown.push(fact);
47+ }
48+ return shown;
1549 }
+76−0
7575 .join(" ")
7676 }
7777
78+/// How alike two memories' words must be (shared over all, as sets) to be
79+/// one memory.
80+pub const SAME_WORDS: f64 = 0.85;
81+/// The fewest distinct words a memory needs before it can be found inside
82+/// another: shorter ones are too general to be the same thing.
83+const CONTAINED_MIN_WORDS: usize = 6;
84+
85+/// Whether two memories say the same thing: the same words once folded
86+/// ([`fingerprint`]); one's words, in order, inside the other's; all of a
87+/// memory's words (at least six of them) among the other's; or most of
88+/// their words shared (Jaccard at or above [`SAME_WORDS`]). "g1t is a
89+/// Cargo workspace (apps/*, crates/*)" and the same sentence listing every
90+/// crate are one memory.
91+pub fn same_memory(a: &str, b: &str) -> bool {
92+ let (a, b) = (fingerprint(a), fingerprint(b));
93+ if a.is_empty() || b.is_empty() {
94+ return false;
95+ }
96+ if a == b {
97+ return true;
98+ }
99+ let (short, long) = if a.len() <= b.len() { (&a, &b) } else { (&b, &a) };
100+ let short_words: std::collections::HashSet<&str> = short.split(' ').collect();
101+ let long_words: std::collections::HashSet<&str> = long.split(' ').collect();
102+ // Contiguous: the shorter one's words, in order, inside the longer one's.
103+ if short.split(' ').count() >= 4 && format!(" {long} ").contains(&format!(" {short} ")) {
104+ return true;
105+ }
106+ if short_words.len() >= CONTAINED_MIN_WORDS && short_words.is_subset(&long_words) {
107+ return true;
108+ }
109+ let shared = short_words.intersection(&long_words).count();
110+ let all = short_words.union(&long_words).count();
111+ all > 0 && shared as f64 / all as f64 >= SAME_WORDS
112+}
113+
78114 /// One thing learned, as a source reports it.
79115 #[derive(Clone, Debug, Serialize, Deserialize)]
80116 #[serde(rename_all = "camelCase")]
229265 pub limit: Option<u32>,
230266 }
231267
268+/// `prune_doc_candidates`: after the context service has read every file
269+/// of a project, removes the project's candidates that came only from its
270+/// docs and manifests and are still waiting, unless one of `texts` (what
271+/// they suggest now) is the same memory ([`same_memory`]). Kept and
272+/// dismissed memory is never touched. Services only. Returns `Pruned`.
273+#[derive(Debug, Serialize, Deserialize)]
274+#[serde(rename_all = "camelCase")]
275+pub struct PruneDocCandidatesArgs {
276+ pub workspace: String,
277+ pub repo_id: String,
278+ #[serde(default)]
279+ pub texts: Vec<String>,
280+}
281+
282+#[derive(Clone, Debug, Default, PartialEq, Eq, Serialize, Deserialize)]
283+pub struct Pruned {
284+ /// Candidates removed.
285+ pub removed: u32,
286+}
287+
232288 /// Memories, newest first, for listing a review queue.
233289 pub type Candidates = Vec<Memory>;
234290
272328 assert_eq!(fingerprint("use pnpm never NPM"), fingerprint("Use pnpm; never npm."));
273329 assert_ne!(fingerprint("use pnpm"), fingerprint("use npm"));
274330 }
331+
332+ #[test]
333+ fn near_duplicates_are_one_memory() {
334+ assert!(same_memory("Use pnpm, never npm.", "use pnpm never NPM"));
335+ // A crate list that grew, and the same fact without the list.
336+ let before = "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing); `cargo test` runs its tests.";
337+ let after = "g1t is a Cargo workspace (apps/api, crates/*, services/actions, services/billing, services/work); `cargo test` runs its tests.";
338+ assert!(same_memory(before, after));
339+ assert!(same_memory(before, "g1t is a Cargo workspace (apps/*, crates/*, services/*); `cargo test` runs its tests."));
340+ // One cut short, inside the whole of it.
341+ assert!(same_memory(
342+ "Roles. Viewer, commenter, planner and approver map onto",
343+ "Roles. Viewer, commenter, planner and approver map onto the five repository roles."
344+ ));
345+ // Different facts that share words are not.
346+ assert!(!same_memory("Use pnpm to install.", "Use npm to install."));
347+ assert!(!same_memory("Run cargo test in the crate you changed.", "Run npm test in the app you changed."));
348+ assert!(!same_memory("use pnpm", "use pnpm in the web app and npm in the docs, which predates it"));
349+ assert!(!same_memory("", "anything"));
350+ }
275351 }
+7−0
237237 ): Promise<Result<Memory>>;
238238 /** For services: candidates from docs and backfills. */
239239 captureMemories(workspace: string, items: CaptureItem[], by?: string): Promise<Captured>;
240+ /**
241+ * For services: removes a project's candidates that came only from its
242+ * docs and are still waiting, unless one of `texts` (what its docs suggest
243+ * now) says the same thing. Kept and dismissed memory is never touched.
244+ */
245+ pruneDocCandidates(workspace: string, repoId: string, texts: string[]): Promise<{ removed: number }>;
240246 /** For services: memories by id, in any status. */
241247 memoriesById(workspace: string, ids: string[]): Promise<Memory[]>;
242248 /** For services: kept memories with every word of `query`; the caller has checked the viewer may read them. */
252258 listCandidates: (viewer, workspace, repo) => call("list_candidates", { viewer, workspace, repo: repo ?? null }),
253259 reviewMemory: (actor, workspace, id, decision, change = {}) => call("review_memory", { actor, workspace, id, decision, ...change }),
254260 captureMemories: (workspace, items, by) => call("capture_memories", { workspace, items, by: by ?? null }),
261+ pruneDocCandidates: (workspace, repoId, texts) => call("prune_doc_candidates", { workspace, repoId, texts }),
255262 memoriesById: (workspace, ids) => call("memories_by_id", { workspace, ids }),
256263 searchMemories: (workspace, query, options = {}) =>
257264 call("search_memories", { workspace, query, repoIds: options.repoIds ?? null, limit: options.limit ?? 20 }),
+7−1
99 import type { EntityKind, RelationKind } from "@g1t/contracts";
1010
1111 import type { FileFacts, Hint } from "./extract";
12+import { harvestable } from "./harvest.ts";
1213
1314 export type ProjectInput = {
1415 id: string;
198199 if (integration.repo?.toLowerCase() === repoPath) relate(me, "uses", { kind: "integration", key: integration.id });
199200 }
200201
201− const hints = files.flatMap((file) => file.facts.hints.map((hint) => ({ ...hint, path: file.path })));
202+ // A doc is a source of memory only when it says how to work here (see
203+ // ./harvest), or project.yml says it is.
204+ const memory = files.find((file) => file.facts.memory)?.facts.memory ?? null;
205+ const hints = files
206+ .filter((file) => !file.facts.doc || harvestable(file.path, file.facts.doc.role, memory))
207+ .flatMap((file) => file.facts.hints.map((hint) => ({ ...hint, path: file.path })));
202208 return { entities, relations, hints, tests };
203209 }
204210
+5−2
5353 assert.deepEqual(crate.packages[0].dependencies, ["serde", "worker.workspace", "pretty"]);
5454 const workspace = extract("Cargo.toml", `[workspace]\nmembers = [\n "crates/a",\n "services/b",\n]\n`, ctx);
5555 assert.deepEqual(workspace.packages[0].members, ["crates/a", "services/b"]);
56− assert.match(workspace.hints[0].text, /Cargo workspace \(crates\/a, services\/b\)/);
56+ // Where crates live, not each one, so adding a crate says nothing new.
57+ assert.match(workspace.hints[0].text, /Cargo workspace \(crates\/\*, services\/\*\)/);
58+ const grown = extract("Cargo.toml", `[workspace]\nmembers = ["crates/a", "crates/c", "services/b"]\n`, ctx);
59+ assert.equal(grown.hints[0].text, workspace.hints[0].text);
5760 });
5861
5962 test("go.mod and pyproject.toml", () => {
8891 assert.equal(agents.doc?.role, "agents");
8992 const kinds = agents.hints.map((hint) => [hint.kind, hint.text, hint.confidence]);
9093 assert.deepEqual(kinds, [
94+ ["fact", "In web, for working here: `npm run db:reset`.", 0.9],
9195 ["convention", "Components live in src/components/ui; import from there.", 0.9],
9296 ["gotcha", "Never edit generated.rs by hand.", 0.9],
93− ["fact", "In web, for working here: `npm run db:reset`.", 0.9],
9497 ]);
9598 const readme = extract("README.md", "# Web\n\nThe storefront.\n\n## Testing\n\n```\n$ TZ=UTC npm test\nnpm test\n```\n\n## Features\n\n- Lots of features that are really great\n", ctx);
9699 assert.equal(readme.doc?.title, "Web");
+71−36
99
1010 import { parse as parseYaml } from "yaml";
1111
12+import { aspirational, bulletsOf, classify, bulletSection, MAX_DOC_HINT_CHARS, planLike, planSection, plain, SETUP_HEADING, wholeSentences, type MemoryDocsConfig } from "./harvest.ts";
13+
1214 export type Ecosystem = "npm" | "cargo" | "go" | "pypi";
1315
1416 export type PackageFact = {
6163 /** Usernames this file names as owners. */
6264 owners: string[];
6365 hints: Hint[];
66+ /** For `.g1t/project.yml`: which docs memory is suggested from. */
67+ memory?: MemoryDocsConfig | null;
68+ /** The `EXTRACT_VERSION` that read it; facts from an older one are read again. */
69+ version?: number;
6470 };
6571
6672 /** What extraction knows besides the file itself. */
7177 siblings: string[];
7278 };
7379
80+/** Bumped when what is read from a file changes, so stored facts are read again. */
81+export const EXTRACT_VERSION = 2;
7482 export const MAX_DEPENDENCIES = 60;
7583 export const CHUNK_CHARS = 1500;
7684 export const MAX_CHUNKS = 12;
7785 const MAX_ROUTES = 40;
7886 const MAX_HINTS_PER_FILE = 12;
79−const MAX_HINT_CHARS = 300;
87+const MAX_HINT_CHARS = MAX_DOC_HINT_CHARS;
8088
8189 const EMPTY: FileFacts = { languages: [], packages: [], apis: [], doc: null, tests: false, owners: [], hints: [] };
8290
279287 const name = str(toml.package?.name);
280288 const members = Array.isArray(toml.workspace?.members) ? (toml.workspace.members as string[]) : [];
281289 const deps = [...keys(toml.dependencies), ...keys(toml["dev-dependencies"]), ...keys(toml["workspace.dependencies"])];
290+ // Where its crates live (apps/*, crates/*), not each one: the list
291+ // changes with every crate added, and the fact does not.
292+ const places = [...new Set(members.map((member) => (member.includes("/") ? `${member.slice(0, member.indexOf("/"))}/*` : member)))];
282293 const hints: Hint[] = [
283294 {
284295 kind: "fact",
285296 text: clip(
286297 members.length
287− ? `${ctx.project} is a Cargo workspace (${members.join(", ")}); \`cargo test\` runs its tests.`
298+ ? `${ctx.project} is a Cargo workspace (${places.join(", ")}); \`cargo test\` runs its tests.`
288299 : `${ctx.project} is written in Rust; \`cargo test\` runs its tests.`,
289300 ),
290301 confidence: 0.9,
431442 }
432443
433444 const COMMAND = /^\s*(?:\$\s*)?((?:npm|pnpm|yarn|bun|npx|cargo|go|make|pytest|python3?|poetry|uv|docker|wrangler|just|mix|bundle|rake|gradle|\.\/gradlew|mvn|dotnet)\b[^\n]{0,160})$/;
434−const SETUP_HEADING = /\b(develop|development|getting started|setup|set up|install|build|test|testing|contribut|running|run locally|local)\b/i;
435−const CONVENTION_HEADING = /\b(convention|guideline|rules|style|standards|gotcha|pitfall|caveat|known issue|do not|don't)\b/i;
436−const GOTCHA = /\b(never|don't|do not|must not|careful|beware|gotcha|warning|avoid)\b/i;
437445
438−/** What a doc says worth remembering: commands in its setup sections, and its conventions. */
446+/**
447+ * What a doc says worth remembering: the commands in its setup sections,
448+ * and the bullets of its conventions and setup sections, each whole (see
449+ * `./harvest`). A doc that reads as a plan, a report or feedback says
450+ * nothing; nor does a plan section in any doc, or a line that says what
451+ * someone wants rather than how things are.
452+ */
439453 function docHints(path: string, text: string, role: DocFact["role"], ctx: ExtractContext): Hint[] {
454+ const agents = role === "agents";
455+ if (!agents && planLike(text)) return [];
440456 const hints: Hint[] = [];
441− const agents = role === "agents";
457+ const where = (heading: string) => `${path}${heading ? ` (${heading})` : ""}`;
458+
459+ // Commands in code blocks under a setup heading (any, in AGENTS.md).
442460 let heading = "";
461+ let skipping = 0;
443462 let fenced = false;
444463 for (const raw of text.split(/\r?\n/)) {
445464 const line = raw.trimEnd();
447466 fenced = !fenced;
448467 continue;
449468 }
450− const h = /^#{1,6}\s+(.*)$/.exec(line);
469+ const h = /^(#{1,6})\s+(.*)$/.exec(line);
451470 if (h && !fenced) {
452− heading = h[1].trim();
471+ const level = h[1].length;
472+ if (skipping && level <= skipping) skipping = 0;
473+ heading = h[2].trim();
474+ if (!skipping && planSection(heading)) skipping = level;
453475 continue;
454476 }
455− if (fenced) {
456− const command = COMMAND.exec(line)?.[1];
457− if (command && (agents || SETUP_HEADING.test(heading))) {
458− const what = heading ? heading.replace(/[:.]$/, "").toLowerCase() : "work on it";
459− hints.push({
460− kind: "fact",
461− text: clip(`In ${ctx.project}, for ${what}: \`${command.trim()}\`.`),
462− confidence: agents ? 0.9 : 0.7,
463− evidence: `${path}${heading ? ` (${heading})` : ""}`,
464− });
465− }
466− continue;
477+ if (!fenced || skipping) continue;
478+ const command = COMMAND.exec(line)?.[1];
479+ if (command && (agents || SETUP_HEADING.test(heading))) {
480+ const what = heading ? plain(heading).replace(/[:.]$/, "").toLowerCase() : "work on it";
481+ const said = wholeSentences(`In ${ctx.project}, for ${what}: \`${command.trim()}\`.`, MAX_HINT_CHARS);
482+ if (said) hints.push({ kind: "fact", text: said, confidence: agents ? 0.9 : 0.7, evidence: where(heading) });
467483 }
468− const bullet = /^\s*[-*+]\s+(.*)$/.exec(line)?.[1]?.trim();
469− if (!bullet || bullet.length < 20 || bullet.length > MAX_HINT_CHARS) continue;
484+ }
485+
486+ // Bullets, whole, from the sections that say how things are done.
487+ for (const bullet of bulletsOf(text)) {
488+ if (!bulletSection(bullet.heading, role)) continue;
470489 // A bullet that is only a link is a table of contents.
471− if (/^\[[^\]]*\]\([^)]*\)\.?$/.test(bullet)) continue;
472− if (agents || CONVENTION_HEADING.test(heading)) {
473− hints.push({
474− kind: GOTCHA.test(bullet) ? "gotcha" : "convention",
475− text: clip(bullet.replace(/\*\*/g, "")),
476− confidence: agents ? 0.9 : 0.7,
477− evidence: `${path}${heading ? ` (${heading})` : ""}`,
478− });
479− }
490+ if (/^\[[^\]]*\]\([^)]*\)\.?$/.test(bullet.text)) continue;
491+ if (aspirational(bullet.text, role)) continue;
492+ const words = plain(bullet.text);
493+ // Too short to mean anything, or a lead-in to a list of its own.
494+ if (words.length < 20 || /:$/.test(words)) continue;
495+ const said = wholeSentences(words, MAX_HINT_CHARS);
496+ if (!said) continue;
497+ hints.push({ ...classify(said, bullet.heading, role), text: said, evidence: where(bullet.heading) });
480498 }
481− // The same command under two headings is one hint.
499+ // The same line under two headings is one hint.
482500 const seen = new Set<string>();
483− return hints.filter((hint) => !seen.has(hint.text) && seen.add(hint.text)).slice(0, MAX_HINTS_PER_FILE);
501+ return hints
502+ .filter((hint) => {
503+ const key = hint.text.toLowerCase().replace(/[^\p{L}\p{N}]+/gu, " ").trim();
504+ return !seen.has(key) && seen.add(key);
505+ })
506+ .slice(0, MAX_HINTS_PER_FILE);
484507 }
485508
486509 function doc(path: string, text: string, ctx: ExtractContext): FileFacts {
504527 };
505528 }
506529
530+/** Strings, from a YAML value that is a string or a list of them. */
531+function strings(value: unknown): string[] {
532+ const list = Array.isArray(value) ? value : typeof value === "string" ? [value] : [];
533+ return list.filter((item): item is string => typeof item === "string" && item.trim() !== "").map((item) => item.trim()).slice(0, 50);
534+}
535+
507536 function owners(path: string, text: string): FileFacts {
508537 const names = new Set<string>();
538+ let memory: MemoryDocsConfig | null = null;
509539 if (path.endsWith("project.yml")) {
510− const parsed = parseYaml(text) as { owners?: unknown } | null;
540+ const parsed = parseYaml(text) as { owners?: unknown; memory?: { docs?: unknown; skip?: unknown } } | null;
541+ if (parsed?.memory && typeof parsed.memory === "object") memory = { docs: strings(parsed.memory.docs), skip: strings(parsed.memory.skip) };
511542 const list = Array.isArray(parsed?.owners) ? parsed.owners : typeof parsed?.owners === "string" ? [parsed.owners] : [];
512543 for (const owner of list) if (typeof owner === "string") names.add(owner.replace(/^@/, "").trim());
513544 } else {
516547 for (const match of line.matchAll(/@([A-Za-z0-9][A-Za-z0-9_.-]*)(?=\s|$)/g)) names.add(match[1]);
517548 }
518549 }
519− return { ...EMPTY, owners: [...names].filter(Boolean).slice(0, 20) };
550+ return { ...EMPTY, owners: [...names].filter(Boolean).slice(0, 20), ...(memory ? { memory } : {}) };
520551 }
521552
522553 const TESTS = /\b(test|tests|pytest|vitest|jest|mocha|cargo test|go test|npm test|playwright|cypress)\b/i;
523554
524555 /** The facts in one file. A file that cannot be parsed says nothing. */
525556 export function extract(path: string, text: string, ctx: ExtractContext): FileFacts {
557+ return { ...read(path, text, ctx), version: EXTRACT_VERSION };
558+}
559+
560+function read(path: string, text: string, ctx: ExtractContext): FileFacts {
526561 const name = base(path);
527562 try {
528563 if (path.startsWith(".g1t/workflows/") || path.startsWith(".github/workflows/")) {
+135−0
1+import { test } from "node:test";
2+import assert from "node:assert/strict";
3+
4+import { assemble } from "./assemble.ts";
5+import { extract } from "./extract.ts";
6+import { aspirational, bulletsOf, classify, harvestable, planLike, sentencesOf, wholeSentences } from "./harvest.ts";
7+
8+const ctx = { project: "g1t", siblings: [] };
9+
10+test("memory comes from docs on how to work here, never plans, feedback or history", () => {
11+ for (const path of ["README.md", "AGENTS.md", "CLAUDE.md", "CONTRIBUTING.md"]) {
12+ const role = path.startsWith("README") ? "readme" : path.startsWith("CONTRIBUTING") ? "contributing" : "agents";
13+ assert.ok(harvestable(path, role), path);
14+ }
15+ for (const path of ["docs/DEPLOYING.md", "docs/testing.md", "docs/architecture.md", "docs/SELF_HOSTING.md", "runbooks/restore.md", "docs/getting-started.md"]) {
16+ assert.ok(harvestable(path, "doc"), path);
17+ }
18+ for (const path of ["docs/PLAN.md", "docs/CLOUDFLARE_FEEDBACK.md", "docs/CHANGELOG.md", "docs/roadmap.md", "docs/DEMO.md", "docs/INCIDENTS.md", "docs/PERFORMANCE.md", "docs/release-notes.md"]) {
19+ assert.ok(!harvestable(path, "doc"), path);
20+ }
21+ // project.yml adds a doc, or leaves one out, and leaving out wins.
22+ assert.ok(harvestable("docs/PERFORMANCE.md", "doc", { docs: ["docs/PERFORMANCE.md"], skip: [] }));
23+ assert.ok(harvestable("docs/notes/x.md", "doc", { docs: ["docs/notes/*"], skip: [] }));
24+ assert.ok(!harvestable("README.md", "readme", { docs: [], skip: ["README.md"] }));
25+ assert.ok(!harvestable("docs/testing.md", "doc", { docs: ["docs/testing.md"], skip: ["docs/*"] }));
26+});
27+
28+const FEEDBACK = `# Building a platform: a field report
29+
30+### A1. What is a billable operation? (blocking)
31+
32+- **What we tried.** Price from cost: pass storage through to each workspace with a modest
33+ uniform overhead, so a bill tracks what it really costs us.
34+- **What we hit.** Pricing defines operations only loosely.
35+- **What it means.** Our capacity model gives a 15x spread.
36+- **Ask.** A table: each binding method and each endpoint, billable or not.
37+
38+### Conventions we would like
39+
40+- **Ask.** Metrics with the same event names as the invoice, per repository.
41+`;
42+
43+const PLAN = `# Plan
44+
45+## For people who do not write code
46+
47+- **Roles.** A person can plan and approve work without ever cloning a
48+ repo. The roles that were sketched here map onto the five repository
49+ roles that are built.
50+
51+### Rules every automation obeys
52+
53+- **Deduplication.** The same Sentry issue firing 500 times maps to one
54+ issue.
55+- **Loop protection.** Work started by an automation cannot retrigger the
56+ same automation without a person in between.
57+
58+## Roadmap
59+
60+## Later
61+
62+- g1t will learn from docs.
63+`;
64+
65+test("a doc that reads as feedback or a plan says nothing, whatever its name", () => {
66+ assert.ok(planLike(FEEDBACK));
67+ assert.deepEqual(extract("CONTRIBUTING.md", FEEDBACK, ctx).hints, []);
68+ assert.ok(planLike(PLAN));
69+ assert.deepEqual(extract("docs/testing.md", PLAN, ctx).hints, []);
70+ const files = [
71+ { path: "docs/PLAN.md", facts: extract("docs/PLAN.md", "# Working on g1t\n\n## Conventions\n\n- Run the whole suite before you push, every time.\n", ctx) },
72+ { path: "docs/CLOUDFLARE_FEEDBACK.md", facts: extract("docs/CLOUDFLARE_FEEDBACK.md", FEEDBACK, ctx) },
73+ ];
74+ const project = { id: "p", workspace: "w", slug: "g1t", name: "g1t", description: null, private: false, repoId: "r", repo: { namespace: "w", name: "g1t" }, rootDir: "", defaultBranch: "main" };
75+ const around = { owners: [], dependsOn: [], deploy: null, integrations: [] };
76+ assert.deepEqual(assemble(project, files, around).hints, [], "docs/PLAN.md is not a source of memory");
77+ // Unless project.yml says it is.
78+ const withConfig = [...files, { path: ".g1t/project.yml", facts: extract(".g1t/project.yml", "memory:\n docs:\n - docs/PLAN.md\n", ctx) }];
79+ assert.deepEqual(assemble(project, withConfig, around).hints.map((hint) => hint.path), ["docs/PLAN.md"]);
80+});
81+
82+test("a bullet is taken whole, with the lines it wraps onto", () => {
83+ const bullets = bulletsOf(
84+ "## Conventions\n\n- **Loop protection.** Work started by an automation cannot retrigger the\n same automation without a person in between.\n- Short one.\n\nA paragraph.\n\n```\n- not a bullet\n```\n",
85+ );
86+ assert.deepEqual(bullets, [
87+ { text: "**Loop protection.** Work started by an automation cannot retrigger the same automation without a person in between.", heading: "Conventions" },
88+ { text: "Short one.", heading: "Conventions" },
89+ ]);
90+ // Bullets in a plan section, and its subsections, are left out.
91+ assert.deepEqual(bulletsOf("## Roadmap\n\n### Next\n\n- One.\n\n## Testing\n\n- Two.\n").map((b) => b.text), ["Two."]);
92+});
93+
94+test("never cut inside a sentence: whole sentences up to the limit, or nothing", () => {
95+ assert.deepEqual(sentencesOf("Run `npm test` first. It reads README.md, e.g. the setup. Done!"), [
96+ "Run `npm test` first.",
97+ "It reads README.md, e.g. the setup.",
98+ "Done!",
99+ ]);
100+ assert.equal(wholeSentences("Short and whole", 40), "Short and whole");
101+ assert.equal(wholeSentences("First sentence is here. Second sentence pushes it past the limit.", 30), "First sentence is here.");
102+ assert.equal(wholeSentences("One very long sentence that never ends before the limit is reached at all", 30), null);
103+});
104+
105+test("plans and wishes are not memory; kinds are earned, not assumed", () => {
106+ assert.ok(aspirational("g1t will learn from docs.", "doc"));
107+ assert.ok(aspirational("Tests should run on every push.", "contributing"));
108+ assert.ok(aspirational("**Ask.** A usage endpoint.", "doc"));
109+ assert.ok(aspirational("**What we tried.** Pass storage through.", "doc"));
110+ assert.ok(!aspirational("Run `cargo test` in the crate you changed.", "contributing"));
111+ // In AGENTS.md "should" states a rule.
112+ assert.ok(!aspirational("You should run the formatter before committing.", "agents"));
113+ assert.ok(aspirational("TODO: document the release flow.", "agents"));
114+
115+ assert.deepEqual(classify("Never edit generated files by hand.", "Conventions", "contributing"), { kind: "gotcha", confidence: 0.6 });
116+ assert.deepEqual(classify("Use the shared client for every call.", "Conventions", "contributing"), { kind: "convention", confidence: 0.6 });
117+ assert.deepEqual(classify("The API lives in apps/api and serves api.g1t.sh.", "Conventions", "readme"), { kind: "fact", confidence: 0.5 });
118+});
119+
120+test("CONTRIBUTING's before-you-push bullets are suggested whole; nothing is cut", () => {
121+ const facts = extract(
122+ "CONTRIBUTING.md",
123+ "# Contributing\n\n## Before you push\n\n- `cargo test` in the crate or service you changed.\n- `npx tsc -b --force` in `apps/web` (the incremental build misses changes\n in `packages/contracts`).\n- Look at what you changed in a browser. Screenshots catch what type\n checks do not.\n\n## Deploying\n\nPushes to `main` deploy themselves.\n",
124+ ctx,
125+ );
126+ assert.deepEqual(
127+ facts.hints.map((hint) => [hint.kind, hint.text, hint.confidence]),
128+ [
129+ ["fact", "`cargo test` in the crate or service you changed.", 0.5],
130+ ["fact", "`npx tsc -b --force` in `apps/web` (the incremental build misses changes in `packages/contracts`).", 0.5],
131+ ["convention", "Look at what you changed in a browser. Screenshots catch what type checks do not.", 0.6],
132+ ],
133+ );
134+ for (const hint of facts.hints) assert.match(hint.text, /[.!?)`]$/);
135+});
+255−0
1+/**
2+ * Which docs, and which of their lines, are worth suggesting as memory.
3+ *
4+ * Memory is for how to work in a project: how to build, test and run it,
5+ * its conventions and its traps. A project's docs also hold plans,
6+ * roadmaps, feedback to vendors and release notes, which say what someone
7+ * wants, not how things are; none of that is suggested. A line is taken
8+ * whole (a bullet with its continuation lines), cut only between sentences
9+ * to fit, and left out when even its first sentence does not fit. Pure.
10+ */
11+
12+/** README, AGENTS.md (or CLAUDE.md), CONTRIBUTING, or another doc. */
13+export type DocRole = "readme" | "agents" | "contributing" | "doc";
14+
15+/** What a project's `.g1t/project.yml` says about memory from its docs. */
16+export type MemoryDocsConfig = {
17+ /** Docs to read for memory besides the ones g1t picks, by path or `dir/*`. */
18+ docs: string[];
19+ /** Docs never read for memory, by path or `dir/*`. Wins over `docs`. */
20+ skip: string[];
21+};
22+
23+/** The longest memory suggested from a doc, in characters. */
24+export const MAX_DOC_HINT_CHARS = 300;
25+
26+/** Doc names that say how to work on a project. */
27+const WORKING_NAME =
28+ /(build|test|setup|set-up|develop|contribut|hacking|architecture|convention|style|guideline|standard|deploy|install|getting[-_ ]?started|local|onboard|troubleshoot|debug|runbook|self[-_ ]?host|coding|structure|layout|workflow|agents|claude)/i;
29+/** Doc names that hold plans, feedback, history or news. */
30+const PLAN_NAME =
31+ /(plan|roadmap|feedback|changelog|changes|history|news|release[-_ ]?notes|rfc|proposal|idea|todo|backlog|vision|strategy|research|demo|incident|postmortem|post-mortem|retro|meeting|minutes|notes|draft|wish|faq|announce|blog|pitch)/i;
32+
33+/** Headings of sections that hold plans, asks or history. */
34+const PLAN_HEADING =
35+ /\b(roadmap|plans?|planned|planning|milestones?|phase \d+|future|later|next steps|what'?s next|what we'?d love|wish ?list|asks?|feedback|changelog|release notes|what'?s new|open questions|non-goals|goals|ideas|proposals?|backlog|todo|executive summary|what we (hit|tried|built|want))\b/i;
36+/** Headings of sections that say how things are done here. */
37+const CONVENTION_HEADING =
38+ /\b(conventions?|guidelines?|code style|coding style|style guide|standards|gotchas?|pitfalls?|caveats?|known issues|troubleshooting|before you (push|commit|open)|house rules|how we work|working (here|on|in)|dos and don'?ts|rules for (contributors|agents|code))\b/i;
39+/** Headings of sections that say how to set up, build or test. */
40+export const SETUP_HEADING =
41+ /\b(develop|development|getting started|setup|set up|install|installing|build|building|test|testing|contribut|running|run locally|local|before you push|deploying)\b/i;
42+
43+/** A bullet's bold label that marks it as a plan, an ask or a report. */
44+const PLAN_LABEL =
45+ /^(ask|asks|what we (tried|hit|built|want|need|'d love)|what it means|proposal|idea|plan|next|later|future|todo|status|why it matters)\b/i;
46+/** Words that make a line a plan or a wish rather than how things are. */
47+const ASPIRATION =
48+ /\b(will|won'?t|would|should|could|might|shall|plans? to|planned|we'?d|we'?ll|eventually|someday|one day|in the future|not yet|coming soon|todo|tbd|wip|roadmap)\b|^ask\b/i;
49+/** The modal words an AGENTS.md uses to state a rule; aspirational elsewhere. */
50+const RULE_MODAL = /\b(will|won'?t|would|should|could|might|shall)\b/gi;
51+
52+/** Words that mark a line as a trap. */
53+const GOTCHA =
54+ /^(never|don'?t|do not|avoid|beware|careful|warning|watch out)\b|\b(must not|must never|gotcha|beware|careful|fails? (unless|if|when|until|without)|breaks? (if|when|unless)|otherwise)\b/i;
55+/** Words that mark a line as a rule of how things are done. */
56+const CONVENTION =
57+ /^(use|run|prefer|keep|put|name|write|import|add|look|always|never|don'?t|do not|avoid|make sure|remember|check|update|commit|open|follow|place|call|read|test|document)\b|\b(we use|must|prefer|instead of|rather than|always|never|convention|is required|are required)\b/i;
58+
59+/** Whether `path` matches one of `patterns`: an exact path, or `dir/*`. */
60+function matches(path: string, patterns: string[]): boolean {
61+ const lower = path.toLowerCase();
62+ return patterns.some((raw) => {
63+ const pattern = raw.trim().replace(/^\.?\//, "").toLowerCase();
64+ if (!pattern) return false;
65+ if (pattern.endsWith("/*")) return lower.startsWith(pattern.slice(0, -1)) && !lower.slice(pattern.length - 1).includes("/");
66+ if (pattern.endsWith("/")) return lower.startsWith(pattern);
67+ return lower === pattern;
68+ });
69+}
70+
71+/**
72+ * Whether memory is suggested from the doc at `path`: a README, AGENTS.md,
73+ * CLAUDE.md or CONTRIBUTING always, a runbook, and another doc when its
74+ * name says how to work on the project (`docs/DEPLOYING.md`, not
75+ * `docs/PLAN.md` or `docs/CHANGELOG.md`). `.g1t/project.yml` can add docs
76+ * or leave them out.
77+ */
78+export function harvestable(path: string, role: DocRole, config?: MemoryDocsConfig | null): boolean {
79+ if (config && matches(path, config.skip)) return false;
80+ if (config && matches(path, config.docs)) return true;
81+ if (role !== "doc") return true;
82+ const name = path.slice(path.lastIndexOf("/") + 1).replace(/\.(md|markdown|txt)$/i, "");
83+ if (PLAN_NAME.test(name)) return false;
84+ if (/(^|\/)runbooks\//i.test(path)) return true;
85+ return WORKING_NAME.test(name);
86+}
87+
88+/**
89+ * Whether a doc reads as a plan, a report or feedback rather than how to
90+ * work: many of its headings are plans or asks, or many of its bullets are
91+ * labelled as asks and findings.
92+ */
93+export function planLike(text: string): boolean {
94+ let headings = 0;
95+ let planHeadings = 0;
96+ let labels = 0;
97+ let aspirations = 0;
98+ let lines = 0;
99+ for (const line of text.split(/\r?\n/)) {
100+ const heading = /^#{1,6}\s+(.*)$/.exec(line)?.[1];
101+ if (heading) {
102+ headings++;
103+ if (PLAN_HEADING.test(heading)) planHeadings++;
104+ continue;
105+ }
106+ const label = /^\s*[-*+]\s+\*\*([^*]+)\*\*/.exec(line)?.[1];
107+ if (label && PLAN_LABEL.test(label.trim())) labels++;
108+ if (line.trim()) {
109+ lines++;
110+ if (/\b(we will|we'?ll|we plan|we'?d love|will be|is planned|are planned|not yet built|coming soon)\b/i.test(line)) aspirations++;
111+ }
112+ }
113+ return labels >= 3 || (headings >= 3 && planHeadings * 3 >= headings) || (lines >= 20 && aspirations * 8 >= lines);
114+}
115+
116+/** Markdown inline marks taken out: bold, italics, links to their text, images. */
117+export function plain(text: string): string {
118+ return text
119+ .replace(/!\[[^\]]*\]\([^)]*\)/g, "")
120+ .replace(/\[([^\]]*)\]\([^)]*\)/g, "$1")
121+ .replace(/\*\*([^*]+)\*\*/g, "$1")
122+ .replace(/__([^_]+)__/g, "$1")
123+ .replace(/(^|\s)\*([^*\s][^*]*)\*(?=\s|[.,;:!?)]|$)/g, "$1$2")
124+ .replace(/\s+/g, " ")
125+ .trim();
126+}
127+
128+const ABBREVIATION = /(\b(e\.g|i\.e|etc|vs|cf|approx|no|fig)\.|\b[A-Z]\.)$/i;
129+
130+/** A line's sentences, never split inside `code` or after an abbreviation. */
131+export function sentencesOf(text: string): string[] {
132+ const out: string[] = [];
133+ let current = "";
134+ let code = false;
135+ for (let i = 0; i < text.length; i++) {
136+ const c = text[i];
137+ current += c;
138+ if (c === "`") code = !code;
139+ if (code || !/[.!?]/.test(c)) continue;
140+ const next = text[i + 1];
141+ // A full stop inside a word (README.md, v1.2) or before more punctuation.
142+ if (next !== undefined && !/\s/.test(next)) continue;
143+ if (ABBREVIATION.test(current.trimEnd())) continue;
144+ out.push(current.trim());
145+ current = "";
146+ }
147+ if (current.trim()) out.push(current.trim());
148+ return out;
149+}
150+
151+/**
152+ * `text` whole if it fits in `max` characters; otherwise as many of its
153+ * first sentences as fit; null when not even the first one does. Never cut
154+ * inside a sentence.
155+ */
156+export function wholeSentences(text: string, max = MAX_DOC_HINT_CHARS): string | null {
157+ const line = text.replace(/\s+/g, " ").trim();
158+ if (line.length <= max) return line;
159+ let out = "";
160+ for (const sentence of sentencesOf(line)) {
161+ const next = out ? `${out} ${sentence}` : sentence;
162+ if (next.length > max) break;
163+ out = next;
164+ }
165+ return out && /[.!?]$/.test(out) ? out : null;
166+}
167+
168+/** Whether a line says what someone wants or plans rather than how things are. */
169+export function aspirational(text: string, role: DocRole): boolean {
170+ const label = /^\*\*([^*]+)\*\*/.exec(text.trim())?.[1]?.trim();
171+ if (label && PLAN_LABEL.test(label)) return true;
172+ const words = plain(text);
173+ if (/^(ask|what we (tried|hit|built))\b/i.test(words)) return true;
174+ // AGENTS.md states its rules with "should" and "will": those are rules.
175+ const checked = role === "agents" ? words.replace(RULE_MODAL, "") : words;
176+ return ASPIRATION.test(checked);
177+}
178+
179+/** The kind of memory a line from a doc is, and how sure g1t is of it. */
180+export function classify(text: string, heading: string, role: DocRole): { kind: "fact" | "convention" | "gotcha"; confidence: number } {
181+ const agents = role === "agents";
182+ if (GOTCHA.test(text)) return { kind: "gotcha", confidence: agents ? 0.9 : 0.6 };
183+ if (CONVENTION.test(text) || (agents && CONVENTION_HEADING.test(heading))) return { kind: "convention", confidence: agents ? 0.9 : 0.6 };
184+ return { kind: "fact", confidence: agents ? 0.85 : 0.5 };
185+}
186+
187+/** Headings of sections that report or explain rather than say how. */
188+const REPORT_HEADING = /\b(speed|performance|benchmarks?|timings?|costs?|numbers|results|findings|lessons|why|background|motivation|history|incidents?|features|status)\b/i;
189+
190+/** Whether a section's bullets are worth reading for memory. */
191+export function bulletSection(heading: string, role: DocRole): boolean {
192+ if (PLAN_HEADING.test(heading) || (role !== "agents" && REPORT_HEADING.test(heading))) return false;
193+ return role === "agents" || CONVENTION_HEADING.test(heading) || SETUP_HEADING.test(heading);
194+}
195+
196+/** Whether a section is a plan, an ask or history, with everything under it. */
197+export function planSection(heading: string): boolean {
198+ return PLAN_HEADING.test(heading);
199+}
200+
201+export type Bullet = { text: string; heading: string };
202+
203+/**
204+ * A doc's bullets, each whole with its continuation lines, with the heading
205+ * it is under. Bullets in code blocks and in plan sections (and their
206+ * subsections) are left out.
207+ */
208+export function bulletsOf(markdown: string): Bullet[] {
209+ const out: Bullet[] = [];
210+ let heading = "";
211+ // The level of the plan section being skipped, if any.
212+ let skipping = 0;
213+ let fenced = false;
214+ let current: { text: string; indent: number } | null = null;
215+ const flush = () => {
216+ if (current && !skipping) out.push({ text: current.text.replace(/\s+/g, " ").trim(), heading });
217+ current = null;
218+ };
219+ for (const raw of markdown.split(/\r?\n/)) {
220+ if (/^\s*(```|~~~)/.test(raw)) {
221+ flush();
222+ fenced = !fenced;
223+ continue;
224+ }
225+ if (fenced) continue;
226+ const h = /^(#{1,6})\s+(.*)$/.exec(raw);
227+ if (h) {
228+ flush();
229+ const level = h[1].length;
230+ if (skipping && level <= skipping) skipping = 0;
231+ heading = h[2].trim();
232+ if (!skipping && planSection(heading)) skipping = level;
233+ continue;
234+ }
235+ const bullet = /^(\s*)(?:[-*+]|\d+[.)])\s+(.*)$/.exec(raw);
236+ if (bullet) {
237+ flush();
238+ current = { text: bullet[2], indent: bullet[1].length };
239+ continue;
240+ }
241+ if (current && raw.trim() && /^\s+/.test(raw) && raw.search(/\S/) > current.indent) {
242+ current.text += ` ${raw.trim()}`;
243+ continue;
244+ }
245+ // A blank line or a paragraph ends the bullet. A bullet's lazy
246+ // continuation (not indented) is still part of it.
247+ if (current && raw.trim() && !/^\s*[|<>]/.test(raw)) {
248+ current.text += ` ${raw.trim()}`;
249+ continue;
250+ }
251+ flush();
252+ }
253+ flush();
254+ return out;
255+}
+29−7
6868 } from "@g1t/contracts";
6969
7070 import { assemble, authorsOf, integrationEntities, type EntityDraft, type FileRecord, type ProjectInput, type Surroundings } from "./assemble";
71−import { extract, interesting, type FileFacts } from "./extract";
71+import { EXTRACT_VERSION, extract, interesting, type FileFacts } from "./extract";
7272 import { composeRunContext, type ContextNote, type ProjectContext } from "./runcontext";
7373 import { evaluate } from "./scorecards";
7474 import { allowedKinds, countVisible, indexFilter, memoryReadable, merge, projectReadable, readable, runMemoryReadable, type IndexMeta, type Reader } from "./visibility";
271271 }
272272 }
273273
274−type ScanStats = { entities: number; candidates: number; kept: number; indexed: number };
274+type ScanStats = { entities: number; candidates: number; kept: number; indexed: number; pruned?: number };
275275
276276 class Context {
277277 constructor(private readonly env: Env) {}
454454 // ---- Building the catalog --------------------------------------------------
455455
456456 /** The files of a project worth reading, with their blobs, found in a few tree reads. */
457− private async candidates(project: Project, actor: User, ref: string): Promise<{ files: { path: string; hash: string }[]; siblings: string[]; head: string | null } | null> {
457+ private async candidates(project: Project, actor: User, ref: string): Promise<{ files: { path: string; hash: string }[]; siblings: string[]; head: string | null; partial: boolean } | null> {
458458 if (project.source.kind !== "hosted") return null;
459459 const repos = reposClient(this.env.REPOS);
460460 const { repo, rootDir: root } = project.source;
472472 };
473473 blobs(top.value.entries, "");
474474 const dirs = new Set(top.value.entries.filter((entry) => entry.kind === "tree").map((entry) => entry.name));
475+ // A folder that could not be listed leaves the list partial.
476+ let partial = false;
475477 const look = async (dir: string) => {
476478 const found = await repos.tree(repo, actor, ref, at(dir)).catch(() => null);
479+ if (!found?.ok) partial = true;
477480 return found?.ok ? found.value.entries : [];
478481 };
479482 for (const dir of ["docs", "doc", "runbooks"]) if (dirs.has(dir)) blobs(await look(dir), dir);
483486 blobs(inside, dir);
484487 if (inside.some((entry) => entry.name === "workflows" && entry.kind === "tree")) blobs(await look(`${dir}/workflows`), `${dir}/workflows`);
485488 }
486− return { files, siblings, head: top.value.head?.hash ?? null };
489+ return { files, siblings, head: top.value.head?.hash ?? null, partial };
487490 }
488491
489492 /**
511514 const writes: D1PreparedStatement[] = [];
512515 let reads = 0;
513516 let workflowReads = 0;
517+ // Whether every file is known as this version of extract reads it: only
518+ // then are the doc candidates it no longer suggests let go.
519+ let complete = !found.partial;
514520 const repos = reposClient(this.env.REPOS);
515521 for (const file of found.files) {
516522 const before = stored.get(file.path);
523+ const kept = before ? (JSON.parse(before.facts) as FileFacts) : null;
517524 const workflow = file.path.includes("/workflows/");
518− const fresh = before && before.hash === file.hash && !force;
525+ // Facts an older extract read are read again, as a changed file is.
526+ const fresh = before && before.hash === file.hash && kept?.version === EXTRACT_VERSION && !force;
519527 const canRead = reads < MAX_READS && (!workflow || workflowReads < MAX_WORKFLOW_READS);
520528 if (fresh || !canRead) {
521− if (before) files.push({ path: file.path, facts: JSON.parse(before.facts) as FileFacts });
529+ if (!fresh) complete = false;
530+ if (kept) files.push({ path: file.path, facts: kept });
522531 continue;
523532 }
524533 reads++;
525534 if (workflow) workflowReads++;
526535 const blob = await repos.blob(repo, cache.actor, found.head ?? ref, [rootDir, file.path].filter(Boolean).join("/")).catch(() => null);
527− if (!blob?.ok || blob.value.text == null) continue;
536+ if (!blob?.ok || blob.value.text == null) {
537+ complete = false;
538+ continue;
539+ }
528540 const facts = extract(file.path, blob.value.text, { project: project.name, siblings: found.siblings });
529541 files.push({ path: file.path, facts });
530542 changed.push(file.path);
724736 stats.candidates += captured?.added ?? 0;
725737 stats.kept += captured?.kept ?? 0;
726738 }
739+ // Candidates from this project's docs that are still waiting but that
740+ // its docs, as read now, no longer suggest: a line since changed, or one
741+ // the rules for what is worth remembering leave out. Kept and dismissed
742+ // memory is never touched.
743+ if (complete) {
744+ const pruned = await memoryReviewClient(this.env.WORK)
745+ .pruneDocCandidates(workspace, repoId, built.hints.map((hint) => hint.text))
746+ .catch(() => null);
747+ stats.pruned = pruned?.removed ?? 0;
748+ }
727749 return stats;
728750 }
729751
+5−0
11 {
22 "extends": "../../tsconfig.base.json",
3+ "compilerOptions": {
4+ // Node runs the tests on these files as they are, which needs the
5+ // extension on relative imports.
6+ "allowImportingTsExtensions": true
7+ },
38 "include": [
49 "src/**/*",
510 "worker-configuration.d.ts"
+188−10
1111 //! decision candidate. No model is asked: the pull request's own words
1212 //! are the decision, and a candidate waits for review anyway.
1313 //! - **Docs.** The context service reads a project's README, AGENTS.md,
14−//! docs and manifests and sends what they say (`capture_memories`).
14+//! CONTRIBUTING, docs on how to work in it, and manifests, and sends
15+//! what they say (`capture_memories`). Once it has read every file, the
16+//! waiting candidates its docs no longer suggest are let go
17+//! (`prune_doc_candidates`).
1518 //!
1619 //! The same thing said again, by an independent source, is one memory seen
17−//! twice, which keeps it (`promotes`). A dismissed memory stays dismissed,
18−//! so the same wording is never suggested again. Nothing that looks like a
19−//! secret is stored, in the text or in its evidence.
20+//! twice, which keeps it (`promotes`). The same thing in other words (a
21+//! sentence that grew, one cut short) is the memory already there
22+//! (`same_memory`). A dismissed memory stays dismissed, so the same thing
23+//! is never suggested again. Nothing that looks like a secret is stored, in
24+//! the text or in its evidence.
2025
2126 use std::collections::HashMap;
2227
7681 "the right way",
7782 ];
7883
79−#[derive(Deserialize)]
84+#[derive(Clone, Deserialize)]
8085 struct Existing {
8186 id: String,
8287 status: String,
8388 sources: String,
8489 confidence: Option<f64>,
90+ text: String,
91+}
92+
93+/// The most memories of one scope compared with a new one for being the
94+/// same thing in other words: every kept and waiting one (at most 500) and
95+/// the most recently dismissed.
96+const MAX_COMPARED: u32 = 2000;
97+
98+/// The memory in `among` that says the same thing as `text`, if any.
99+fn same_as<'a>(among: &'a [Existing], text: &str) -> Option<&'a Existing> {
100+ among.iter().find(|existing| same_memory(&existing.text, text))
85101 }
86102
103+/// Whether `existing` is a waiting candidate that the source `reference`
104+/// said before and now says in other words (`print`).
105+fn reworded_by(existing: &Existing, reference: &str, print: &str) -> bool {
106+ let sources: Vec<String> = serde_json::from_str(&existing.sources).unwrap_or_default();
107+ existing.status == MemoryStatus::Candidate.as_str()
108+ && fingerprint(&existing.text) != print
109+ && sources.iter().any(|source| source == reference)
110+}
111+
112+/// Whether a candidate came from docs and manifests alone.
113+fn only_from_docs(sources: &str) -> bool {
114+ let list: Vec<String> = serde_json::from_str(sources).unwrap_or_default();
115+ !list.is_empty() && list.iter().all(|source| source.starts_with("doc:"))
116+}
117+
118+/// Of a project's waiting doc candidates, the ids its docs no longer
119+/// suggest: none of `texts` is the same memory.
120+fn unsuggested(candidates: &[Existing], texts: &[String]) -> Vec<String> {
121+ candidates
122+ .iter()
123+ .filter(|candidate| only_from_docs(&candidate.sources))
124+ .filter(|candidate| !texts.iter().any(|text| same_memory(&candidate.text, text)))
125+ .map(|candidate| candidate.id.clone())
126+ .collect()
127+}
128+
87129 #[derive(Deserialize)]
88130 struct RunTicket {
89131 id: String,
281323 let workspace = workspace.to_lowercase();
282324 let mut done = Captured::default();
283325 let mut paths = HashMap::new();
326+ // Each scope's memories, read once a capture, for near-duplicates.
327+ let mut scopes: HashMap<(&'static str, String), Vec<Existing>> = HashMap::new();
284328 for item in items.iter().take(MAX_CAPTURE) {
285329 let Ok((text, evidence)) = checked(item) else {
286330 done.refused += 1;
307351 }
308352 let print = fingerprint(&text);
309353 let now = rfc3339(now_ms());
310− let existing = self
354+ let mut existing = self
311355 .db
312356 .prepare(
313− "SELECT id, status, sources, confidence FROM memories
357+ "SELECT id, status, sources, confidence, text FROM memories
314358 WHERE scope = ? AND scope_key = ? AND (fingerprint = ? OR text = ?) LIMIT 1",
315359 )
316360 .bind(&[item.scope.as_str().into(), key.as_str().into(), print.as_str().into(), text.as_str().into()])?
317361 .first::<Existing>(None)
318362 .await?;
363+ // The same thing in other words: a sentence that grew, or one
364+ // cut short, is the memory already there, kept, waiting or
365+ // dismissed.
366+ let scope_key = (item.scope.as_str(), key.clone());
367+ if existing.is_none() {
368+ if !scopes.contains_key(&scope_key) {
369+ let rows = self
370+ .db
371+ .prepare(
372+ "SELECT id, status, sources, confidence, text FROM memories
373+ WHERE scope = ? AND scope_key = ?
374+ ORDER BY status = 'dismissed', updated_at DESC LIMIT ?",
375+ )
376+ .bind(&[item.scope.as_str().into(), key.as_str().into(), MAX_COMPARED.into()])?
377+ .all()
378+ .await?
379+ .results::<Existing>()?;
380+ scopes.insert(scope_key.clone(), rows);
381+ }
382+ existing = scopes.get(&scope_key).and_then(|among| same_as(among, &text)).cloned();
383+ }
319384 if let Some(existing) = existing {
320385 done.merged += 1;
321386 if existing.status == MemoryStatus::Dismissed.as_str() {
322387 continue;
323388 }
324389 let sources = with_source(&existing.sources, &item.reference);
390+ // A waiting candidate its own source now says in other words
391+ // (a doc's line read whole, or edited) takes the new words,
392+ // kind and confidence. Kept memory keeps its words.
393+ let reworded = reworded_by(&existing, &item.reference, &print);
325394 let confidence = match (existing.confidence, item.confidence) {
395+ _ if reworded => item.confidence,
326396 (Some(a), Some(b)) => Some(a.max(b)),
327397 (a, b) => a.or(b),
328398 };
332402 done.kept += 1;
333403 }
334404 let status = if keep { MemoryStatus::Kept.as_str() } else { existing.status.as_str() };
405+ let (new_text, new_print) =
406+ if reworded { (text.clone(), print.clone()) } else { (existing.text.clone(), fingerprint(&existing.text)) };
335407 self.db
336408 .prepare(
337− "UPDATE memories SET sources = ?, confidence = ?, status = ?, fingerprint = ?, updated_at = ?
409+ "UPDATE memories SET sources = ?, confidence = ?, status = ?, text = ?, fingerprint = ?,
410+ kind = COALESCE(?, kind), evidence = COALESCE(?, evidence), updated_at = ?
338411 WHERE id = ?",
339412 )
340413 .bind(&[
341414 serde_json::to_string(&sources)?.into(),
342415 confidence.map_or(JsValue::NULL, JsValue::from),
343416 status.into(),
344− print.as_str().into(),
417+ new_text.as_str().into(),
418+ new_print.as_str().into(),
419+ if reworded { item.kind.as_str().into() } else { JsValue::NULL },
420+ if reworded { evidence.clone().map_or(JsValue::NULL, JsValue::from) } else { JsValue::NULL },
345421 now.as_str().into(),
346422 existing.id.as_str().into(),
347423 ])?
403479 if keep {
404480 done.kept += 1;
405481 }
482+ // The same thing twice in one capture is one memory.
483+ if let Some(among) = scopes.get_mut(&scope_key) {
484+ among.push(Existing {
485+ id: id.clone(),
486+ status: status.as_str().to_owned(),
487+ sources: serde_json::to_string(&[&item.reference])?,
488+ confidence: item.confidence,
489+ text: text.clone(),
490+ });
491+ }
406492 self.memory_changed(&id, &workspace, status.as_str(), repo.as_deref()).await;
407493 }
408494 Ok(done)
409495 }
410496
497+ /// Removes a project's waiting candidates that came from its docs alone
498+ /// and that its docs, as read now, no longer suggest. Kept and dismissed
499+ /// memory is never touched, nor a candidate another source saw too.
500+ pub(crate) async fn prune_doc_candidates(&self, a: PruneDocCandidatesArgs) -> Result<Pruned> {
501+ let workspace = a.workspace.to_lowercase();
502+ let candidates = self
503+ .db
504+ .prepare(
505+ "SELECT id, status, sources, confidence, text FROM memories
506+ WHERE workspace = ? AND scope = 'project' AND scope_key = ? AND status = 'candidate'
507+ AND source_kind = 'doc'
508+ LIMIT ?",
509+ )
510+ .bind(&[workspace.as_str().into(), a.repo_id.as_str().into(), MAX_COMPARED.into()])?
511+ .all()
512+ .await?
513+ .results::<Existing>()?;
514+ let gone = unsuggested(&candidates, &a.texts);
515+ if gone.is_empty() {
516+ return Ok(Pruned::default());
517+ }
518+ let removed = self
519+ .db
520+ .prepare(
521+ "DELETE FROM memories WHERE workspace = ? AND scope_key = ? AND status = 'candidate'
522+ AND id IN (SELECT value FROM json_each(?)) RETURNING id",
523+ )
524+ .bind(&[workspace.as_str().into(), a.repo_id.as_str().into(), serde_json::to_string(&gone)?.into()])?
525+ .all()
526+ .await?
527+ .results::<serde_json::Value>()?
528+ .len();
529+ worker::console_log!("pruned {removed} doc candidates of {} in {workspace}", a.repo_id);
530+ Ok(Pruned { removed: removed as u32 })
531+ }
532+
411533 pub(crate) async fn capture_memories(&self, a: CaptureMemoriesArgs) -> Result<Captured> {
412534 self.capture(&a.workspace, &a.items, a.by.as_deref().unwrap_or("g1t")).await
413535 }
744866 }
745867
746868 /// The methods this module serves at `POST /rpc/<method>`.
747−pub(crate) const METHODS: [&str; 7] = [
869+pub(crate) const METHODS: [&str; 8] = [
748870 "capture_memories",
871+ "prune_doc_candidates",
749872 "report_learned",
750873 "list_candidates",
751874 "review_memory",
758881 use g1t_kit::{args, reply};
759882 match method {
760883 "capture_memories" => reply(&work.capture_memories(args(body)?).await?),
884+ "prune_doc_candidates" => reply(&work.prune_doc_candidates(args(body)?).await?),
761885 "report_learned" => reply(&work.report_learned(args(body)?).await?),
762886 "list_candidates" => reply(&work.list_candidates(args(body)?).await?),
763887 "review_memory" => reply(&work.review_memory(args(body)?).await?),
834958 assert_eq!(with_source("not json", "doc:x"), ["doc:x"]);
835959 }
836960
961+ fn existing(id: &str, text: &str, sources: &str) -> Existing {
962+ Existing { id: id.into(), status: "candidate".into(), sources: sources.into(), confidence: Some(0.7), text: text.into() }
963+ }
964+
965+ #[test]
966+ fn the_same_thing_in_other_words_is_found() {
967+ let among = [
968+ existing("a", "Run cargo test in the crate you changed.", "[]"),
969+ existing("b", "g1t is a Cargo workspace (apps/api, crates/*, services/actions); `cargo test` runs its tests.", "[]"),
970+ ];
971+ assert_eq!(same_as(&among, "run cargo test in the crate you changed").map(|e| e.id.as_str()), Some("a"));
972+ assert_eq!(
973+ same_as(&among, "g1t is a Cargo workspace (apps/*, crates/*, services/*); `cargo test` runs its tests.").map(|e| e.id.as_str()),
974+ Some("b")
975+ );
976+ assert!(same_as(&among, "Use pnpm, never npm.").is_none());
977+ }
978+
979+ #[test]
980+ fn a_candidate_cut_short_takes_its_doc_s_whole_words() {
981+ let cut = existing("a", "Work started by an automation cannot retrigger the", "[\"doc:r:CONTRIBUTING.md\"]");
982+ let whole = fingerprint("Work started by an automation cannot retrigger the same automation.");
983+ assert!(reworded_by(&cut, "doc:r:CONTRIBUTING.md", &whole));
984+ // Another source saying it is a sighting, not new words.
985+ assert!(!reworded_by(&cut, "run:x", &whole));
986+ // Kept memory keeps the words someone kept.
987+ let kept = Existing { status: "kept".into(), ..cut.clone() };
988+ assert!(!reworded_by(&kept, "doc:r:CONTRIBUTING.md", &whole));
989+ // The same words are not new ones.
990+ assert!(!reworded_by(&cut, "doc:r:CONTRIBUTING.md", &fingerprint(&cut.text)));
991+ }
992+
993+ #[test]
994+ fn waiting_doc_candidates_the_docs_no_longer_suggest_are_let_go() {
995+ let candidates = [
996+ // Cut mid-sentence by the old reader; the docs now say it whole.
997+ existing("cut", "Run the whole suite before you push, every", "[\"doc:r:CONTRIBUTING.md\"]"),
998+ // From a plan, which is no longer read for memory.
999+ existing("plan", "Loop protection. Work started by an automation cannot retrigger the", "[\"doc:r:docs/PLAN.md\"]"),
1000+ // Still suggested.
1001+ existing("still", "`cargo test` in the crate or service you changed.", "[\"doc:r:CONTRIBUTING.md\"]"),
1002+ // Seen by a run too: not the docs' alone to take back.
1003+ existing("run", "Deduplication. The same Sentry issue firing 500 times maps to one", "[\"doc:r:docs/PLAN.md\",\"run:x\"]"),
1004+ ];
1005+ let texts = [
1006+ "Run the whole suite before you push, every time, on every branch.".to_owned(),
1007+ "`cargo test` in the crate or service you changed.".to_owned(),
1008+ ];
1009+ assert_eq!(unsuggested(&candidates, &texts), ["plan"]);
1010+ assert!(only_from_docs("[\"doc:a\",\"doc:b\"]"));
1011+ assert!(!only_from_docs("[]"));
1012+ assert!(!only_from_docs("[\"doc:a\",\"review:b\"]"));
1013+ }
1014+
8371015 fn item(text: &str, evidence: Option<&str>) -> CaptureItem {
8381016 CaptureItem {
8391017 scope: MemoryScope::Project,