pr_01m47d15m3e54sn21z27rpy5n9/docs/PLAN.md
Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.
| Initial g1t: services, event bus, intents and attempts | 1 | # g1t plan |
| 2 | ||
| 3 | g1t is a git forge for agents, built on Cloudflare Workers and Artifacts for | |
| 4 | the "Build the Next-Gen Git Platform on Cloudflare" competition. | |
| 5 | ||
| 6 | - Submission closes **October 14, 2026, 11:59 PM PDT**: a 5–10 minute demo | |
| 7 | video, this repository (MIT) and run instructions. | |
| 8 | - Judging: 50% originality and quality of the prototype for agent-oriented | |
| 9 | collaboration; 25% multi-agent concurrency, coordination, context | |
| 10 | preservation, review and conflict handling; 25% ease of use. | |
| 11 | ||
| 12 | ## Product model | |
| 13 | ||
| 14 | A pull request assumes one author and one change. g1t assumes many agents | |
| 15 | working at once, in two shapes: several agents racing on the same goal, and | |
| 16 | many different goals in flight that all have to land on `main`. | |
| 17 | ||
| 18 | | Concept | What it is | | |
| 19 | | --- | --- | | |
| 20 | | **Intent** | A goal stated against a repo, with acceptance checks (commands that must pass). Replaces the issue and the pull request. | | |
| 21 | | **Attempt** | One agent's run at an intent, in its own Artifacts fork. Any number run in parallel. | | |
| 22 | | **Session** | The agent's full context for an attempt: prompt, messages, tool calls, cost. Stored with the attempt and linked from every commit it produced. | | |
| 23 | | **Arena** | The compare view for an intent: every attempt side by side with diff, check results, conflicts against main and against each other, and a reviewer agent's summary. | | |
| 24 | | **Ship** | A person or a policy picks an attempt. A per-repo merge queue lands it; the other attempts are rebased by their agents or closed. | | |
| 25 | ||
| 26 | Features that fall out of the model: | |
| 27 | ||
| 28 | - **Why-blame.** Click a line and see the prompt and reasoning that produced | |
| 29 | it, not only the commit. | |
| 30 | - **Overlap radar.** Attempts that touch the same files are flagged while the | |
| 31 | agents are still working, and the agents are told. | |
| 32 | - **Live lanes.** Watch every attempt progress in real time. | |
| 33 | ||
| 34 | ## Converging on main | |
| 35 | ||
| 36 | Twelve intents started together will finish at different times and touch | |
| 37 | overlapping code. Getting them all into `main` without a person refereeing | |
| 38 | is the hard part, and it is handled in four places. | |
| 39 | ||
| 40 | 1. **Before work starts: plan the overlap away.** A project is a graph of | |
| 41 | intents. A planner agent can split a large goal into intents, predict | |
| 42 | which files each will touch, and add a dependency where two would collide, | |
| 43 | so one starts from the other's result instead of from `main`. | |
| 44 | 2. **While agents work: overlap radar.** Each attempt's changed files and | |
| 45 | symbols are tracked as it pushes. When two attempts from different intents | |
| 46 | enter the same area, both agents are told what the other is doing there. | |
| 47 | 3. **When `main` moves: the author resolves.** Every open attempt is | |
| 48 | trial-merged against the new `main`. A clean merge updates the attempt | |
| 49 | silently. A conflict resumes that attempt's agent with its original | |
| 50 | session and the incoming change, so the conflict is resolved by the agent | |
| 51 | that wrote the code and still knows why. | |
| 52 | 4. **At landing: a speculative queue.** Shipped attempts enter the repo's | |
| 53 | queue. g1t builds the combined states (`main`+A, `main`+A+B, …) and runs | |
| 54 | their checks in parallel. Attempts land in order as their combined state | |
| 55 | passes; one that fails is ejected back to its agent and the states behind | |
| 56 | it are rebuilt. `main` only ever receives a state that passed. | |
| 57 | ||
| 58 | Landing can be fully automatic: a repo policy such as "checks pass and the | |
| 59 | reviewer agent approves" ships without a person. | |
| 60 | ||
| 61 | ## Agents aware of each other | |
| 62 | ||
| 63 | Each repo keeps a live **work registry**: for every running attempt, its | |
| 64 | intent, a running summary of what it has done, and the files and symbols it | |
| 65 | has touched or plans to touch. Agents use it through MCP tools; g1t also | |
| 66 | acts on it without being asked. | |
| 67 | ||
| 68 | - **Before starting.** When an intent is opened, or an agent is about to | |
| 69 | begin a task, g1t searches open intents and running attempts for the same | |
| 70 | goal (by meaning, not wording) and for the same area of code. If a match | |
| 71 | exists the agent is told who is on it and how far along, and chooses: join | |
| 72 | as a deliberate racer, wait for the result, or drop the task. Duplicate | |
| 73 | intents are offered for merging. | |
| 74 | - **Finding out-of-scope work.** An agent that discovers something outside | |
| 75 | its intent asks the registry who works there. If another attempt owns that | |
| 76 | area, it **hands off**: a note, the relevant excerpt of its session, and | |
| 77 | optionally commits the receiver can take. If nobody does, it opens a child | |
| 78 | intent instead of widening its own change. | |
| 79 | - **Asking.** An agent can put a question or a request to another attempt. | |
| 80 | The receiver gets it at its next turn. | |
| 81 | - **Waiting.** An agent that needs another attempt's result parks itself. | |
| 82 | Its sandbox sleeps, spend stops, and it resumes from the new state when | |
| 83 | that attempt ships. | |
| 84 | - **Agents that do not cooperate.** For pushes from tools that never call | |
| 85 | these tools, g1t compares the pushed change against running attempts and | |
| 86 | flags near-duplicates itself. | |
| 87 | ||
| 88 | Every handoff, question and wait has a state (offered, accepted, declined, | |
| 89 | done), appears in the timeline, and is visible to people. A handoff declined | |
| 90 | twice, or two agents passing work back and forth, goes to the "needs you" | |
| 91 | inbox. | |
| 92 | ||
| 93 | ## Review at scale | |
| 94 | ||
| 95 | Cloudflare's brief asks "how do you review everything they produce?". With | |
| 96 | hundreds of agents, a person cannot read every diff, so review is by | |
| 97 | exception. | |
| 98 | ||
| 99 | - **Evidence, not diffs.** Every attempt carries a proof bundle: checks run | |
| 100 | and their output, a preview URL, a plain-language summary, and the | |
| 101 | behaviour that changed. | |
| 102 | - **Two agent reviewers.** One reviews the change against the intent. A | |
| 103 | second is adversarial: it tries to break the change and reports what it | |
| 104 | found. | |
| 105 | - **Risk tiers.** Each change is scored from what it touches, how large it | |
| 106 | is, and how the reviewers ruled. Low risk ships on policy; high risk goes | |
| 107 | to a person with the evidence already assembled. | |
| 108 | - **Trust is earned.** An agent's record on a path (shipped, reverted, caught | |
| 109 | by review) raises or lowers the tier its changes land in. | |
| 110 | - **Sampling.** A share of auto-shipped changes is sent to a person anyway, | |
| 111 | to keep the policy honest. | |
| 112 | ||
| 113 | ## Rethinking the git primitives | |
| 114 | ||
| 115 | - **No branches for agents.** An attempt is a fork; `main` is the only | |
| 116 | long-lived line. There is nothing to name, clean up or go stale. | |
| 117 | - **Projected main.** New attempts start from `main` plus everything already | |
| 118 | in the landing queue, so they are built on the state they will land on. | |
| 119 | - **Structural merge.** The merge engine merges by syntax tree, not by line, | |
| 120 | for supported languages. Two agents adding different functions to the same | |
| 121 | file do not conflict. | |
| 122 | - **Forkable sessions.** A session can be forked at any turn: the code as it | |
| 123 | was at that moment plus the conversation up to it, continued with a | |
| 124 | different instruction. Branching applies to the reasoning as well as the | |
| 125 | code. | |
| 126 | - **Provenance in history.** Every commit records its intent, session, | |
| 127 | agent, model and cost, and is signed with a key issued to that attempt. The | |
| 128 | history can be audited by machine. | |
| 129 | ||
| 130 | ## People in the loop | |
| 131 | ||
| 132 | ### Code that arrives from outside | |
| 133 | ||
| 134 | People will keep pushing with plain git, their editor, or another tool. Every | |
| 135 | push goes through g1t's git front end, so none of it bypasses the model. | |
| 136 | ||
| 137 | - **A push to a branch becomes an attempt.** g1t adopts it with the pusher as | |
| 138 | author. A reviewer agent writes the intent it appears to serve and offers | |
| 139 | to attach it to an open intent it matches. From there it gets the same | |
| 140 | checks, arena and queue as agent work. | |
| 141 | - **A push to `main` follows repo policy.** Protected: refused with a message | |
| 142 | saying which ref to push to instead, so it enters the queue. Open: accepted | |
| 143 | and treated as "`main` moved", which re-verifies the queue and triggers | |
| 144 | resolve-on-move for every open attempt. | |
| 145 | - **Context is an open format.** A commit trailer names the session that | |
| 146 | produced it, so any tool can attach its transcript. Commits without one are | |
| 147 | shown in why-blame as "pushed by a person, no session". | |
| 148 | - **Approval rules.** Per repo and per path: ship automatically, require a | |
| 149 | named person, or require a person when the change is large or the reviewer | |
| 150 | agent is unsure. | |
| 151 | ||
| 152 | ### Joining work that is already running | |
| 153 | ||
| 154 | - **Every session has a live page** that works on a phone: the transcript as | |
| 155 | it streams, the current diff, check results. | |
| 156 | - **Steer.** Send a message, pause, or redirect. Hosted agents receive it | |
| 157 | immediately; a person's own Claude Code receives it at its next turn | |
| 158 | through the CLI hooks. | |
| 159 | - **Answer.** When an agent is blocked on a question, it appears in a "needs | |
| 160 | you" inbox and as a notification. The answer resumes the agent. | |
| 161 | - **Take over and hand back.** Check out the attempt's fork, commit by hand, | |
| 162 | push, and let the agent continue from there. | |
| 163 | ||
| 164 | ### Planning by writing | |
| 165 | ||
| 166 | - **Brief.** Write the outcome in prose on the site, or commit it as a | |
| 167 | markdown file. A planner agent turns it into a project: intents, acceptance | |
| 168 | checks, dependencies. The person edits the graph before anything starts. | |
| 169 | - **Plan from their own agent.** The same operations are MCP tools, so a | |
| 170 | person can plan in their own Claude Code session and create the project | |
| 171 | from there. | |
| 172 | - **The brief stays the source of truth.** Editing it later re-plans: new | |
| 173 | intents are added, obsolete ones are closed. | |
| 174 | ||
| 175 | ### Seeing what moved | |
| 176 | ||
| 177 | - **Project page.** The outcome, the intent graph coloured by state, and how | |
| 178 | many acceptance checks pass now compared with when the project started. | |
| 179 | - **Digest.** An agent-written summary per project and per person: what | |
| 180 | shipped, what is blocked on whom, which conflicts were resolved, what it | |
| 181 | cost. | |
| 182 | - **Timeline.** Every event (push, steer, check, conflict, ship) in order, | |
| 183 | each linked to the session and the person or agent behind it. | |
| 184 | ||
| 185 | ## One session, any surface | |
| 186 | ||
| 187 | A session belongs to g1t, not to the device it started on. The browser, a | |
| 188 | phone and Claude Code are views of the same session. | |
| 189 | ||
| 190 | - **Browser and phone.** The site is a responsive, installable web app with | |
| 191 | push notifications. Everything a person does (brief, steer, answer, | |
| 192 | approve, ship) works there. | |
| 193 | - **Claude Code.** Through `mcp.g1t.sh` and the CLI hooks, a local session is | |
| 194 | a g1t session: its transcript syncs as it runs and it appears in mission | |
| 195 | control like any other. | |
| 196 | - **Moving a session.** A local session can be sent to the cloud: a hosted | |
| 197 | agent takes over the fork and the transcript and continues, so the laptop | |
| 198 | can close. A hosted session can be pulled down: the CLI checks out the fork | |
| 199 | and resumes it in local Claude Code with its history. | |
| 200 | - **Limit.** A session running only on a laptop stops when the laptop does. | |
| 201 | It can be steered between turns but not continued until it is moved or the | |
| 202 | laptop is back. | |
| 203 | ||
| 204 | ## For people who do not write code | |
| 205 | ||
| 206 | - **Documents are first-class.** Specs, guides, policies and decisions live | |
| 207 | in repos as markdown, shown in a Docs view: rendered pages, edited in the | |
| 208 | browser like a document, with inline comments. "Suggest a change" is an | |
| 209 | attempt and "publish" is ship, without git vocabulary. | |
| 210 | - **Document intents.** "Write the onboarding guide for the billing API" is | |
| 211 | an intent. Its acceptance checks are a checklist judged by a reviewer agent | |
| 212 | instead of commands. Agents draft and revise; people comment and approve. | |
| 213 | - **Templates.** Product brief, RFC, decision record. A filled-in template is | |
| 214 | a brief the planner can turn into a project. | |
| 215 | - **Explain.** Ask about any repo, project or change in plain language and | |
| 216 | get an answer with links to the code and sessions behind it. | |
| 217 | - **Living documentation.** g1t generates "how this works" pages from the | |
| 218 | code and keeps them current. When a shipped change contradicts a document, | |
| 219 | an intent opens to update it. | |
| 220 | - **See it, don't read it.** Every attempt on a deployable repo gets a | |
| 221 | preview URL (Workers Builds from the attempt's fork), so an approver clicks | |
| 222 | through the result instead of reading a diff. Changes are also summarised | |
| 223 | in plain language. | |
| 224 | - **Roles.** Viewer, commenter, planner, approver: a person can plan and | |
| 225 | approve work without ever cloning a repo. | |
| 226 | ||
| 227 | ## The macro view | |
| 228 | ||
| 229 | The hierarchy above a single repo: | |
| 230 | ||
| 231 | | Level | What it is | | |
| 232 | | --- | --- | | |
| 233 | | **Workspace** | A company or team: its people, repos, agents, budget and policies. | | |
| 234 | | **Initiative** | A business outcome with an owner and measurable results, e.g. "move billing to usage-based pricing". Spans any number of repos. | | |
| 235 | | **Project** | One deliverable inside an initiative: a brief and its graph of intents. | | |
| 236 | | **Intent / Attempt** | As above. An intent may touch several repos; its attempt then holds one fork per repo and they land together. | | |
| 237 | ||
| 238 | ### Portfolio | |
| 239 | ||
| 240 | One page answers "where is the business" across every initiative: | |
| 241 | ||
| 242 | - **Health** per initiative: on track, at risk, or blocked, derived from | |
| 243 | facts (checks passing, intents stalled, questions waiting on a person), | |
| 244 | not self-reported. | |
| 245 | - **Progress** as measurable results: acceptance checks passing, intents | |
| 246 | shipped out of planned, and the trend since the start. | |
| 247 | - **Forecast** from actual throughput: at the current rate, when the | |
| 248 | remaining intents land. | |
| 249 | - **Spend** in tokens and dollars against a budget, per initiative. | |
| 250 | - **Waiting on people**: every decision or approval a person owes, by name. | |
| 251 | - **Roadmap**: initiatives laid out as now, next, later, with optional | |
| 252 | time-boxed cycles for teams that work in sprints. | |
| 253 | ||
| 254 | ### Status without asking | |
| 255 | ||
| 256 | - **Standup.** An agent writes a daily report per initiative and one for the | |
| 257 | whole workspace: what shipped, what changed direction, what is at risk and | |
| 258 | why, what needs a person. Delivered by email or webhook. | |
| 259 | - **Ask.** A question box over the full event log and all sessions: "what | |
| 260 | happened on the billing migration since Monday?" answers with links to the | |
| 261 | sessions and commits behind each claim. | |
| 262 | ||
| 263 | ### Long-running agents | |
| 264 | ||
| 265 | Work that runs for days needs supervision that does not depend on someone | |
| 266 | watching. | |
| 267 | ||
| 268 | - **Checkpoints.** A long attempt reports milestones against its intent, so | |
| 269 | progress is visible before anything ships. | |
| 270 | - **Stall and drift detection.** An attempt with no meaningful progress, or | |
| 271 | whose changes have wandered away from its intent, is flagged and can be | |
| 272 | stopped or re-briefed automatically. | |
| 273 | - **Budgets.** Hard limits on spend and time per attempt, project and | |
| 274 | initiative. | |
| 275 | ||
| 276 | ### Context hub | |
| 277 | ||
| 278 | Agents working across repos and days need context that outlives any one | |
| 279 | session and reaches beyond the code. The context hub is one place an agent | |
| 280 | asks, whatever the source. | |
| 281 | ||
| 282 | | Source | What it holds | How it gets there | | |
| 283 | | --- | --- | --- | | |
| 284 | | **Memory** | Decisions, conventions, gotchas, facts about systems | Written by agents and people in g1t | | |
| 285 | | **Code and sessions** | The repos, and the reasoning behind every change | Already in g1t | | |
| 286 | | **Connected sources** | Jira and Linear tickets, Notion and Confluence pages, Google Drive documents, Slack threads, Sentry issues | Connectors, authorised per workspace | | |
| 287 | ||
| 288 | How it behaves: | |
| 289 | ||
| 290 | - **One search.** An agent asks a question and gets ranked results across | |
| 291 | all sources, each labelled with where it came from, who wrote it, and how | |
| 292 | fresh it is. | |
| 293 | - **Connected sources stay where they are.** g1t indexes them for search and | |
| 294 | fetches the current version when an agent opens one. The external system | |
| 295 | remains the source of truth, and a link placed on an intent ("see | |
| 296 | JIRA-482", a Notion URL) is pulled into the agent's starting context. | |
| 297 | - **Permissions carry over.** A connector only exposes what the connecting | |
| 298 | account can see, and a workspace admin chooses which spaces, projects or | |
| 299 | channels are included. | |
| 300 | - **External content is untrusted.** A ticket or page can contain text meant | |
| 301 | to manipulate an agent. It is marked as reference material, never treated | |
| 302 | as instructions. | |
| 303 | - **Documentation is separate.** Context is what agents know; documentation | |
| 304 | is what people read, and it is generated from context and code. | |
| 305 | ||
| 306 | Memory is the part of the hub that g1t owns and agents write to: | |
| 307 | ||
| 308 | - **Memory is written freely.** Any agent or person adds an entry with one | |
| 309 | call: a decision, a convention, a gotcha, a fact about a system. No review | |
| 310 | gate. Each entry records who wrote it, from which session, and when. | |
| 311 | - **It is still a repository.** Each workspace has a memory repo in | |
| 312 | Artifacts, so every write is a commit: versioned, attributable, and | |
| 313 | revertible. | |
| 314 | - **It is kept healthy by an agent.** A consolidation agent merges | |
| 315 | duplicates, retires entries that newer ones contradict, and flags | |
| 316 | conflicts it cannot settle. People can pin an entry (agents may not change | |
| 317 | it), correct it, or retract it. | |
| 318 | - **Agents read it.** Every session starts with the context relevant to its | |
| 319 | intent, found by search, and can query more through MCP. | |
| 320 | - **Updates are events.** A memory write or a change in a connected source | |
| 321 | is an event, so "when context changes, update the affected docs" is an | |
| 322 | automation, on by default. | |
| 323 | - **It is scoped inside the workspace.** Some context applies to the whole | |
| 324 | workspace, some to one initiative, project or repo, so an agent gets what | |
| 325 | applies to its work. | |
| 326 | - **It never crosses workspaces.** A workspace is the isolation boundary: its | |
| 327 | context, sessions and private repos are invisible to every other | |
| 328 | workspace, and an agent's token is bound to one workspace. | |
| 329 | ||
| 330 | ## Working in g1t | |
| 331 | ||
| 332 | - **Mission control.** The signed-in home page: every running session, every | |
| 333 | intent waiting on a decision, and what shipped, across all repos. | |
| 334 | - **Projects.** Group intents across repos toward one outcome and track how | |
| 335 | many are open, racing, or shipped. | |
| 336 | - **Steering.** Send a message to a running attempt, or to all attempts on an | |
| 337 | intent at once, without stopping them. | |
| 338 | - **Automations.** Rules that start work without a person (next section). | |
| 339 | ||
| 340 | ## Automations and integrations | |
| 341 | ||
| 342 | An automation is **when** an event happens, **if** conditions hold, **do** | |
| 343 | something. They are defined as files in the repo (`.g1t/automations/`), the | |
| 344 | way GitHub Actions workflows are, and can also be built in the UI. | |
| 345 | ||
| 346 | ### Events that can trigger one | |
| 347 | ||
| 348 | | Source | Examples | | |
| 349 | | --- | --- | | |
| 350 | | Git | push, ship, check failed, `main` moved | | |
| 351 | | g1t | intent opened, attempt stalled, context updated, handoff declined, budget reached | | |
| 352 | | Time | cron schedule | | |
| 353 | | Integrations | Sentry issue, PagerDuty incident, Linear or Jira ticket, Slack message or mention, GitHub issue, Stripe event | | |
| 354 | | Anything else | a signed generic webhook, or an email to a per-repo address | | |
| 355 | ||
| 356 | **Actions**: open an intent (optionally racing N attempts with a named | |
| 357 | agent), message a running attempt, update documentation, notify, call a | |
| 358 | webhook, write back to the source system. | |
| 359 | ||
| 360 | **Example: Sentry.** A new production error arrives. The automation opens an | |
| 361 | intent with the stack trace, release and frequency as its brief. Why-blame | |
| 362 | finds the session that wrote the failing line, so the fixing agent starts | |
| 363 | with the original reasoning. When the fix ships, g1t comments on the Sentry | |
| 364 | issue and resolves it. | |
| 365 | ||
| 366 | ### Rules every automation obeys | |
| 367 | ||
| 368 | - **Deduplication.** The same Sentry issue firing 500 times maps to one | |
| 369 | intent. | |
| 370 | - **Limits.** Concurrency and budget caps per automation. | |
| 371 | - **Loop protection.** Work started by an automation cannot retrigger the | |
| 372 | same automation without a person in between. | |
| 373 | - **External input is untrusted.** A webhook payload can contain text written | |
| 374 | by an attacker. Agents started by external events run with reduced | |
| 375 | permissions and cannot ship without the repo's approval rule passing. | |
| 376 | ||
| 377 | **Checks** are the other half of what GitHub Actions does: build and test | |
| 378 | commands declared in `.g1t/checks.yaml`, run in sandboxes on every attempt | |
| 379 | and on every combined state in the landing queue. | |
| 380 | ||
| 381 | Agents can also reach integrations directly: an agent definition lists MCP | |
| 382 | servers (Sentry, Linear and so on) it may use while working. | |
| 383 | ||
| 384 | ## Agents and models | |
| 385 | ||
| 386 | ### Defining an agent | |
| 387 | ||
| 388 | An agent is a file (`.g1t/agents/<name>.md`, or in the workspace library): | |
| 389 | instructions, the harness and model to run, the tools and MCP servers it may | |
| 390 | use, its sandbox image, permissions and budget. Agents take roles: planner, | |
| 391 | implementer, reviewer, conflict resolver, documenter, memory consolidator. | |
| 392 | Each role has a default that a repo can replace. | |
| 393 | ||
| 394 | ### Where it runs, and on whose model | |
| 395 | ||
| 396 | | Option | How it works | Fits | | |
| 397 | | --- | --- | --- | | |
| 398 | | Hosted, g1t's model | g1t runs the sandbox and bills usage | Getting started; no keys to manage | | |
| 399 | | Hosted, your API key | Same sandbox, your Anthropic, OpenAI or Google key | Teams with existing contracts | | |
| 400 | | Hosted, your endpoint | Any OpenAI-compatible URL: Bedrock, Vertex, Azure, a self-hosted model | Private or fine-tuned models | | |
| 401 | | Your runner | A g1t runner daemon on your own machines picks up attempts | Code or models that may not leave your network | | |
| 402 | | Your own session | Local Claude Code, Cursor or any MCP client joins through `mcp.g1t.sh` | Individuals; subscription plans | | |
| 403 | ||
| 404 | Decisions behind this: | |
| 405 | ||
| 406 | - **g1t does not build its own agent loop.** It runs existing harnesses | |
| 407 | (Claude Code first, through its headless mode) behind a small runner | |
| 408 | contract: a container image, an entry command, and session events reported | |
| 409 | through the CLI. Other harnesses plug in by meeting the contract. | |
| 410 | - **All hosted model traffic goes through Cloudflare AI Gateway.** That gives | |
| 411 | one place for spend tracking, budgets, rate limits, fallback and logs, | |
| 412 | whichever provider or endpoint is behind it. | |
| 413 | - **Subscriptions stay local.** A Claude subscription cannot be used by a | |
| 414 | hosted sandbox; it needs an API key. People on subscriptions use their own | |
| 415 | Claude Code session, which is a full participant. | |
| 416 | - **Keys are secrets.** Stored in Cloudflare Secrets Store, injected into the | |
| 417 | sandbox for one attempt, never shown again. | |
| 418 | ||
| 419 | ### Choosing the right agent automatically | |
| 420 | ||
| 421 | Because several agents can race on the same intent, every arena is an | |
| 422 | evaluation on real work. g1t records, per repo and per kind of intent, each | |
| 423 | agent's win rate, cost and time. That produces a leaderboard, and a routing | |
| 424 | policy: send each new intent to the agent that wins that kind most often, | |
| 425 | start with the cheapest that is good enough, and escalate to a stronger one | |
| 426 | when checks fail. | |
| 427 | ||
| 428 | ## What GitHub ships today, and where g1t differs | |
| 429 | ||
| 430 | GitHub's Agent HQ and Copilot app give each agent session its own git | |
| 431 | worktree and branch, list sessions in a mission-control view grouped by | |
| 432 | project, and let a task be assigned to several agents so their output can be | |
| 433 | compared. Underneath, the unit of work is still a branch and a pull request. | |
| 434 | ||
| 435 | | | GitHub | g1t | | |
| 436 | | --- | --- | --- | | |
| 437 | | Where a session works | A worktree on one developer's machine, or a cloud sandbox | A server-side fork that any agent on any machine can join and anyone can open | | |
| 438 | | Agent context | Lives in the app's session view | Stored with the repository and linked from each commit (why-blame) | | |
| 439 | | Several agents on one task | Separate pull requests to compare by hand | One intent, one arena, ranked attempts | | |
| 440 | | Collisions between agents | Found as merge conflicts at the end | Flagged during the work (overlap radar) | | |
| 441 | | Landing changes | One pull request at a time | A merge queue that ships the winner and rebases or closes the rest | | |
| 442 | | Which agents | Those offered through a Copilot subscription | Any MCP client, plus hosted agents | | |
| 443 | ||
| 444 | ## How agents connect | |
| 445 | ||
| 446 | 1. **Bring your own agent.** A remote MCP server at `mcp.g1t.sh` lets Claude | |
| 447 | Code (or any MCP client) list intents, claim one, get a clone URL and | |
| 448 | token, report progress and submit. Adding it is one command; sign-in is a | |
| 449 | browser OAuth flow with no token to paste. The `g1t` CLI installs Claude | |
| 450 | Code hooks that upload the session transcript as the agent works. | |
| 451 | 2. **Hosted agents.** Press "Run 10 attempts" on an intent. g1t starts | |
| 452 | sandboxes (Cloudflare Sandbox SDK), each running a coding agent headless | |
| 453 | against its own fork. | |
| 454 | 3. **API and CLI.** Everything above is available at `api.g1t.sh` and | |
| 455 | through `g1t`. | |
| 456 | ||
| 457 | ## Public surfaces | |
| 458 | ||
| 459 | | Host | What it serves | | |
| 460 | | --- | --- | | |
| 461 | | `g1t.sh` | The site, git over HTTPS, git over SSH | | |
| 462 | | `api.g1t.sh` | Versioned REST API with a published OpenAPI document, cursor pagination, rate-limit headers, idempotency keys on writes, server-sent events for live attempt state, and signed webhooks | | |
| 463 | | `mcp.g1t.sh` | Remote MCP server over streamable HTTP | | |
| 464 | ||
| 465 | g1t is its own OAuth 2.1 authorization server: authorization code with PKCE, | |
| 466 | dynamic client registration, discovery metadata, refresh tokens, and scopes | |
| 467 | per resource (`repo:read`, `repo:write`, `intent:write`, `attempt:write`). | |
| 468 | MCP clients, the CLI (device flow) and third-party apps all use it. Access | |
| 469 | tokens and SSH keys remain for git itself. | |
| 470 | ||
| 471 | ## Architecture | |
| 472 | ||
| 473 | | Component | Language | Runs on | Responsibility | | |
| 474 | | --- | --- | --- | --- | | |
| 475 | | `packages/contracts` | TypeScript | — | The interface of every service, the event catalogue, shared types. Services and clients depend on this, never on each other's code. | | |
| 476 | | `services/identity` | TypeScript | Worker + D1 | Accounts, sessions, SSH keys, access tokens; later the OAuth server | | |
| 477 | | `services/repos` | TypeScript | Worker + D1 + Artifacts | Repository registry, contents, forks, git over HTTPS. Storage sits behind a `GitStore` port with an Artifacts adapter. | | |
| 478 | | `services/events` | TypeScript | Worker + Queues + D1 | The event bus: durable log, and one queue per subscribing service | | |
| 479 | | `services/work` | TypeScript | Worker + D1 | Intents, attempts, sessions; later a Durable Object per repo for the landing queue and live state | | |
| 480 | | `apps/web` | TypeScript | Worker | Server-rendered site. Holds no data; calls services over RPC. | | |
| 481 | | `apps/api` (next) | TypeScript | Worker | REST API (`api.g1t.sh`) and MCP server (`mcp.g1t.sh`) over the same services | | |
| 482 | | `crates/sshd` | Rust | Container | Git over SSH, bridged to Artifacts | | |
| 483 | | `crates/merged` | Rust | Container | Trial merges, conflict matrix, landing merges (needs real git; the Artifacts binding is read-only) | | |
| 484 | | `crates/core` | Rust | native and WASM | pkt-line, packfile and diff code shared by the above and by the Worker | | |
| 485 | | `crates/g1t` | Rust | user's machine | CLI: auth, SSH proxy, Claude Code hooks, intents and attempts | | |
| 486 | | runner image | — | Sandbox | Hosted agent environment | | |
| 487 | ||
| 488 | Storage: Artifacts for repositories (one fork per attempt), D1 for accounts | |
| 489 | and metadata, R2 for session transcripts and logs, Durable Object SQLite for | |
| 490 | per-repo coordination state. | |
| 491 | ||
| 492 | How the services fit together: | |
| 493 | ||
| 494 | - **Each service is its own Worker with its own database.** It deploys, | |
| 495 | scales and fails on its own. Callers reach it through a typed RPC binding | |
| 496 | to the interface in `packages/contracts`. | |
| 497 | - **Expected failures are values.** Every call returns a `Result`, so "not | |
| 498 | found" or "forbidden" crosses a service boundary as data. | |
| 499 | - **Side effects travel as events.** A service publishes what happened | |
| 500 | (`git.push`, `intent.opened`, `attempt.started`, …) to the bus and does not | |
| 501 | call other services to react. Each subscriber consumes from its own queue. | |
| 502 | Timelines, webhooks and automations read the same stream, which is what | |
| 503 | lets something like GitHub Actions be built on top. | |
| 504 | - **Every read takes the viewer.** Authorization is decided inside the | |
| 505 | service that owns the data, not by its callers. | |
| 506 | ||
| 507 | The Workers runtime scales request handling on its own, so the edge layer | |
| 508 | stays in TypeScript. Rust is used where there is real computation or a real | |
| 509 | protocol to implement. | |
| 510 | ||
| API and MCP server, Rust identity service, registration, site redesign | 511 | ## Languages |
| 512 | ||
| 513 | The site is TypeScript. Everything behind it is Rust, compiled to | |
| 514 | WebAssembly for Workers and natively for containers and the CLI. Services | |
| 515 | are being ported one at a time; identity is done. Rust services speak a | |
| 516 | small JSON protocol over service bindings (`POST /rpc/<method>`), with the | |
| 517 | types in `crates/contracts`. | |
| 518 | ||
| 519 | ## Identifiers | |
| 520 | ||
| 521 | Every id is a [TypeID](https://github.com/jetify-com/typeid): a prefix naming | |
| 522 | the kind of thing, then a UUIDv7 in lowercase base32, such as | |
| 523 | `att_01jb2k7x9hfq0b3zj0f5s2m8ra`. | |
| 524 | ||
| 525 | - The prefix makes an id self-describing and stops ids of different kinds | |
| 526 | being mixed up. | |
| 527 | - Ids sort by creation time as plain strings. In SQLite (D1 and Durable | |
| 528 | Objects) that keeps inserts at the end of the primary-key index instead of | |
| 529 | scattering them, and gives time-ordered paging for free. | |
| 530 | - The suffix decodes to a standard UUIDv7 for any system that wants one. | |
| 531 | - Ids are made by the service that creates the record, not by the database, | |
| 532 | so they work across services and can be assigned before a write. | |
| 533 | ||
| 534 | ## Events at scale, and audit | |
| 535 | ||
| 536 | The current event log is a single D1 database. That is fine for a | |
| 537 | prototype and wrong for the target: D1 is one writer and 10 GB. The design | |
| 538 | for volume splits storage by how the data is read. | |
| 539 | ||
| 540 | | Tier | Store | Holds | Read by | | |
| 541 | | --- | --- | --- | --- | | |
| 542 | | Hot | A Durable Object per repository, with SQLite | Recent events for that repo | Timelines, live pages over WebSocket | | |
| 543 | | Complete | Cloudflare Pipelines into R2 as Apache Iceberg | Every event, forever, partitioned by day and workspace | Analytics, standups, "ask", export | | |
| 544 | | Audit | The same R2 store, under object lock | Who did what, from where, with which credential | Compliance, investigation | | |
| 545 | ||
| 546 | - **No single hot database.** Each repository's recent events live with that | |
| 547 | repository, so load spreads across as many objects as there are repos. | |
| 548 | - **The complete record is files, not rows.** Iceberg on R2 has no practical | |
| 549 | size limit and is queried with SQL. | |
| 550 | - **Audit is a property of every event.** The envelope carries the actor | |
| 551 | (person, agent, token or system), the credential used, the request id and | |
| 552 | the source address. Audit entries for a workspace are hash-chained, so a | |
| 553 | removed or altered entry is detectable, and are written under a retention | |
| 554 | lock. | |
| 555 | - **Delivery is at least once.** Consumers are idempotent on the event id. | |
| 556 | ||
| Initial g1t: services, event bus, intents and attempts | 557 | ## Accounts and forge basics |
| 558 | ||
| 559 | - Registration with email verification, sign-in, forgot password (Cloudflare | |
| 560 | Email Sending), Turnstile on public forms. | |
| 561 | - GitHub sign-in, SSH keys, access tokens, active sessions. | |
| 562 | - Profiles, public and private repositories, repository search (D1 full-text). | |
| 563 | - Rendered README, syntax highlighting, commit history, diffs. | |
| 564 | ||
| 565 | ## Built on Cloudflare | |
| 566 | ||
| 567 | | Need | Product | | |
| 568 | | --- | --- | | |
| 569 | | Repositories; a fork per attempt; data residency per workspace | Artifacts (forks, jurisdictions) | | |
| 570 | | Reacting to pushes | Artifacts event subscriptions on Queues | | |
| 571 | | Preview URL per attempt; deploy on ship | Workers Builds and previews | | |
| 572 | | Site, API, MCP, git front end | Workers | | |
| 573 | | Per-repo coordination, live updates | Durable Objects | | |
| 574 | | Attempt lifecycles, automations | Workflows, Cron Triggers | | |
| 575 | | Agent sandboxes, SSH server, merge engine | Sandbox SDK and Containers | | |
| 576 | | Fast starts on large repos | ArtifactFS | | |
| 577 | | Model traffic, spend, budgets | AI Gateway | | |
| 578 | | Summaries, embeddings | Workers AI | | |
| 579 | | Context hub search | Vectorize | | |
| 580 | | Accounts and metadata | D1 | | |
| 581 | | Transcripts and logs | R2 | | |
| 582 | | Email, bot protection, keys | Email Sending, Turnstile, Secrets Store | | |
| 583 | ||
| 584 | ## The submission | |
| 585 | ||
| 586 | - **g1t is built on g1t.** This repository is hosted on g1t.sh, its features | |
| 587 | are opened as intents and built by racing agents, and it deploys from | |
| 588 | Artifacts through Workers Builds. The history is the proof. | |
| 589 | - **The demo follows one story.** A brief becomes a project; twelve intents | |
| 590 | fan out to dozens of agents; agents notice each other, hand off, and | |
| 591 | resolve a conflict; reviewers triage; the queue lands everything on | |
| 592 | `main`; why-blame explains a line; the portfolio shows where it all | |
| 593 | stands. Then the same thing at a thousand agents. | |
| 594 | - **Judges can try it in a minute.** Open registration on g1t.sh, one-click | |
| 595 | import of a GitHub repo, one command to connect Claude Code, a seeded demo | |
| 596 | workspace, and a single deploy command for running their own copy. | |
| 597 | - **The formats are open.** The commit trailers, session format and runner | |
| 598 | contract are published so other tools can interoperate. | |
| 599 | ||
| 600 | ## Build order | |
| 601 | ||
| 602 | Done: site on g1t.sh; git over HTTPS; accounts and private repos; the four | |
| 603 | services and the event bus; intents, attempts (a fork each) and session | |
| 604 | storage, with pages for each. SSH server written, not deployed. | |
| 605 | ||
| 606 | 1. `api.g1t.sh` and `mcp.g1t.sh`; CLI with Claude Code hooks, so an outside | |
| 607 | agent can claim an intent, push, and record its session. | |
| 608 | 2. OAuth server. | |
| 609 | 3. SSH deployed; registration, email verification, forgot password; search; | |
| 610 | GitHub import; this repo hosted on g1t. | |
| 611 | 4. Hosted agents in sandboxes behind the runner contract; agent | |
| 612 | definitions; AI Gateway; push events; checks. | |
| 613 | 5. Merge engine with structural merge; landing queue with speculative | |
| 614 | checks; projected main; resolve-on-move. | |
| 615 | 6. Arena, diffs, proof bundles, reviewer and adversarial reviewer, risk | |
| 616 | tiers; work registry with overlap radar, | |
| 617 | duplicate detection, handoff, ask and wait. | |
| 618 | 7. Adopting outside pushes; protected `main`; approval rules. | |
| 619 | 8. Projects with briefs, planner and dependency graph; mission control with | |
| 620 | the "needs you" inbox; live session pages and steering. | |
| 621 | 9. Why-blame, signed provenance, forkable sessions, digest and timeline, | |
| 622 | landing page. | |
| 623 | 10. Workspaces and initiatives; context hub (memory, unified search, | |
| 624 | Jira and Notion connectors); multi-repo intents. | |
| 625 | 11. Portfolio, standup and ask; checkpoints, stall detection, budgets. | |
| 626 | 12. Moving sessions between local and hosted; installable web app with | |
| 627 | notifications; preview URLs per attempt. | |
| 628 | 13. Docs view, document intents, templates, explain, living documentation, | |
| 629 | roles. | |
| 630 | 14. Automations: event bus, triggers, Sentry and generic webhook | |
| 631 | integrations, write-back. | |
| 632 | 15. Own keys, own endpoints, self-hosted runners; agent leaderboard and | |
| 633 | routing. | |
| 634 | 16. Large run (100+ agents across many intents), hardening, README, demo. | |
| 635 | ||
| 636 | Later: code search, mirroring to GitHub, passkeys, SSH | |
| 637 | on port 22 without the CLI proxy (needs the Workers inbound TCP private | |
| 638 | beta). |