Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.
| status.g1t.sh with incident management, invites that land you in the workspace, settings as pages, usage without quotas | 1 | # Incident runbook |
| 2 | ||
| 3 | For g1t staff. How to declare, run, resolve and write up an incident, and | |
| 4 | how to schedule maintenance, in sudo (**Platform → Incidents**, | |
| 5 | `https://sudo.g1t.sh/incidents`). Everything here is recorded in sudo's | |
| 6 | **Audit log** with your email. The public side is described in | |
| 7 | `apps/docs/src/content/docs/guides/status.md`. | |
| 8 | ||
| 9 | ## The short version | |
| 10 | ||
| 11 | 1. **Declare** as soon as people are affected. A wrong first guess is | |
| 12 | fine; silence is not. | |
| 13 | 2. **Update** publicly on the cadence the severity asks for, even when | |
| 14 | nothing changed ("Still working on it; next update by 15:30 UTC"). | |
| 15 | 3. **Resolve** when it is fixed, with one sentence on what fixed it. | |
| 16 | 4. **Write the postmortem** within five working days for SEV1 and SEV2. | |
| 17 | ||
| 18 | ## Severity | |
| 19 | ||
| 20 | | Severity | When | Public updates | Email subscribers | Postmortem | | |
| 21 | | --- | --- | --- | --- | --- | | |
| 22 | | SEV1 | g1t is down or unusable for most people, or data is at risk | Every 30 minutes at least | Yes | Required | | |
| 23 | | SEV2 | A core part (sign-in, git, the API, agents) is broken or badly degraded for many people | Hourly at least | Yes | Required | | |
| 24 | | SEV3 | One part is degraded or broken for some people; there is a way around it | As things change | Your call | If there is a lesson | | |
| 25 | | SEV4 | Little or no customer impact | As things change | No | No | | |
| 26 | ||
| 27 | When unsure, pick the higher one. Lowering it later is one change. | |
| 28 | ||
| 29 | ## Roles | |
| 30 | ||
| 31 | - **Incident commander (IC)**: runs the response. Decides, assigns work, | |
| 32 | keeps the timeline honest. Not necessarily the person fixing it. | |
| 33 | - **Communications**: writes the public updates and keeps the cadence. | |
| 34 | For a small incident the IC does both. | |
| 35 | ||
| 36 | Set both on the declare form or in **Roles** on the incident page. A | |
| 37 | change of role is written on the timeline. | |
| 38 | ||
| 39 | ## Declaring | |
| 40 | ||
| 41 | **Declare incident** on the Incidents page: | |
| 42 | ||
| 43 | 1. **Title**: what people notice, not the internals. "Pushes over HTTPS | |
| 44 | failing", not "pack-receiver 502s on eu-2". | |
| 45 | 2. **Severity**: see the table. | |
| 46 | 3. **Status**: usually **Investigating**. | |
| 47 | 4. **Impact began (UTC)**: leave empty for now, or backdate to when it | |
| 48 | really started. The time-to-acknowledge, mitigate and resolve figures | |
| 49 | count from here. | |
| 50 | 5. **Parts affected**: set each affected part to Degraded performance, | |
| 51 | Partial outage or Major outage. The status page shows that (or worse, | |
| 52 | if the checks say worse) until the incident is resolved. | |
| 53 | 6. **First public update**: what is broken, for whom, and what still works. | |
| 54 | 7. **Email subscribers**: "As the severity says" emails for SEV1 and SEV2. | |
| 55 | 8. **Declare and publish**. It is on status.g1t.sh within 30 seconds. | |
| 56 | ||
| 57 | ## Detected drafts | |
| 58 | ||
| Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders | 59 | When a part fails or is slow on four of its last five checks (a check a |
| 60 | minute), the status worker makes a **draft** incident and emails | |
| 61 | `STATUS_ALERT_EMAIL` (hey@flagon.io) with a link. A draft is not on the | |
| 62 | status page; the part's own state already is, from the checks. | |
| 63 | ||
| 64 | How the checks decide (`apps/status/src/detect.ts`, `probe.ts`): | |
| 65 | ||
| 66 | - **A slow answer is asked again at once.** It counts as slow only if the | |
| 67 | second answer is slow too, so one cold start or cache refill is not a | |
| 68 | slow check. A check that fails (an error or a timeout) is not asked | |
| 69 | again; four of five decides. | |
| 70 | - **Four of five, not three in a row.** A good check in between does not | |
| 71 | hide trouble that keeps coming; three good checks in a row end a run of | |
| 72 | trouble, however long it lasted, and the next trouble starts a new one | |
| 73 | with its own start. A draft's lines say "on 4 checks in a row" or "on 4 | |
| 74 | of 5 checks". | |
| 75 | - **Page loads are timed to the first byte.** The site's pages are loaded | |
| 76 | with a browser's user agent (ending `g1t-status/1.0 (+status.g1t.sh)`), | |
| 77 | because the site renders the whole page first for a crawler. See | |
| 78 | `docs/PERFORMANCE.md`. | |
| 79 | - **Every check is kept for 7 days**: how long it took, what it meant, and | |
| 80 | the Cloudflare data centre it ran from (from the answer's `cf-ray`). The | |
| 81 | incident page shows a **Checks** chart per part, from 30 minutes before | |
| 82 | the impact began to 30 minutes after it ended (at most a day), with the | |
| 83 | slow line, slow checks dotted and failures marked, and the latest 30 | |
| 84 | checks as a table. | |
| 85 | - **A draft nobody acknowledges is raised again**: the alert goes out once | |
| 86 | more after 45 minutes, then every 6 hours while it waits, with a note on | |
| 87 | its timeline each time. Publishing it, posting a note or dismissing it | |
| 88 | acknowledges it and stops the reminders. | |
| 89 | - **Deploys are announced.** `scripts/deploy.mjs` tells the status worker | |
| 90 | when a deploy starts and finishes (`docs/DEPLOYING.md`); during one, and | |
| 91 | for 3 minutes after, trouble is counted but not drafted unless it | |
| 92 | outlasts the deploy. | |
| status.g1t.sh with incident management, invites that land you in the workspace, settings as pages, usage without quotas | 93 | |
| 94 | Open it from the **Drafts** tab (the sidebar's Incidents count includes | |
| 95 | drafts) and either: | |
| 96 | ||
| 97 | - **Publish to the status page**, with a public title (the "Detected:" | |
| 98 | prefix is dropped) and a first update; or | |
| 99 | - **Dismiss draft** with a reason, if it was a blip. Dismissed drafts are | |
| 100 | listed under Resolved and never appear publicly. | |
| 101 | ||
| 102 | While an incident is open on a part, more failures on it add a line to | |
| 103 | that incident's timeline instead of a new draft, and the part answering | |
| Merge status detection: first-byte speed probe, deploy windows, 4 of 5 with a re-check, check history, reminders | 104 | again (three good checks in a row) adds an "answering again" line. |
| status.g1t.sh with incident management, invites that land you in the workspace, settings as pages, usage without quotas | 105 | |
| 106 | ## Running it | |
| 107 | ||
| 108 | The incident page has the timers at the top (open for, to acknowledge, | |
| 109 | to mitigate, to resolve) and two columns: the timeline on the left, and | |
| 110 | resolve, parts, roles, follow-ups and links on the right. | |
| 111 | ||
| 112 | **Post an update** (left): | |
| 113 | ||
| 114 | - **Public update**: shown on the status page and in the feeds; emailed | |
| 115 | to subscribers when the box is ticked (ticked by default for SEV1 and | |
| 116 | SEV2). Changing the status always needs a public update, because the | |
| 117 | status page shows each status with words. | |
| 118 | - **Internal note**: staff only. Use it for findings, links to logs, | |
| 119 | who is doing what. Notes never reach the status page, the feeds or | |
| 120 | email, but the postmortem's timeline starts from them, so keep them | |
| 121 | factual. | |
| 122 | - **Status**: Investigating → Identified (cause known) → Monitoring (fix | |
| 123 | out, watching) → Resolved. Reaching Monitoring marks it mitigated. | |
| 124 | - **Severity** and **Change the parts' impact**: change them as you learn | |
| 125 | more; each change is a line on the timeline. | |
| 126 | ||
| 127 | The timeline shows everything, newest first: public updates highlighted | |
| 128 | in lavender with the number of subscribers emailed, notes in gray, and | |
| 129 | one-line entries for every status, severity, impact, role and follow-up | |
| 130 | change. | |
| 131 | ||
| 132 | **Follow-ups**: add anything that should change so it does not happen | |
| 133 | again, with an owner. Tick them off as they are done; they fill the | |
| 134 | postmortem's action items. | |
| 135 | ||
| 136 | ## Resolving | |
| 137 | ||
| 138 | **Resolve** (right column) with one or two sentences: what is fixed and, | |
| 139 | if known, what fixed it. It posts the last public update, sets the | |
| 140 | resolved time, and the parts stop showing the incident's impact. If it | |
| 141 | comes back, post an update with an earlier status: that reopens it. | |
| 142 | ||
| 143 | ## Postmortem | |
| 144 | ||
| 145 | From a resolved incident, **Write the postmortem**. The editor starts | |
| 146 | with: | |
| 147 | ||
| 148 | - **Impact**: the parts, how badly and for how long; | |
| 149 | - **Timeline**: every timeline line in UTC, internal notes included; | |
| 150 | - **Action items**: the follow-ups. | |
| 151 | ||
| 152 | **Edit out anything internal** (hostnames, customer names, people's | |
| 153 | names in blame) before publishing. Write the **Summary** and **Root | |
| 154 | cause** (both required to publish), and what went well and badly. A blank | |
| 155 | line starts a paragraph; lines starting with `- ` become a list. | |
| 156 | ||
| 157 | **Save draft** shows the preview beside the editor. **Save and publish** | |
| 158 | (or **Publish the saved draft**) puts it on the incident's public page, | |
| 159 | `https://status.g1t.sh/incidents/<id>#postmortem`, and marks the incident | |
| 160 | "Postmortem published". **Take down** removes it again. | |
| 161 | ||
| 162 | Blameless: say what the system allowed, not who slipped. | |
| 163 | ||
| 164 | ## Scheduled maintenance | |
| 165 | ||
| 166 | **Schedule maintenance** on the Incidents page: | |
| 167 | ||
| 168 | 1. Title, window (UTC, at most 72 hours, at most 180 days ahead), the | |
| 169 | parts affected, and a message saying what people will notice. | |
| 170 | 2. **Email subscribers** emails them now, when it starts and when it ends. | |
| 171 | ||
| 172 | It is listed as upcoming at once. The minutely job starts it when the | |
| 173 | window opens (its parts show "Under maintenance" and their checks stop | |
| 174 | counting against uptime) and completes it when the window ends. From its | |
| 175 | page you can post updates, **Start now**, **Complete** early, or | |
| 176 | **Cancel**. | |
| 177 | ||
| 178 | If the work breaks something beyond what was announced, declare an | |
| 179 | incident: an incident's impact shows over maintenance. | |
| 180 | ||
| 181 | ## Configuration | |
| 182 | ||
| 183 | The status worker (`apps/status/wrangler.jsonc`): | |
| 184 | ||
| 185 | | Setting | What it does | | |
| 186 | | --- | --- | | |
| 187 | | `send_email` binding `EMAIL` | Cloudflare Email Sending, as identity uses. Without it nothing is emailed; the feeds still work. | | |
| 188 | | `STATUS_SECRET` (secret) | Signs unsubscribe links. Without it, the subscribe form is off. `npx wrangler secret put STATUS_SECRET` | | |
| 189 | | `STATUS_ALERT_EMAIL` | Who hears about detected drafts. Empty sends none. | | |
| 190 | | `STATUS_URL`, `SUDO_URL` | Links in email and feeds, and the alert's link to sudo. | | |
| 191 | | `STATUS_FROM` | The From address. | | |
| 192 | ||
| 193 | sudo reaches the status worker only through its `STATUS` service binding | |
| 194 | (`StatusAdmin` entrypoint); status.g1t.sh itself has no way to write. |
This file's history is long; its oldest lines are credited to the oldest commit read.