| 1 | # Incident runbook |
| 2 | |
| 3 | For g1t staff. How to declare, run, resolve and write up an incident, and |
| 4 | how to schedule maintenance, in sudo (**Platform → Incidents**, |
| 5 | `https://sudo.g1t.sh/incidents`). Everything here is recorded in sudo's |
| 6 | **Audit log** with your email. The public side is described in |
| 7 | `apps/docs/src/content/docs/guides/status.md`. |
| 8 | |
| 9 | ## The short version |
| 10 | |
| 11 | 1. **Declare** as soon as people are affected. A wrong first guess is |
| 12 | fine; silence is not. |
| 13 | 2. **Update** publicly on the cadence the severity asks for, even when |
| 14 | nothing changed ("Still working on it; next update by 15:30 UTC"). |
| 15 | 3. **Resolve** when it is fixed, with one sentence on what fixed it. |
| 16 | 4. **Write the postmortem** within five working days for SEV1 and SEV2. |
| 17 | |
| 18 | ## Severity |
| 19 | |
| 20 | | Severity | When | Public updates | Email subscribers | Postmortem | |
| 21 | | --- | --- | --- | --- | --- | |
| 22 | | SEV1 | g1t is down or unusable for most people, or data is at risk | Every 30 minutes at least | Yes | Required | |
| 23 | | SEV2 | A core part (sign-in, git, the API, agents) is broken or badly degraded for many people | Hourly at least | Yes | Required | |
| 24 | | SEV3 | One part is degraded or broken for some people; there is a way around it | As things change | Your call | If there is a lesson | |
| 25 | | SEV4 | Little or no customer impact | As things change | No | No | |
| 26 | |
| 27 | When unsure, pick the higher one. Lowering it later is one change. |
| 28 | |
| 29 | ## Roles |
| 30 | |
| 31 | - **Incident commander (IC)**: runs the response. Decides, assigns work, |
| 32 | keeps the timeline honest. Not necessarily the person fixing it. |
| 33 | - **Communications**: writes the public updates and keeps the cadence. |
| 34 | For a small incident the IC does both. |
| 35 | |
| 36 | Set both on the declare form or in **Roles** on the incident page. A |
| 37 | change of role is written on the timeline. |
| 38 | |
| 39 | ## Declaring |
| 40 | |
| 41 | **Declare incident** on the Incidents page: |
| 42 | |
| 43 | 1. **Title**: what people notice, not the internals. "Pushes over HTTPS |
| 44 | failing", not "pack-receiver 502s on eu-2". |
| 45 | 2. **Severity**: see the table. |
| 46 | 3. **Status**: usually **Investigating**. |
| 47 | 4. **Impact began (UTC)**: leave empty for now, or backdate to when it |
| 48 | really started. The time-to-acknowledge, mitigate and resolve figures |
| 49 | count from here. |
| 50 | 5. **Parts affected**: set each affected part to Degraded performance, |
| 51 | Partial outage or Major outage. The status page shows that (or worse, |
| 52 | if the checks say worse) until the incident is resolved. |
| 53 | 6. **First public update**: what is broken, for whom, and what still works. |
| 54 | 7. **Email subscribers**: "As the severity says" emails for SEV1 and SEV2. |
| 55 | 8. **Declare and publish**. It is on status.g1t.sh within 30 seconds. |
| 56 | |
| 57 | ## Detected drafts |
| 58 | |
| 59 | When a part fails or is slow on four of its last five checks (a check a |
| 60 | minute), the status worker makes a **draft** incident and emails |
| 61 | `STATUS_ALERT_EMAIL` (hey@flagon.io) with a link. A draft is not on the |
| 62 | status page; the part's own state already is, from the checks. |
| 63 | |
| 64 | How the checks decide (`apps/status/src/detect.ts`, `probe.ts`): |
| 65 | |
| 66 | - **A slow answer is asked again at once.** It counts as slow only if the |
| 67 | second answer is slow too, so one cold start or cache refill is not a |
| 68 | slow check. A check that fails (an error or a timeout) is not asked |
| 69 | again; four of five decides. |
| 70 | - **Four of five, not three in a row.** A good check in between does not |
| 71 | hide trouble that keeps coming; three good checks in a row end a run of |
| 72 | trouble, however long it lasted, and the next trouble starts a new one |
| 73 | with its own start. A draft's lines say "on 4 checks in a row" or "on 4 |
| 74 | of 5 checks". |
| 75 | - **Page loads are timed to the first byte.** The site's pages are loaded |
| 76 | with a browser's user agent (ending `g1t-status/1.0 (+status.g1t.sh)`), |
| 77 | because the site renders the whole page first for a crawler. See |
| 78 | `docs/PERFORMANCE.md`. |
| 79 | - **Every check is kept for 7 days**: how long it took, what it meant, and |
| 80 | the Cloudflare data centre it ran from (from the answer's `cf-ray`). The |
| 81 | incident page shows a **Checks** chart per part, from 30 minutes before |
| 82 | the impact began to 30 minutes after it ended (at most a day), with the |
| 83 | slow line, slow checks dotted and failures marked, and the latest 30 |
| 84 | checks as a table. |
| 85 | - **A draft nobody acknowledges is raised again**: the alert goes out once |
| 86 | more after 45 minutes, then every 6 hours while it waits, with a note on |
| 87 | its timeline each time. Publishing it, posting a note or dismissing it |
| 88 | acknowledges it and stops the reminders. |
| 89 | - **Deploys are announced.** `scripts/deploy.mjs` tells the status worker |
| 90 | when a deploy starts and finishes (`docs/DEPLOYING.md`); during one, and |
| 91 | for 3 minutes after, trouble is counted but not drafted unless it |
| 92 | outlasts the deploy. |
| 93 | |
| 94 | Open it from the **Drafts** tab (the sidebar's Incidents count includes |
| 95 | drafts) and either: |
| 96 | |
| 97 | - **Publish to the status page**, with a public title (the "Detected:" |
| 98 | prefix is dropped) and a first update; or |
| 99 | - **Dismiss draft** with a reason, if it was a blip. Dismissed drafts are |
| 100 | listed under Resolved and never appear publicly. |
| 101 | |
| 102 | While an incident is open on a part, more failures on it add a line to |
| 103 | that incident's timeline instead of a new draft, and the part answering |
| 104 | again (three good checks in a row) adds an "answering again" line. |
| 105 | |
| 106 | ## Running it |
| 107 | |
| 108 | The incident page has the timers at the top (open for, to acknowledge, |
| 109 | to mitigate, to resolve) and two columns: the timeline on the left, and |
| 110 | resolve, parts, roles, follow-ups and links on the right. |
| 111 | |
| 112 | **Post an update** (left): |
| 113 | |
| 114 | - **Public update**: shown on the status page and in the feeds; emailed |
| 115 | to subscribers when the box is ticked (ticked by default for SEV1 and |
| 116 | SEV2). Changing the status always needs a public update, because the |
| 117 | status page shows each status with words. |
| 118 | - **Internal note**: staff only. Use it for findings, links to logs, |
| 119 | who is doing what. Notes never reach the status page, the feeds or |
| 120 | email, but the postmortem's timeline starts from them, so keep them |
| 121 | factual. |
| 122 | - **Status**: Investigating → Identified (cause known) → Monitoring (fix |
| 123 | out, watching) → Resolved. Reaching Monitoring marks it mitigated. |
| 124 | - **Severity** and **Change the parts' impact**: change them as you learn |
| 125 | more; each change is a line on the timeline. |
| 126 | |
| 127 | The timeline shows everything, newest first: public updates highlighted |
| 128 | in lavender with the number of subscribers emailed, notes in gray, and |
| 129 | one-line entries for every status, severity, impact, role and follow-up |
| 130 | change. |
| 131 | |
| 132 | **Follow-ups**: add anything that should change so it does not happen |
| 133 | again, with an owner. Tick them off as they are done; they fill the |
| 134 | postmortem's action items. |
| 135 | |
| 136 | ## Resolving |
| 137 | |
| 138 | **Resolve** (right column) with one or two sentences: what is fixed and, |
| 139 | if known, what fixed it. It posts the last public update, sets the |
| 140 | resolved time, and the parts stop showing the incident's impact. If it |
| 141 | comes back, post an update with an earlier status: that reopens it. |
| 142 | |
| 143 | ## Postmortem |
| 144 | |
| 145 | From a resolved incident, **Write the postmortem**. The editor starts |
| 146 | with: |
| 147 | |
| 148 | - **Impact**: the parts, how badly and for how long; |
| 149 | - **Timeline**: every timeline line in UTC, internal notes included; |
| 150 | - **Action items**: the follow-ups. |
| 151 | |
| 152 | **Edit out anything internal** (hostnames, customer names, people's |
| 153 | names in blame) before publishing. Write the **Summary** and **Root |
| 154 | cause** (both required to publish), and what went well and badly. A blank |
| 155 | line starts a paragraph; lines starting with `- ` become a list. |
| 156 | |
| 157 | **Save draft** shows the preview beside the editor. **Save and publish** |
| 158 | (or **Publish the saved draft**) puts it on the incident's public page, |
| 159 | `https://status.g1t.sh/incidents/<id>#postmortem`, and marks the incident |
| 160 | "Postmortem published". **Take down** removes it again. |
| 161 | |
| 162 | Blameless: say what the system allowed, not who slipped. |
| 163 | |
| 164 | ## Scheduled maintenance |
| 165 | |
| 166 | **Schedule maintenance** on the Incidents page: |
| 167 | |
| 168 | 1. Title, window (UTC, at most 72 hours, at most 180 days ahead), the |
| 169 | parts affected, and a message saying what people will notice. |
| 170 | 2. **Email subscribers** emails them now, when it starts and when it ends. |
| 171 | |
| 172 | It is listed as upcoming at once. The minutely job starts it when the |
| 173 | window opens (its parts show "Under maintenance" and their checks stop |
| 174 | counting against uptime) and completes it when the window ends. From its |
| 175 | page you can post updates, **Start now**, **Complete** early, or |
| 176 | **Cancel**. |
| 177 | |
| 178 | If the work breaks something beyond what was announced, declare an |
| 179 | incident: an incident's impact shows over maintenance. |
| 180 | |
| 181 | ## Configuration |
| 182 | |
| 183 | The status worker (`apps/status/wrangler.jsonc`): |
| 184 | |
| 185 | | Setting | What it does | |
| 186 | | --- | --- | |
| 187 | | `send_email` binding `EMAIL` | Cloudflare Email Sending, as identity uses. Without it nothing is emailed; the feeds still work. | |
| 188 | | `STATUS_SECRET` (secret) | Signs unsubscribe links. Without it, the subscribe form is off. `npx wrangler secret put STATUS_SECRET` | |
| 189 | | `STATUS_ALERT_EMAIL` | Who hears about detected drafts. Empty sends none. | |
| 190 | | `STATUS_URL`, `SUDO_URL` | Links in email and feeds, and the alert's link to sudo. | |
| 191 | | `STATUS_FROM` | The From address. | |
| 192 | |
| 193 | sudo reaches the status worker only through its `STATUS` service binding |
| 194 | (`StatusAdmin` entrypoint); status.g1t.sh itself has no way to write. |