| 1 | # Incident runbook |
| 2 | |
| 3 | For g1t staff. How to declare, run, resolve and write up an incident, and |
| 4 | how to schedule maintenance, in sudo (**Platform → Incidents**, |
| 5 | `https://sudo.g1t.sh/incidents`). Everything here is recorded in sudo's |
| 6 | **Audit log** with your email. The public side is described in |
| 7 | `apps/docs/src/content/docs/guides/status.md`. |
| 8 | |
| 9 | ## The short version |
| 10 | |
| 11 | 1. **Declare** as soon as people are affected. A wrong first guess is |
| 12 | fine; silence is not. |
| 13 | 2. **Update** publicly on the cadence the severity asks for, even when |
| 14 | nothing changed ("Still working on it; next update by 15:30 UTC"). |
| 15 | 3. **Resolve** when it is fixed, with one sentence on what fixed it. |
| 16 | 4. **Write the postmortem** within five working days for SEV1 and SEV2. |
| 17 | |
| 18 | ## Severity |
| 19 | |
| 20 | | Severity | When | Public updates | Email subscribers | Postmortem | |
| 21 | | --- | --- | --- | --- | --- | |
| 22 | | SEV1 | g1t is down or unusable for most people, or data is at risk | Every 30 minutes at least | Yes | Required | |
| 23 | | SEV2 | A core part (sign-in, git, the API, agents) is broken or badly degraded for many people | Hourly at least | Yes | Required | |
| 24 | | SEV3 | One part is degraded or broken for some people; there is a way around it | As things change | Your call | If there is a lesson | |
| 25 | | SEV4 | Little or no customer impact | As things change | No | No | |
| 26 | |
| 27 | When unsure, pick the higher one. Lowering it later is one change. |
| 28 | |
| 29 | ## Roles |
| 30 | |
| 31 | - **Incident commander (IC)**: runs the response. Decides, assigns work, |
| 32 | keeps the timeline honest. Not necessarily the person fixing it. |
| 33 | - **Communications**: writes the public updates and keeps the cadence. |
| 34 | For a small incident the IC does both. |
| 35 | |
| 36 | Set both on the declare form or in **Roles** on the incident page. A |
| 37 | change of role is written on the timeline. |
| 38 | |
| 39 | ## Declaring |
| 40 | |
| 41 | **Declare incident** on the Incidents page: |
| 42 | |
| 43 | 1. **Title**: what people notice, not the internals. "Pushes over HTTPS |
| 44 | failing", not "pack-receiver 502s on eu-2". |
| 45 | 2. **Severity**: see the table. |
| 46 | 3. **Status**: usually **Investigating**. |
| 47 | 4. **Impact began (UTC)**: leave empty for now, or backdate to when it |
| 48 | really started. The time-to-acknowledge, mitigate and resolve figures |
| 49 | count from here. |
| 50 | 5. **Parts affected**: set each affected part to Degraded performance, |
| 51 | Partial outage or Major outage. The status page shows that (or worse, |
| 52 | if the checks say worse) until the incident is resolved. |
| 53 | 6. **First public update**: what is broken, for whom, and what still works. |
| 54 | 7. **Email subscribers**: "As the severity says" emails for SEV1 and SEV2. |
| 55 | 8. **Declare and publish**. It is on status.g1t.sh within 30 seconds. |
| 56 | |
| 57 | ## Detected drafts |
| 58 | |
| 59 | When a part fails or is slow three checks in a row (three minutes), the |
| 60 | status worker makes a **draft** incident and emails `STATUS_ALERT_EMAIL` |
| 61 | (hey@flagon.io) with a link. A draft is not on the status page; the |
| 62 | part's own state already is, from the checks. |
| 63 | |
| 64 | Open it from the **Drafts** tab (the sidebar's Incidents count includes |
| 65 | drafts) and either: |
| 66 | |
| 67 | - **Publish to the status page**, with a public title (the "Detected:" |
| 68 | prefix is dropped) and a first update; or |
| 69 | - **Dismiss draft** with a reason, if it was a blip. Dismissed drafts are |
| 70 | listed under Resolved and never appear publicly. |
| 71 | |
| 72 | While an incident is open on a part, more failures on it add a line to |
| 73 | that incident's timeline instead of a new draft, and the part answering |
| 74 | again adds a "answering again" line. |
| 75 | |
| 76 | ## Running it |
| 77 | |
| 78 | The incident page has the timers at the top (open for, to acknowledge, |
| 79 | to mitigate, to resolve) and two columns: the timeline on the left, and |
| 80 | resolve, parts, roles, follow-ups and links on the right. |
| 81 | |
| 82 | **Post an update** (left): |
| 83 | |
| 84 | - **Public update**: shown on the status page and in the feeds; emailed |
| 85 | to subscribers when the box is ticked (ticked by default for SEV1 and |
| 86 | SEV2). Changing the status always needs a public update, because the |
| 87 | status page shows each status with words. |
| 88 | - **Internal note**: staff only. Use it for findings, links to logs, |
| 89 | who is doing what. Notes never reach the status page, the feeds or |
| 90 | email, but the postmortem's timeline starts from them, so keep them |
| 91 | factual. |
| 92 | - **Status**: Investigating → Identified (cause known) → Monitoring (fix |
| 93 | out, watching) → Resolved. Reaching Monitoring marks it mitigated. |
| 94 | - **Severity** and **Change the parts' impact**: change them as you learn |
| 95 | more; each change is a line on the timeline. |
| 96 | |
| 97 | The timeline shows everything, newest first: public updates highlighted |
| 98 | in lavender with the number of subscribers emailed, notes in gray, and |
| 99 | one-line entries for every status, severity, impact, role and follow-up |
| 100 | change. |
| 101 | |
| 102 | **Follow-ups**: add anything that should change so it does not happen |
| 103 | again, with an owner. Tick them off as they are done; they fill the |
| 104 | postmortem's action items. |
| 105 | |
| 106 | ## Resolving |
| 107 | |
| 108 | **Resolve** (right column) with one or two sentences: what is fixed and, |
| 109 | if known, what fixed it. It posts the last public update, sets the |
| 110 | resolved time, and the parts stop showing the incident's impact. If it |
| 111 | comes back, post an update with an earlier status: that reopens it. |
| 112 | |
| 113 | ## Postmortem |
| 114 | |
| 115 | From a resolved incident, **Write the postmortem**. The editor starts |
| 116 | with: |
| 117 | |
| 118 | - **Impact**: the parts, how badly and for how long; |
| 119 | - **Timeline**: every timeline line in UTC, internal notes included; |
| 120 | - **Action items**: the follow-ups. |
| 121 | |
| 122 | **Edit out anything internal** (hostnames, customer names, people's |
| 123 | names in blame) before publishing. Write the **Summary** and **Root |
| 124 | cause** (both required to publish), and what went well and badly. A blank |
| 125 | line starts a paragraph; lines starting with `- ` become a list. |
| 126 | |
| 127 | **Save draft** shows the preview beside the editor. **Save and publish** |
| 128 | (or **Publish the saved draft**) puts it on the incident's public page, |
| 129 | `https://status.g1t.sh/incidents/<id>#postmortem`, and marks the incident |
| 130 | "Postmortem published". **Take down** removes it again. |
| 131 | |
| 132 | Blameless: say what the system allowed, not who slipped. |
| 133 | |
| 134 | ## Scheduled maintenance |
| 135 | |
| 136 | **Schedule maintenance** on the Incidents page: |
| 137 | |
| 138 | 1. Title, window (UTC, at most 72 hours, at most 180 days ahead), the |
| 139 | parts affected, and a message saying what people will notice. |
| 140 | 2. **Email subscribers** emails them now, when it starts and when it ends. |
| 141 | |
| 142 | It is listed as upcoming at once. The minutely job starts it when the |
| 143 | window opens (its parts show "Under maintenance" and their checks stop |
| 144 | counting against uptime) and completes it when the window ends. From its |
| 145 | page you can post updates, **Start now**, **Complete** early, or |
| 146 | **Cancel**. |
| 147 | |
| 148 | If the work breaks something beyond what was announced, declare an |
| 149 | incident: an incident's impact shows over maintenance. |
| 150 | |
| 151 | ## Configuration |
| 152 | |
| 153 | The status worker (`apps/status/wrangler.jsonc`): |
| 154 | |
| 155 | | Setting | What it does | |
| 156 | | --- | --- | |
| 157 | | `send_email` binding `EMAIL` | Cloudflare Email Sending, as identity uses. Without it nothing is emailed; the feeds still work. | |
| 158 | | `STATUS_SECRET` (secret) | Signs unsubscribe links. Without it, the subscribe form is off. `npx wrangler secret put STATUS_SECRET` | |
| 159 | | `STATUS_ALERT_EMAIL` | Who hears about detected drafts. Empty sends none. | |
| 160 | | `STATUS_URL`, `SUDO_URL` | Links in email and feeds, and the alert's link to sudo. | |
| 161 | | `STATUS_FROM` | The From address. | |
| 162 | |
| 163 | sudo reaches the status worker only through its `STATUS` service binding |
| 164 | (`StatusAdmin` entrypoint); status.g1t.sh itself has no way to write. |