Skip to content
194 linesCodeBlameRaw
1# Incident runbook
2
3For g1t staff. How to declare, run, resolve and write up an incident, and
4how to schedule maintenance, in sudo (**Platform → Incidents**,
5`https://sudo.g1t.sh/incidents`). Everything here is recorded in sudo's
6**Audit log** with your email. The public side is described in
7`apps/docs/src/content/docs/guides/status.md`.
8
9## The short version
10
111. **Declare** as soon as people are affected. A wrong first guess is
12 fine; silence is not.
132. **Update** publicly on the cadence the severity asks for, even when
14 nothing changed ("Still working on it; next update by 15:30 UTC").
153. **Resolve** when it is fixed, with one sentence on what fixed it.
164. **Write the postmortem** within five working days for SEV1 and SEV2.
17
18## Severity
19
20| Severity | When | Public updates | Email subscribers | Postmortem |
21| --- | --- | --- | --- | --- |
22| SEV1 | g1t is down or unusable for most people, or data is at risk | Every 30 minutes at least | Yes | Required |
23| SEV2 | A core part (sign-in, git, the API, agents) is broken or badly degraded for many people | Hourly at least | Yes | Required |
24| SEV3 | One part is degraded or broken for some people; there is a way around it | As things change | Your call | If there is a lesson |
25| SEV4 | Little or no customer impact | As things change | No | No |
26
27When unsure, pick the higher one. Lowering it later is one change.
28
29## Roles
30
31- **Incident commander (IC)**: runs the response. Decides, assigns work,
32 keeps the timeline honest. Not necessarily the person fixing it.
33- **Communications**: writes the public updates and keeps the cadence.
34 For a small incident the IC does both.
35
36Set both on the declare form or in **Roles** on the incident page. A
37change of role is written on the timeline.
38
39## Declaring
40
41**Declare incident** on the Incidents page:
42
431. **Title**: what people notice, not the internals. "Pushes over HTTPS
44 failing", not "pack-receiver 502s on eu-2".
452. **Severity**: see the table.
463. **Status**: usually **Investigating**.
474. **Impact began (UTC)**: leave empty for now, or backdate to when it
48 really started. The time-to-acknowledge, mitigate and resolve figures
49 count from here.
505. **Parts affected**: set each affected part to Degraded performance,
51 Partial outage or Major outage. The status page shows that (or worse,
52 if the checks say worse) until the incident is resolved.
536. **First public update**: what is broken, for whom, and what still works.
547. **Email subscribers**: "As the severity says" emails for SEV1 and SEV2.
558. **Declare and publish**. It is on status.g1t.sh within 30 seconds.
56
57## Detected drafts
58
59When a part fails or is slow on four of its last five checks (a check a
60minute), the status worker makes a **draft** incident and emails
61`STATUS_ALERT_EMAIL` (hey@flagon.io) with a link. A draft is not on the
62status page; the part's own state already is, from the checks.
63
64How the checks decide (`apps/status/src/detect.ts`, `probe.ts`):
65
66- **A slow answer is asked again at once.** It counts as slow only if the
67 second answer is slow too, so one cold start or cache refill is not a
68 slow check. A check that fails (an error or a timeout) is not asked
69 again; four of five decides.
70- **Four of five, not three in a row.** A good check in between does not
71 hide trouble that keeps coming; three good checks in a row end a run of
72 trouble, however long it lasted, and the next trouble starts a new one
73 with its own start. A draft's lines say "on 4 checks in a row" or "on 4
74 of 5 checks".
75- **Page loads are timed to the first byte.** The site's pages are loaded
76 with a browser's user agent (ending `g1t-status/1.0 (+status.g1t.sh)`),
77 because the site renders the whole page first for a crawler. See
78 `docs/PERFORMANCE.md`.
79- **Every check is kept for 7 days**: how long it took, what it meant, and
80 the Cloudflare data centre it ran from (from the answer's `cf-ray`). The
81 incident page shows a **Checks** chart per part, from 30 minutes before
82 the impact began to 30 minutes after it ended (at most a day), with the
83 slow line, slow checks dotted and failures marked, and the latest 30
84 checks as a table.
85- **A draft nobody acknowledges is raised again**: the alert goes out once
86 more after 45 minutes, then every 6 hours while it waits, with a note on
87 its timeline each time. Publishing it, posting a note or dismissing it
88 acknowledges it and stops the reminders.
89- **Deploys are announced.** `scripts/deploy.mjs` tells the status worker
90 when a deploy starts and finishes (`docs/DEPLOYING.md`); during one, and
91 for 3 minutes after, trouble is counted but not drafted unless it
92 outlasts the deploy.
93
94Open it from the **Drafts** tab (the sidebar's Incidents count includes
95drafts) and either:
96
97- **Publish to the status page**, with a public title (the "Detected:"
98 prefix is dropped) and a first update; or
99- **Dismiss draft** with a reason, if it was a blip. Dismissed drafts are
100 listed under Resolved and never appear publicly.
101
102While an incident is open on a part, more failures on it add a line to
103that incident's timeline instead of a new draft, and the part answering
104again (three good checks in a row) adds an "answering again" line.
105
106## Running it
107
108The incident page has the timers at the top (open for, to acknowledge,
109to mitigate, to resolve) and two columns: the timeline on the left, and
110resolve, parts, roles, follow-ups and links on the right.
111
112**Post an update** (left):
113
114- **Public update**: shown on the status page and in the feeds; emailed
115 to subscribers when the box is ticked (ticked by default for SEV1 and
116 SEV2). Changing the status always needs a public update, because the
117 status page shows each status with words.
118- **Internal note**: staff only. Use it for findings, links to logs,
119 who is doing what. Notes never reach the status page, the feeds or
120 email, but the postmortem's timeline starts from them, so keep them
121 factual.
122- **Status**: Investigating → Identified (cause known) → Monitoring (fix
123 out, watching) → Resolved. Reaching Monitoring marks it mitigated.
124- **Severity** and **Change the parts' impact**: change them as you learn
125 more; each change is a line on the timeline.
126
127The timeline shows everything, newest first: public updates highlighted
128in lavender with the number of subscribers emailed, notes in gray, and
129one-line entries for every status, severity, impact, role and follow-up
130change.
131
132**Follow-ups**: add anything that should change so it does not happen
133again, with an owner. Tick them off as they are done; they fill the
134postmortem's action items.
135
136## Resolving
137
138**Resolve** (right column) with one or two sentences: what is fixed and,
139if known, what fixed it. It posts the last public update, sets the
140resolved time, and the parts stop showing the incident's impact. If it
141comes back, post an update with an earlier status: that reopens it.
142
143## Postmortem
144
145From a resolved incident, **Write the postmortem**. The editor starts
146with:
147
148- **Impact**: the parts, how badly and for how long;
149- **Timeline**: every timeline line in UTC, internal notes included;
150- **Action items**: the follow-ups.
151
152**Edit out anything internal** (hostnames, customer names, people's
153names in blame) before publishing. Write the **Summary** and **Root
154cause** (both required to publish), and what went well and badly. A blank
155line starts a paragraph; lines starting with `- ` become a list.
156
157**Save draft** shows the preview beside the editor. **Save and publish**
158(or **Publish the saved draft**) puts it on the incident's public page,
159`https://status.g1t.sh/incidents/<id>#postmortem`, and marks the incident
160"Postmortem published". **Take down** removes it again.
161
162Blameless: say what the system allowed, not who slipped.
163
164## Scheduled maintenance
165
166**Schedule maintenance** on the Incidents page:
167
1681. Title, window (UTC, at most 72 hours, at most 180 days ahead), the
169 parts affected, and a message saying what people will notice.
1702. **Email subscribers** emails them now, when it starts and when it ends.
171
172It is listed as upcoming at once. The minutely job starts it when the
173window opens (its parts show "Under maintenance" and their checks stop
174counting against uptime) and completes it when the window ends. From its
175page you can post updates, **Start now**, **Complete** early, or
176**Cancel**.
177
178If the work breaks something beyond what was announced, declare an
179incident: an incident's impact shows over maintenance.
180
181## Configuration
182
183The status worker (`apps/status/wrangler.jsonc`):
184
185| Setting | What it does |
186| --- | --- |
187| `send_email` binding `EMAIL` | Cloudflare Email Sending, as identity uses. Without it nothing is emailed; the feeds still work. |
188| `STATUS_SECRET` (secret) | Signs unsubscribe links. Without it, the subscribe form is off. `npx wrangler secret put STATUS_SECRET` |
189| `STATUS_ALERT_EMAIL` | Who hears about detected drafts. Empty sends none. |
190| `STATUS_URL`, `SUDO_URL` | Links in email and feeds, and the alert's link to sudo. |
191| `STATUS_FROM` | The From address. |
192
193sudo reaches the status worker only through its `STATUS` service binding
194(`StatusAdmin` entrypoint); status.g1t.sh itself has no way to write.