flagon-io/g1t

public

Where people and agents ship software together. The open-source git platform for the whole job: issues, agents, checks and deploys to the edge.

g1t/docs/INCIDENTS.md

164 lines7,360 bytesCodeBlame
1# Incident runbook
2
3For g1t staff. How to declare, run, resolve and write up an incident, and
4how to schedule maintenance, in sudo (**Platform → Incidents**,
5`https://sudo.g1t.sh/incidents`). Everything here is recorded in sudo's
6**Audit log** with your email. The public side is described in
7`apps/docs/src/content/docs/guides/status.md`.
8
9## The short version
10
111. **Declare** as soon as people are affected. A wrong first guess is
12 fine; silence is not.
132. **Update** publicly on the cadence the severity asks for, even when
14 nothing changed ("Still working on it; next update by 15:30 UTC").
153. **Resolve** when it is fixed, with one sentence on what fixed it.
164. **Write the postmortem** within five working days for SEV1 and SEV2.
17
18## Severity
19
20| Severity | When | Public updates | Email subscribers | Postmortem |
21| --- | --- | --- | --- | --- |
22| SEV1 | g1t is down or unusable for most people, or data is at risk | Every 30 minutes at least | Yes | Required |
23| SEV2 | A core part (sign-in, git, the API, agents) is broken or badly degraded for many people | Hourly at least | Yes | Required |
24| SEV3 | One part is degraded or broken for some people; there is a way around it | As things change | Your call | If there is a lesson |
25| SEV4 | Little or no customer impact | As things change | No | No |
26
27When unsure, pick the higher one. Lowering it later is one change.
28
29## Roles
30
31- **Incident commander (IC)**: runs the response. Decides, assigns work,
32 keeps the timeline honest. Not necessarily the person fixing it.
33- **Communications**: writes the public updates and keeps the cadence.
34 For a small incident the IC does both.
35
36Set both on the declare form or in **Roles** on the incident page. A
37change of role is written on the timeline.
38
39## Declaring
40
41**Declare incident** on the Incidents page:
42
431. **Title**: what people notice, not the internals. "Pushes over HTTPS
44 failing", not "pack-receiver 502s on eu-2".
452. **Severity**: see the table.
463. **Status**: usually **Investigating**.
474. **Impact began (UTC)**: leave empty for now, or backdate to when it
48 really started. The time-to-acknowledge, mitigate and resolve figures
49 count from here.
505. **Parts affected**: set each affected part to Degraded performance,
51 Partial outage or Major outage. The status page shows that (or worse,
52 if the checks say worse) until the incident is resolved.
536. **First public update**: what is broken, for whom, and what still works.
547. **Email subscribers**: "As the severity says" emails for SEV1 and SEV2.
558. **Declare and publish**. It is on status.g1t.sh within 30 seconds.
56
57## Detected drafts
58
59When a part fails or is slow three checks in a row (three minutes), the
60status worker makes a **draft** incident and emails `STATUS_ALERT_EMAIL`
61(hey@flagon.io) with a link. A draft is not on the status page; the
62part's own state already is, from the checks.
63
64Open it from the **Drafts** tab (the sidebar's Incidents count includes
65drafts) and either:
66
67- **Publish to the status page**, with a public title (the "Detected:"
68 prefix is dropped) and a first update; or
69- **Dismiss draft** with a reason, if it was a blip. Dismissed drafts are
70 listed under Resolved and never appear publicly.
71
72While an incident is open on a part, more failures on it add a line to
73that incident's timeline instead of a new draft, and the part answering
74again adds a "answering again" line.
75
76## Running it
77
78The incident page has the timers at the top (open for, to acknowledge,
79to mitigate, to resolve) and two columns: the timeline on the left, and
80resolve, parts, roles, follow-ups and links on the right.
81
82**Post an update** (left):
83
84- **Public update**: shown on the status page and in the feeds; emailed
85 to subscribers when the box is ticked (ticked by default for SEV1 and
86 SEV2). Changing the status always needs a public update, because the
87 status page shows each status with words.
88- **Internal note**: staff only. Use it for findings, links to logs,
89 who is doing what. Notes never reach the status page, the feeds or
90 email, but the postmortem's timeline starts from them, so keep them
91 factual.
92- **Status**: Investigating → Identified (cause known) → Monitoring (fix
93 out, watching) → Resolved. Reaching Monitoring marks it mitigated.
94- **Severity** and **Change the parts' impact**: change them as you learn
95 more; each change is a line on the timeline.
96
97The timeline shows everything, newest first: public updates highlighted
98in lavender with the number of subscribers emailed, notes in gray, and
99one-line entries for every status, severity, impact, role and follow-up
100change.
101
102**Follow-ups**: add anything that should change so it does not happen
103again, with an owner. Tick them off as they are done; they fill the
104postmortem's action items.
105
106## Resolving
107
108**Resolve** (right column) with one or two sentences: what is fixed and,
109if known, what fixed it. It posts the last public update, sets the
110resolved time, and the parts stop showing the incident's impact. If it
111comes back, post an update with an earlier status: that reopens it.
112
113## Postmortem
114
115From a resolved incident, **Write the postmortem**. The editor starts
116with:
117
118- **Impact**: the parts, how badly and for how long;
119- **Timeline**: every timeline line in UTC, internal notes included;
120- **Action items**: the follow-ups.
121
122**Edit out anything internal** (hostnames, customer names, people's
123names in blame) before publishing. Write the **Summary** and **Root
124cause** (both required to publish), and what went well and badly. A blank
125line starts a paragraph; lines starting with `- ` become a list.
126
127**Save draft** shows the preview beside the editor. **Save and publish**
128(or **Publish the saved draft**) puts it on the incident's public page,
129`https://status.g1t.sh/incidents/<id>#postmortem`, and marks the incident
130"Postmortem published". **Take down** removes it again.
131
132Blameless: say what the system allowed, not who slipped.
133
134## Scheduled maintenance
135
136**Schedule maintenance** on the Incidents page:
137
1381. Title, window (UTC, at most 72 hours, at most 180 days ahead), the
139 parts affected, and a message saying what people will notice.
1402. **Email subscribers** emails them now, when it starts and when it ends.
141
142It is listed as upcoming at once. The minutely job starts it when the
143window opens (its parts show "Under maintenance" and their checks stop
144counting against uptime) and completes it when the window ends. From its
145page you can post updates, **Start now**, **Complete** early, or
146**Cancel**.
147
148If the work breaks something beyond what was announced, declare an
149incident: an incident's impact shows over maintenance.
150
151## Configuration
152
153The status worker (`apps/status/wrangler.jsonc`):
154
155| Setting | What it does |
156| --- | --- |
157| `send_email` binding `EMAIL` | Cloudflare Email Sending, as identity uses. Without it nothing is emailed; the feeds still work. |
158| `STATUS_SECRET` (secret) | Signs unsubscribe links. Without it, the subscribe form is off. `npx wrangler secret put STATUS_SECRET` |
159| `STATUS_ALERT_EMAIL` | Who hears about detected drafts. Empty sends none. |
160| `STATUS_URL`, `SUDO_URL` | Links in email and feeds, and the alert's link to sudo. |
161| `STATUS_FROM` | The From address. |
162
163sudo reaches the status worker only through its `STATUS` service binding
164(`StatusAdmin` entrypoint); status.g1t.sh itself has no way to write.