Pick any line to see why it is the way it is: the commit, the pull request and issue it came from, and what the agent was thinking.
| status.g1t.sh with incident management, invites that land you in the workspace, settings as pages, usage without quotas | 1 | # Incident runbook |
| 2 | ||
| 3 | For g1t staff. How to declare, run, resolve and write up an incident, and | |
| 4 | how to schedule maintenance, in sudo (**Platform → Incidents**, | |
| 5 | `https://sudo.g1t.sh/incidents`). Everything here is recorded in sudo's | |
| 6 | **Audit log** with your email. The public side is described in | |
| 7 | `apps/docs/src/content/docs/guides/status.md`. | |
| 8 | ||
| 9 | ## The short version | |
| 10 | ||
| 11 | 1. **Declare** as soon as people are affected. A wrong first guess is | |
| 12 | fine; silence is not. | |
| 13 | 2. **Update** publicly on the cadence the severity asks for, even when | |
| 14 | nothing changed ("Still working on it; next update by 15:30 UTC"). | |
| 15 | 3. **Resolve** when it is fixed, with one sentence on what fixed it. | |
| 16 | 4. **Write the postmortem** within five working days for SEV1 and SEV2. | |
| 17 | ||
| 18 | ## Severity | |
| 19 | ||
| 20 | | Severity | When | Public updates | Email subscribers | Postmortem | | |
| 21 | | --- | --- | --- | --- | --- | | |
| 22 | | SEV1 | g1t is down or unusable for most people, or data is at risk | Every 30 minutes at least | Yes | Required | | |
| 23 | | SEV2 | A core part (sign-in, git, the API, agents) is broken or badly degraded for many people | Hourly at least | Yes | Required | | |
| 24 | | SEV3 | One part is degraded or broken for some people; there is a way around it | As things change | Your call | If there is a lesson | | |
| 25 | | SEV4 | Little or no customer impact | As things change | No | No | | |
| 26 | ||
| 27 | When unsure, pick the higher one. Lowering it later is one change. | |
| 28 | ||
| 29 | ## Roles | |
| 30 | ||
| 31 | - **Incident commander (IC)**: runs the response. Decides, assigns work, | |
| 32 | keeps the timeline honest. Not necessarily the person fixing it. | |
| 33 | - **Communications**: writes the public updates and keeps the cadence. | |
| 34 | For a small incident the IC does both. | |
| 35 | ||
| 36 | Set both on the declare form or in **Roles** on the incident page. A | |
| 37 | change of role is written on the timeline. | |
| 38 | ||
| 39 | ## Declaring | |
| 40 | ||
| 41 | **Declare incident** on the Incidents page: | |
| 42 | ||
| 43 | 1. **Title**: what people notice, not the internals. "Pushes over HTTPS | |
| 44 | failing", not "pack-receiver 502s on eu-2". | |
| 45 | 2. **Severity**: see the table. | |
| 46 | 3. **Status**: usually **Investigating**. | |
| 47 | 4. **Impact began (UTC)**: leave empty for now, or backdate to when it | |
| 48 | really started. The time-to-acknowledge, mitigate and resolve figures | |
| 49 | count from here. | |
| 50 | 5. **Parts affected**: set each affected part to Degraded performance, | |
| 51 | Partial outage or Major outage. The status page shows that (or worse, | |
| 52 | if the checks say worse) until the incident is resolved. | |
| 53 | 6. **First public update**: what is broken, for whom, and what still works. | |
| 54 | 7. **Email subscribers**: "As the severity says" emails for SEV1 and SEV2. | |
| 55 | 8. **Declare and publish**. It is on status.g1t.sh within 30 seconds. | |
| 56 | ||
| 57 | ## Detected drafts | |
| 58 | ||
| 59 | When a part fails or is slow three checks in a row (three minutes), the | |
| 60 | status worker makes a **draft** incident and emails `STATUS_ALERT_EMAIL` | |
| 61 | (hey@flagon.io) with a link. A draft is not on the status page; the | |
| 62 | part's own state already is, from the checks. | |
| 63 | ||
| 64 | Open it from the **Drafts** tab (the sidebar's Incidents count includes | |
| 65 | drafts) and either: | |
| 66 | ||
| 67 | - **Publish to the status page**, with a public title (the "Detected:" | |
| 68 | prefix is dropped) and a first update; or | |
| 69 | - **Dismiss draft** with a reason, if it was a blip. Dismissed drafts are | |
| 70 | listed under Resolved and never appear publicly. | |
| 71 | ||
| 72 | While an incident is open on a part, more failures on it add a line to | |
| 73 | that incident's timeline instead of a new draft, and the part answering | |
| 74 | again adds a "answering again" line. | |
| 75 | ||
| 76 | ## Running it | |
| 77 | ||
| 78 | The incident page has the timers at the top (open for, to acknowledge, | |
| 79 | to mitigate, to resolve) and two columns: the timeline on the left, and | |
| 80 | resolve, parts, roles, follow-ups and links on the right. | |
| 81 | ||
| 82 | **Post an update** (left): | |
| 83 | ||
| 84 | - **Public update**: shown on the status page and in the feeds; emailed | |
| 85 | to subscribers when the box is ticked (ticked by default for SEV1 and | |
| 86 | SEV2). Changing the status always needs a public update, because the | |
| 87 | status page shows each status with words. | |
| 88 | - **Internal note**: staff only. Use it for findings, links to logs, | |
| 89 | who is doing what. Notes never reach the status page, the feeds or | |
| 90 | email, but the postmortem's timeline starts from them, so keep them | |
| 91 | factual. | |
| 92 | - **Status**: Investigating → Identified (cause known) → Monitoring (fix | |
| 93 | out, watching) → Resolved. Reaching Monitoring marks it mitigated. | |
| 94 | - **Severity** and **Change the parts' impact**: change them as you learn | |
| 95 | more; each change is a line on the timeline. | |
| 96 | ||
| 97 | The timeline shows everything, newest first: public updates highlighted | |
| 98 | in lavender with the number of subscribers emailed, notes in gray, and | |
| 99 | one-line entries for every status, severity, impact, role and follow-up | |
| 100 | change. | |
| 101 | ||
| 102 | **Follow-ups**: add anything that should change so it does not happen | |
| 103 | again, with an owner. Tick them off as they are done; they fill the | |
| 104 | postmortem's action items. | |
| 105 | ||
| 106 | ## Resolving | |
| 107 | ||
| 108 | **Resolve** (right column) with one or two sentences: what is fixed and, | |
| 109 | if known, what fixed it. It posts the last public update, sets the | |
| 110 | resolved time, and the parts stop showing the incident's impact. If it | |
| 111 | comes back, post an update with an earlier status: that reopens it. | |
| 112 | ||
| 113 | ## Postmortem | |
| 114 | ||
| 115 | From a resolved incident, **Write the postmortem**. The editor starts | |
| 116 | with: | |
| 117 | ||
| 118 | - **Impact**: the parts, how badly and for how long; | |
| 119 | - **Timeline**: every timeline line in UTC, internal notes included; | |
| 120 | - **Action items**: the follow-ups. | |
| 121 | ||
| 122 | **Edit out anything internal** (hostnames, customer names, people's | |
| 123 | names in blame) before publishing. Write the **Summary** and **Root | |
| 124 | cause** (both required to publish), and what went well and badly. A blank | |
| 125 | line starts a paragraph; lines starting with `- ` become a list. | |
| 126 | ||
| 127 | **Save draft** shows the preview beside the editor. **Save and publish** | |
| 128 | (or **Publish the saved draft**) puts it on the incident's public page, | |
| 129 | `https://status.g1t.sh/incidents/<id>#postmortem`, and marks the incident | |
| 130 | "Postmortem published". **Take down** removes it again. | |
| 131 | ||
| 132 | Blameless: say what the system allowed, not who slipped. | |
| 133 | ||
| 134 | ## Scheduled maintenance | |
| 135 | ||
| 136 | **Schedule maintenance** on the Incidents page: | |
| 137 | ||
| 138 | 1. Title, window (UTC, at most 72 hours, at most 180 days ahead), the | |
| 139 | parts affected, and a message saying what people will notice. | |
| 140 | 2. **Email subscribers** emails them now, when it starts and when it ends. | |
| 141 | ||
| 142 | It is listed as upcoming at once. The minutely job starts it when the | |
| 143 | window opens (its parts show "Under maintenance" and their checks stop | |
| 144 | counting against uptime) and completes it when the window ends. From its | |
| 145 | page you can post updates, **Start now**, **Complete** early, or | |
| 146 | **Cancel**. | |
| 147 | ||
| 148 | If the work breaks something beyond what was announced, declare an | |
| 149 | incident: an incident's impact shows over maintenance. | |
| 150 | ||
| 151 | ## Configuration | |
| 152 | ||
| 153 | The status worker (`apps/status/wrangler.jsonc`): | |
| 154 | ||
| 155 | | Setting | What it does | | |
| 156 | | --- | --- | | |
| 157 | | `send_email` binding `EMAIL` | Cloudflare Email Sending, as identity uses. Without it nothing is emailed; the feeds still work. | | |
| 158 | | `STATUS_SECRET` (secret) | Signs unsubscribe links. Without it, the subscribe form is off. `npx wrangler secret put STATUS_SECRET` | | |
| 159 | | `STATUS_ALERT_EMAIL` | Who hears about detected drafts. Empty sends none. | | |
| 160 | | `STATUS_URL`, `SUDO_URL` | Links in email and feeds, and the alert's link to sudo. | | |
| 161 | | `STATUS_FROM` | The From address. | | |
| 162 | ||
| 163 | sudo reaches the status worker only through its `STATUS` service binding | |
| 164 | (`StatusAdmin` entrypoint); status.g1t.sh itself has no way to write. |