· Planned game day, not a real outage
Postmortem: game day 1
Summary
On 1 October I broke this site's scheduled job on purpose, to test the incident desk end to end. The site noticed after two failed runs, raised an incident by itself, and an AI triaged it while every decision after that stayed with me. Visitors saw nothing. I found eight problems; six are fixed.
- Fault
- The scheduled job's call to GitHub forced to fail (HTTP 503)
- User impact
- None. Every page served throughout
- Time to detect
- 17 min (two failed runs, by design)
- Time to restore
- 27 min
- Error budget used
- Scheduled job: 2 of 86 failures allowed this month (2.3%). Outside probe: 0 failed checks
- Incident record
- GitHub issue #2
Timeline
Toronto time (EDT), 1 October.
| 1:10 a.m. | Last healthy run before the fault |
|---|---|
| 1:13 | Fault deployed through the normal pipeline |
| 1:20 | Run 1 fails. Counted, no incident: one blip is not an incident |
| 1:30 | Run 2 fails. The site raises the alert on its own |
| 1:30 | The alert lands on an open test incident instead of a new one (finding 1) |
| 1:33 | Restore deployed |
| 1:40 | For about 15 seconds the snapshot is exactly 30 minutes old and the page says "Heartbeat late"; no probe lands in the window (finding 3) |
| 1:40 | Healthy run. The site reports recovery; the engine notes that resolving is a human decision |
| 1:47 | My out-of-order /resolve is refused by the lifecycle, but the issue gets closed by the button (finding 5) |
| 2:39 | A command with a leading space is skipped without a word (finding 6) |
What the outside view and the site's own records each showed
Three views watched the same failure, and they disagreed.
The site's own records saw it first. Scheduled job success dropped to 98.71%, two failed runs out of 78, logged within seconds.
The outside probe stayed green the whole time. It asks one question: is the status snapshot fresh? The last good snapshot was still young enough, so the answer was yes. That is correct. Visitors were fine.
The dashboard's headline believed the probe. For about twenty minutes it said "Operating normally" while the job behind it was failing.
None of these were wrong. They answer different questions: are users affected, and is the system healthy? The headline only asked the first. It now asks both, and says "Degraded" when the last run failed, even if visitors are still fine.
What went well
- The site raised the alert by itself; the incident engine ran with no human involved.
- One failed run did not raise an incident. The anti-flap rule held.
- The AI triaged it, drafted updates marked "never sent", and flagged two real gaps in its own setup.
- The human gate held: an out-of-order resolve was refused.
- Recovery was reported automatically; resolving stayed a human decision.
What hurt
My own test incident swallowed the real alert. An open "manual test" ticket for the same site was treated as the same problem, so the game day's alert landed as a comment on a test. Deduplication now matches the monitor as well as the service.
The timing was luck, not design. Two failed runs, ten minutes apart, put the snapshot at exactly thirty minutes old, the limit. For fifteen seconds the page said "Heartbeat late". The outside probe checks every five minutes, so it missed that window. Had it landed there, I would have been paged for a failure that was already recovering. The limit is now 35 minutes.
The gate held, but the interface around it didn't. I tried to resolve the incident before acknowledging it. The lifecycle refused, which is what it is for. But the same click closed the GitHub issue, so GitHub said "closed" while the incident record said "triaging". Closing an incident issue now reopens it until the record agrees. Then a command with one leading space was skipped without a word. A strict gate is fine. A silent one isn't. Commands now tolerate whitespace, and refusals say what to do next.
Two things are still open. Alerts go to a mailbox that is almost full, and a full mailbox is a silent pager. And this game day never tested paging at all: detection came from the site itself. The next one will run long enough to page me on purpose.
Findings and actions
| # | Finding | Action | Status |
|---|---|---|---|
| 1 | An open test incident swallowed the real alert | Deduplicate on monitor and service | Fixed |
| 2 | The headline said "Operating normally" while the job failed | Show "Degraded: last scheduled run failed" | Fixed |
| 3 | Two failed runs land exactly on the 30-minute staleness limit | Limit raised to 35 minutes | Fixed |
| 4 | The incident engine didn't know this site: no owner, no blast radius, no runbook hint | Site and its dependencies added to its CMDB | Fixed |
| 5 | Closing the issue with the button bypassed the incident record | Reopen until the record agrees; refusals name the next step | Fixed |
| 6 | A leading space made a command disappear silently | Commands tolerate whitespace | Fixed |
| 7 | The alert mailbox is nearly full: a silent pager | Clear space or add a second channel | Open (me) |
| 8 | This game day never exercised paging | Next game day runs long enough to page on purpose | Open |
Live reliability · The incident on GitHub · How this site is run