Vikrant Singh

· Planned game day, not a real outage

Postmortem: game day 1

Summary

On 1 October I broke this site's scheduled job on purpose, to test the incident desk end to end. The site noticed after two failed runs, raised an incident by itself, and an AI triaged it while every decision after that stayed with me. Visitors saw nothing. I found eight problems; six are fixed.

Fault
The scheduled job's call to GitHub forced to fail (HTTP 503)
User impact
None. Every page served throughout
Time to detect
17 min (two failed runs, by design)
Time to restore
27 min
Error budget used
Scheduled job: 2 of 86 failures allowed this month (2.3%). Outside probe: 0 failed checks
Incident record
GitHub issue #2

Timeline

Toronto time (EDT), 1 October.

1:10 a.m.Last healthy run before the fault
1:13Fault deployed through the normal pipeline
1:20Run 1 fails. Counted, no incident: one blip is not an incident
1:30Run 2 fails. The site raises the alert on its own
1:30The alert lands on an open test incident instead of a new one (finding 1)
1:33Restore deployed
1:40For about 15 seconds the snapshot is exactly 30 minutes old and the page says "Heartbeat late"; no probe lands in the window (finding 3)
1:40Healthy run. The site reports recovery; the engine notes that resolving is a human decision
1:47My out-of-order /resolve is refused by the lifecycle, but the issue gets closed by the button (finding 5)
2:39A command with a leading space is skipped without a word (finding 6)

What the outside view and the site's own records each showed

Three views watched the same failure, and they disagreed.

The site's own records saw it first. Scheduled job success dropped to 98.71%, two failed runs out of 78, logged within seconds.

The outside probe stayed green the whole time. It asks one question: is the status snapshot fresh? The last good snapshot was still young enough, so the answer was yes. That is correct. Visitors were fine.

The dashboard's headline believed the probe. For about twenty minutes it said "Operating normally" while the job behind it was failing.

None of these were wrong. They answer different questions: are users affected, and is the system healthy? The headline only asked the first. It now asks both, and says "Degraded" when the last run failed, even if visitors are still fine.

What went well

What hurt

My own test incident swallowed the real alert. An open "manual test" ticket for the same site was treated as the same problem, so the game day's alert landed as a comment on a test. Deduplication now matches the monitor as well as the service.

The timing was luck, not design. Two failed runs, ten minutes apart, put the snapshot at exactly thirty minutes old, the limit. For fifteen seconds the page said "Heartbeat late". The outside probe checks every five minutes, so it missed that window. Had it landed there, I would have been paged for a failure that was already recovering. The limit is now 35 minutes.

The gate held, but the interface around it didn't. I tried to resolve the incident before acknowledging it. The lifecycle refused, which is what it is for. But the same click closed the GitHub issue, so GitHub said "closed" while the incident record said "triaging". Closing an incident issue now reopens it until the record agrees. Then a command with one leading space was skipped without a word. A strict gate is fine. A silent one isn't. Commands now tolerate whitespace, and refusals say what to do next.

Two things are still open. Alerts go to a mailbox that is almost full, and a full mailbox is a silent pager. And this game day never tested paging at all: detection came from the site itself. The next one will run long enough to page me on purpose.

Findings and actions

#FindingActionStatus
1An open test incident swallowed the real alertDeduplicate on monitor and serviceFixed
2The headline said "Operating normally" while the job failedShow "Degraded: last scheduled run failed"Fixed
3Two failed runs land exactly on the 30-minute staleness limitLimit raised to 35 minutesFixed
4The incident engine didn't know this site: no owner, no blast radius, no runbook hintSite and its dependencies added to its CMDBFixed
5Closing the issue with the button bypassed the incident recordReopen until the record agrees; refusals name the next stepFixed
6A leading space made a command disappear silentlyCommands tolerate whitespaceFixed
7The alert mailbox is nearly full: a silent pagerClear space or add a second channelOpen (me)
8This game day never exercised pagingNext game day runs long enough to page on purposeOpen

Live reliability · The incident on GitHub · How this site is run