Vikrant Singh

· Display bug, no data lost

The failure list that lost its history

The reliability page stopped showing two game days and an outage. Nothing was lost. One read had a cap, and the wrong records filled it.

Summary

On 4 October I noticed the Failures list on Live reliability showed two entries, both from 3 October. The game days of 1 and 2 October, which used to be there, were gone. The 30-day strip right above it still showed those days in red and amber, so the page disagreed with itself.

No record was lost. The site keeps every detailed record for 400 days. The page asked for the newest 1,000 of them across every source at once, and the multiplayer game had produced most of the recent ones: 590 slow-turn records and 382 round-trip timings. Everything older than midday on 2 October fell outside the 1,000, including every game-day failure. The fix caps each source on its own, and a new page lists the full history.

Found
By eye: the list and the strip above it disagreed
Cause
One cap of 1,000 records shared by every source, filled by the game's records
Data lost
None. Daily totals and every detailed record were intact
User impact
The failure history was hidden from the page; the objectives and budgets were right throughout
Error budget used
None. Nothing failed; a list was incomplete
Fixed
4 October, 6:15 p.m. Toronto time: all 12 failures of the window back on the page

What it looked like

The Failures list on 3 October, showing ten entries from 1 and 2 October: probe checks and scheduled jobs that failed, most tagged game day.
3 October: the list as it was, with both game days tagged.
The 30-day strip on 4 October showing red and amber days, outlined, above a Failures list, outlined, that shows only two entries from 3 October.
4 October, before the fix: the strip shows failures on three days; the list below shows two entries.
The same strip after the fix, above a Failures list showing the latest five failures, game days tagged, and the line: Showing the latest 5 of 12 in the last 30 days. Full failure history.
After: the latest five, a count of the rest, and a link to all of them.

Why it happened

The page read its records with one rule: newest first, at most 1,000. That made sense when the site's own checks were the only thing writing detailed records. Then the multiplayer game started keeping records too, and a day of play produces hundreds. The cap didn't belong to any one kind of record, so the busiest one used all of it. The list had no way to know it was incomplete, so it didn't say so.

The fix

The new Failure history page: 12 failures on 3 days since 1 October 2026, grouped under 3, 2 and 1 October, with game-day failures tagged.
The full history, grouped by day.

Findings and actions

#FindingActionStatus
1One cap was shared by every source, so the busiest took it allA cap per sourceFixed
2A shortened list didn't say it was shortenedThe list says "latest 5 of 12" and links to the restFixed
3There was no way to see older failures on the siteThe failure history pageFixed
4The page's own checks test how records are drawn, not how they are read from storageA check that runs the storage read itselfOpen

What this is not

It wasn't an outage or an incident. Pages served, the objectives and budgets were right, and nothing paged. It was a page that quietly showed less than it had.

The rule that came out of it

A list that is cut short says so, and says where the rest is.

Failure history · Live reliability · Postmortem: game day 1 · Postmortem: the day's database writes