· Real incident, self-inflicted
Postmortem: a load test used up the day's database writes
Summary
I load-tested the multiplayer game on a separate test copy of this site. The test copy shares the account's daily allowance of database writes with the live site, and an hour later that allowance ran out. For six and a half hours the live reliability page couldn't show its numbers, new games couldn't start, and the site's own records missed their readings. Pages, the heartbeat and the outside monitor were fine, and nothing raised a false alarm.
The worst part wasn't the outage. It was the page saying "the live numbers appear here once 30 days of data exist", which looked like a choice rather than a failure. Six minutes after the cause was found, it said what was actually wrong. Within two hours, every place that depends on that storage explained itself and said when it would be back, the game wrote a third as much per move, and missed readings had a way back.
- Cause
- Two load tests and the runs after them used up the daily 100,000 database writes on Cloudflare's free tier, shared by the test and live copies
- User impact
- Live reliability numbers and new Ludo games unavailable for 6 h 24 min. Every page served
- Time to detect
- Minutes, by eye. No signal raised it (finding 4)
- Time to explain
- 6 minutes after the cause was confirmed: the page said what was wrong
- Time to restore
- The daily reset at 8:00 p.m.; the free tier has no earlier way back
- Error budget used
- None. The outside probe's checks all passed; the heartbeat never failed
- Data
- Missed readings rebuilt from logs after the reset; 140 of the day's 144 scheduled runs are on record
Timeline
Toronto time (EDT), 2 October.
| 12:21 p.m. | Load test on the test copy: 40 games of four players for a minute, about 11,600 moves |
|---|---|
| 12:30 | Second load test after a scaling change; then a round of regression games |
| 1:10 | The scaling change ships. Each move now writes about six database rows (finding 3) |
| 1:36 | The day's allowance runs out. Every database write on the account is refused |
| 1:36 | The reliability page falls back to "numbers appear once 30 days exist" (finding 5) |
| ~1:40 | I notice the page is wrong and ask why |
| 1:42 | The cause is in the logs word for word: "Exceeded allowed rows written … free tier" |
| 1:48 | The page now says the numbers are temporarily unavailable, and that the site itself is fine |
| 2:09 | Every affected place names the cause and the reset time; the game cuts its writes to about two per move |
| 2:58 | Rebuilding missed readings from logs becomes a standing rule; today's gap is queued |
| 3:35 | Checked: with writes refused, could the records at least be read? No. Even a bare read is refused (finding 8) |
| 3:45 | A read copy of the records now lives outside that storage; the page shows it, labelled "Recording paused since…" |
| 8:00 p.m. | The allowance resets; the rebuild runs, half an hour of readings at a time |
What it looked like
Screenshots taken at the time, on a phone and by a scheduled capture. Cropped, otherwise as they were.



What went well
- Degrade open held. Every page served, and the heartbeat and the outside monitor never noticed.
- No false alarm. Nothing that can raise an incident depends on this storage, by design.
- The cause was in the logs verbatim, so diagnosis took minutes, not guesses.
- Nothing was really lost. Every reading is also written to the logs, which live somewhere else.
What hurt
I checked the wrong limit. Before the second test I looked at the daily request allowance and had room. The one that ran out was rows written, which is smaller in practice. A test plan should name the resource it might exhaust, and this one named the wrong one.
The test copy wasn't as separate as it looked. It had its own address, its own storage and no access to the live data, but it drew on the same account-wide allowance. Isolation that stops at the account boundary isn't isolation for capacity.
The scaling change made this worse. It removed a real bottleneck, every move queuing on one shared ledger, and cut the slowest 1% of moves from 240 to 80 ms. But it did so by writing more rows per move, and on this plan rows written are the scarce resource. It was optimised for the wrong constraint. Each move now writes one record and one small log row, a third of before.
And the failure looked like a decision. "The live numbers appear here once 30 days of data exist" is a sensible message for a flag that's off and a misleading one for storage that's refusing writes. A page that can't say why it's empty will be read as hiding something.
Findings and actions
| # | Finding | Action | Status |
|---|---|---|---|
| 1 | The test plan checked request limits, not rows written | Capacity budget named in every test plan; no large load tests on the free tier | Done |
| 2 | Test and live copies share account-wide allowances | Separate account or a paid plan before the next load test | Open (me) |
| 3 | A scaling change raised writes per move from about 4 to 6 | One record plus one log row per move, about 2 | Fixed |
| 4 | Detected by eye; nothing signalled the refused writes | The heartbeat notices refused writes and queues a rebuild (never a page) | Fixed; a quiet notice is open |
| 5 | The empty page looked like a choice | Every affected place names the cause and the reset time | Fixed |
| 6 | The game failed as a dead connection | A "Ludo is napping" screen with the reset in your own time | Fixed |
| 7 | Six and a half hours of readings missed | Rebuilt from logs after any outage, automatically, back into each day's totals | Fixed; the rebuild ran on 3 Oct |
| 8 | The page went dark although only writes were refused; it turned out reads are refused too | A read copy of the 30-day records, refreshed every 10 minutes, shown with "Recording paused since…" when the store can't be read | Fixed |
| 9 | The site said rebuilt readings were "marked as backfilled", but only failed or slow ones keep the mark: good ones go into the day's totals unmarked. For 2 Oct, two marked readings exist | Every page now says what really happens. A visible count of rebuilt readings per day comes with the next change to how records are stored | Wording fixed; count open |
How it was worked
The fixes came from a handful of questions, asked in this order. Each one changed something.
- Why is the page empty? It looked like a choice. The logs said otherwise, word for word, within minutes.
- What exactly ran out, and who shares it? Rows written, not requests; and the test copy shares the live site's allowance. (Findings 1 and 2.)
- Does this page anyone? It must not: a budget running out isn't an incident. I checked that nothing on the incident path depends on this storage, and watched the heartbeat stay green throughout.
- What do people see meanwhile? The cause and when it's back, everywhere it shows, instead of a blank or a dead connection. (Findings 5 and 6.)
- Is anything lost for good? No. Every reading is also in the logs, so missed ones are rebuilt, automatically, every time. (Finding 7.)
- If only writes are refused, why can't we still read? A fair question, so I tested it rather than assumed: reads are refused too. The records now have a copy that can always be read. (Finding 8.)
- What stops it happening this way again? A third of the writes per move, load tests that name the resource they could exhaust, and both written into the repo's rules. (Findings 1 and 3.)
The rule that came out of it
The fixes that matter most are the ones that change what happens next time. When this site's records can't be written, for any reason, each missed reading is logged in full. The heartbeat notices the refusal, and on the first healthy run afterwards the missed readings are rebuilt from the logs, back into each day's totals; any that failed or were slow are kept in full and marked as rebuilt, so the record stays honest about them. Running out of a budget is never treated as an incident, and never hidden either.
Live reliability · The raw record · Postmortem: game day 1 · How this site is run