Vikrant Singh

· Real incident, self-inflicted

Postmortem: a load test used up the day's database writes

Summary

I load-tested the multiplayer game on a separate test copy of this site. The test copy shares the account's daily allowance of database writes with the live site, and an hour later that allowance ran out. For six and a half hours the live reliability page couldn't show its numbers, new games couldn't start, and the site's own records missed their readings. Pages, the heartbeat and the outside monitor were fine, and nothing raised a false alarm.

The worst part wasn't the outage. It was the page saying "the live numbers appear here once 30 days of data exist", which looked like a choice rather than a failure. Six minutes after the cause was found, it said what was actually wrong. Within two hours, every place that depends on that storage explained itself and said when it would be back, the game wrote a third as much per move, and missed readings had a way back.

Cause
Two load tests and the runs after them used up the daily 100,000 database writes on Cloudflare's free tier, shared by the test and live copies
User impact
Live reliability numbers and new Ludo games unavailable for 6 h 24 min. Every page served
Time to detect
Minutes, by eye. No signal raised it (finding 4)
Time to explain
6 minutes after the cause was confirmed: the page said what was wrong
Time to restore
The daily reset at 8:00 p.m.; the free tier has no earlier way back
Error budget used
None. The outside probe's checks all passed; the heartbeat never failed
Data
Missed readings rebuilt from logs after the reset; 140 of the day's 144 scheduled runs are on record

Timeline

Toronto time (EDT), 2 October.

12:21 p.m.Load test on the test copy: 40 games of four players for a minute, about 11,600 moves
12:30Second load test after a scaling change; then a round of regression games
1:10The scaling change ships. Each move now writes about six database rows (finding 3)
1:36The day's allowance runs out. Every database write on the account is refused
1:36The reliability page falls back to "numbers appear once 30 days exist" (finding 5)
~1:40I notice the page is wrong and ask why
1:42The cause is in the logs word for word: "Exceeded allowed rows written … free tier"
1:48The page now says the numbers are temporarily unavailable, and that the site itself is fine
2:09Every affected place names the cause and the reset time; the game cuts its writes to about two per move
2:58Rebuilding missed readings from logs becomes a standing rule; today's gap is queued
3:35Checked: with writes refused, could the records at least be read? No. Even a bare read is refused (finding 8)
3:45A read copy of the records now lives outside that storage; the page shows it, labelled "Recording paused since…"
8:00 p.m.The allowance resets; the rebuild runs, half an hour of readings at a time

What it looked like

Screenshots taken at the time, on a phone and by a scheduled capture. Cropped, otherwise as they were.

The Live reliability page on a phone. A banner reads: Live numbers are paused until the daily allowance resets. Nothing is broken. Today's free allowance of database writes is used up. It resets at 00:00 UTC (20:00 EDT), in about 2 min, and the numbers come back on their own.
7:58 p.m.: the reliability page names the cause and the reset time (finding 5).
The Ludo page on a phone, showing Ludo is napping. It's not you, and it's not the server. It's my wallet. The free allowance of database writes is used up. Back at 20:00 your time, in 2 min.
7:58 p.m.: the game says why it's resting and when it's back, instead of a dead connection (finding 6).
The Live reliability page after the reset. A banner reads: Catching up on readings since 2026-10-02 13:36 EDT; the readings are being rebuilt from the site's logs, half an hour at a time. Below it, the status is Operating normally and both objectives are within budget.
8:30 p.m., after the reset: missed readings are being rebuilt from the logs while everything else reads normal (finding 7). Nobody touched anything between this and the pictures above. One line overstated: only failed or slow rebuilt readings keep the "backfilled" mark (finding 9).

What went well

What hurt

I checked the wrong limit. Before the second test I looked at the daily request allowance and had room. The one that ran out was rows written, which is smaller in practice. A test plan should name the resource it might exhaust, and this one named the wrong one.

The test copy wasn't as separate as it looked. It had its own address, its own storage and no access to the live data, but it drew on the same account-wide allowance. Isolation that stops at the account boundary isn't isolation for capacity.

The scaling change made this worse. It removed a real bottleneck, every move queuing on one shared ledger, and cut the slowest 1% of moves from 240 to 80 ms. But it did so by writing more rows per move, and on this plan rows written are the scarce resource. It was optimised for the wrong constraint. Each move now writes one record and one small log row, a third of before.

And the failure looked like a decision. "The live numbers appear here once 30 days of data exist" is a sensible message for a flag that's off and a misleading one for storage that's refusing writes. A page that can't say why it's empty will be read as hiding something.

Findings and actions

#FindingActionStatus
1The test plan checked request limits, not rows writtenCapacity budget named in every test plan; no large load tests on the free tierDone
2Test and live copies share account-wide allowancesSeparate account or a paid plan before the next load testOpen (me)
3A scaling change raised writes per move from about 4 to 6One record plus one log row per move, about 2Fixed
4Detected by eye; nothing signalled the refused writesThe heartbeat notices refused writes and queues a rebuild (never a page)Fixed; a quiet notice is open
5The empty page looked like a choiceEvery affected place names the cause and the reset timeFixed
6The game failed as a dead connectionA "Ludo is napping" screen with the reset in your own timeFixed
7Six and a half hours of readings missedRebuilt from logs after any outage, automatically, back into each day's totalsFixed; the rebuild ran on 3 Oct
8The page went dark although only writes were refused; it turned out reads are refused tooA read copy of the 30-day records, refreshed every 10 minutes, shown with "Recording paused since…" when the store can't be readFixed
9The site said rebuilt readings were "marked as backfilled", but only failed or slow ones keep the mark: good ones go into the day's totals unmarked. For 2 Oct, two marked readings existEvery page now says what really happens. A visible count of rebuilt readings per day comes with the next change to how records are storedWording fixed; count open

How it was worked

The fixes came from a handful of questions, asked in this order. Each one changed something.

  1. Why is the page empty? It looked like a choice. The logs said otherwise, word for word, within minutes.
  2. What exactly ran out, and who shares it? Rows written, not requests; and the test copy shares the live site's allowance. (Findings 1 and 2.)
  3. Does this page anyone? It must not: a budget running out isn't an incident. I checked that nothing on the incident path depends on this storage, and watched the heartbeat stay green throughout.
  4. What do people see meanwhile? The cause and when it's back, everywhere it shows, instead of a blank or a dead connection. (Findings 5 and 6.)
  5. Is anything lost for good? No. Every reading is also in the logs, so missed ones are rebuilt, automatically, every time. (Finding 7.)
  6. If only writes are refused, why can't we still read? A fair question, so I tested it rather than assumed: reads are refused too. The records now have a copy that can always be read. (Finding 8.)
  7. What stops it happening this way again? A third of the writes per move, load tests that name the resource they could exhaust, and both written into the repo's rules. (Findings 1 and 3.)

The rule that came out of it

The fixes that matter most are the ones that change what happens next time. When this site's records can't be written, for any reason, each missed reading is logged in full. The heartbeat notices the refusal, and on the first healthy run afterwards the missed readings are rebuilt from the logs, back into each day's totals; any that failed or were slow are kept in full and marked as rebuilt, so the record stays honest about them. Running out of a budget is never treated as an incident, and never hidden either.

Live reliability · The raw record · Postmortem: game day 1 · How this site is run