Engineering · Performance · The back office

The seconds were the morning’s; the shape was ours

One morning the database answered slowly and nine screens went red at once. The milliseconds belonged to the morning. The waits in a row were ours to fix.


Our speed board lists every page by how long it took to load, measured in the browser that loaded it, and counts a page as over budget when three or more real loads on the current release average above three hundred milliseconds. Most mornings it shows nothing, because nobody has opened a screen yet. One morning it showed nine pages over budget out of ten measured, and the worst of them was taking over two seconds.

The first instinct is to take the worst number and chase it. We did not, and the reason is worth writing down.

Reading the gate

Every request in the back office pays one database hop at the front door, where the server decides whether you may see the venue you named. That hop is the same on every route and has not changed in months, so its timing is a clock for the wire between the browser's edge and the database. Most days it reads about ninety milliseconds from the UK. That morning every row on the board read between two hundred and forty and six hundred and thirty for the same single hop, and the difference between what the browser timed and what the worker timed was anywhere up to a second per call.

So the database was slow, and the wire was slow, and both were the morning's doing. No route change moves either. If we had taken the two-second dashboard as a job we would have spent the day rewriting a route that was already at its floor.

What the morning could not hide

What a slow morning does do is make shape visible. When one hop costs five hundred milliseconds, a route that makes three hops in a row costs fifteen hundred, and a route that makes one costs five hundred. The worker's own timing, minus the gate, divided by the gate, is roughly the number of serial hops. Two rows on the board came out at three where one would have done: the kitchen policy screen and the prep-areas list. Those were the job, and they were fixed that morning, with a counting harness confirming the hop count went from four to two on one and from three or four to two on the other.

The dashboard, at the top of the list by raw seconds, came out at exactly one hop after the gate. That is the floor for a route that must read something. It stayed as it was.

The rule we took from it

When everything is slow at once, the absolute numbers tell you about the day. The ratios tell you about the code. Take jobs from the ratios. A route that costs one hop on a bad day costs one hop on a good day too, and the one that costs three will still cost three when the wire recovers, just quietly enough that nobody looks.


Try it on tonight’s service.

Nothing to install, no card. Not better by the weekend? Close the tab.