Engineering · Performance · The till

One hop deep and twenty-two wide

The till menu waited once for everything, and the tests said so. From London it still took nearly two seconds. Depth was fixed. Width was the cost.


The office opened at six this morning and our speed board, which had read nothing for a week because nobody had opened a screen, lit seven rows red at once. The worst with enough readings to trust was the Menu screen in the back office: five loads, two and a quarter seconds on average, four seconds at the worst. The call it waited on longest was the same one the till asks for when it boots: the sellable menu.

That route has had two performance passes already. The first put every read it makes into one wave instead of fourteen in a line. The second went inside the helpers it calls and made each of those one wait too. There is a test that runs the route against a real database behind a counting face and refuses to pass if any read waits on another when it does not have to. It was green. By every measure we had, the route was one round trip deep.

The same route, two numbers

We asked the server how long it had spent on those five loads, from its own clock. It said 743, 910, 1,300 and 1,835 milliseconds. Then we asked the same route, same venue, same hundred and twelve kilobytes, from the sandbox that sits beside the database. It said 85. Same code. The only difference was where the server was standing when it ran, and this time the chain of waits that usually explains that gap was not there.

What was there was the width of the wave. The route fires about twenty-two database statements at the same instant: eleven of its own and eleven inside the six helpers. Beside the database each of those costs ten milliseconds and they overlap almost perfectly. From a server in London each is a crossing of the Atlantic, and twenty-two crossings started together do not come back together. They queue at the far end. We had measured that effect weeks ago on a different route and written it down: twelve copies of one call fired at once came back at 13, 13, 28, 37, 39, 67, 69, 75, 84, 94, 96 and 111 milliseconds. Scale a ten millisecond hop to a hundred and the slowest of twenty-two is most of a second. That is what the board was reading.

One request carrying eleven

The database can take a batch: one request with several statements in it, answered in order. The eleven statements that belong to the route itself now travel that way. Each answer comes back under the name the rest of the route already used for it, so not one line after the wave changed. The six helpers are shared with other screens and stay as they were, which leaves twelve requests in flight instead of twenty-two.

Measured on the live site after the deploy, from beside the database, the route went from 70 to 85 milliseconds of server time down to 43 to 55. Six copies fired at once, the shape the Menu screen actually produces, went from a spread of 104 to 221 down to 50 to 125. Those are the small end of the gain, because here each removed request was a ten millisecond hop. For the person who opened the office this morning each one was ten times that. The next warm loads of that screen, and of the till, are the real reading, and the board will show them against the release they ran on.

The lesson is a narrow one. A test that counts how deep a route waits will stay green while the route grows wider, and width has its own price once the server and the database are an ocean apart. Depth first, always. Then count the width.


Try it on tonight’s service.

Nothing to install, no card. Not better by the weekend? Close the tab.