Engineering · Reliability · How we work

Six of seven findings were the tool

We wrote something to find a class of bug, it reported seven, and we had nearly fixed all seven before noticing that six of them were already correct.


Our live error log holds a cluster of sixty-one entries that all say the same unhelpful thing: a background request failed, and there is no record of which request it was. They are all dated a day in September, the day before we added the labelling that would have identified them. Pre-fix, and still a mystery.

So we wrote something to scan our own code for the pattern that produces them: a request fired off where nobody is waiting for the answer and nobody handles a failure. It found seven.

We went through them, wrapped all seven in the reporting helper, and wrote up the conclusion — including a confident sentence naming which call had caused the sixty-one live errors.

Then a test we had written weeks earlier went red

It was asserting that one specific call already handled its own failures. And it did. It always had — it is one line, and the handling is on the end of it, plainly visible.

The scanner had been reading each call only as far as the first semicolon. For a request with a bit of work attached to the result, the first semicolon lands inside that work — several lines above where the failure handling actually sits. So it looked at six calls that were already correct, could not see the part that made them correct, and reported them.

Six of seven findings were artefacts of the tool. And the sentence we had written about the live errors could not possibly be true, because the call we blamed cannot fail in the way described.

We reverted all of it

Every one of those six changes came back out. The release shipped no product code at all. What it shipped was the corrected scanner, now reading each call properly — which reports exactly one, the one that was genuinely unhandled — and a test holding the old broken reading alongside the right one, so that the next person's version of this mistake fails instead of landing.

We are writing this up because the failure mode generalises well beyond software. A measuring tool that is subtly wrong does not return nothing. It returns a confident, specific, plausible list — and the list gets written down, and the write-up becomes the thing the next person trusts. Six unnecessary changes to working code would have been the smaller cost. The larger one was very nearly a permanent note in our records attributing sixty-one real errors to a call that did not cause them.

The only reason we caught it was that a test written for an entirely different reason, weeks before, happened to assert the thing the tool had just got wrong. That is not a process. We would rather it had been one.


Try it on tonight’s service.

Nothing to install, no card. Not better by the weekend? Close the tab.