A quiet night and a failed one look the same
A job that asks staff for their availability runs every hour. If the database refused, it returned the same result as a night with nobody to ask.
There is a small job in this system that runs on a schedule and asks staff who have not yet said when they can work next week. It looks up who is outstanding, sends the nudges, and reports what it did: so many checked, so many asked, so many skipped.
If the database did not answer, it reported nought checked, nought asked, nought skipped — and carried on.
Why that is worse than crashing
Those three zeros are also exactly what a perfectly healthy night produces when everybody has already submitted their availability. The two situations are not merely similar; the result is byte for byte identical. There is no reader — human or machine — that could ever have told them apart.
The code that calls this job was already correct. It logs a line when nudges go out, and it logs an error if the job fails. But the job never failed: it caught the problem and returned a tidy answer. And because the success line only prints when something was actually sent, a swallowed failure printed nothing at all. A hundred and sixty-eight runs a week, and a broken one leaves no trace anywhere.
The visible consequence is a rota. Staff do not get asked, managers chase people who think they already answered, and the week gets built on incomplete information — with every log in the system saying the night was quiet.
The sibling that was already right
The thing that makes this worth writing up is that we had the correct pattern in the next file along. A different nightly job handles each venue separately, counts its failures, names the venue in the log when one goes wrong, and deliberately does not catch the error on its own initial lookup — with a comment above it explaining that an unanswered list becoming an empty one would report "this estate has no venues" instead of "the database did not answer".
That comment described, precisely, the bug sitting in the file beside it. Someone had learned the lesson, written it down clearly, and it had not travelled one file sideways.
A failure that reports itself as nothing-to-do is the most expensive kind to own, because the cost is paid by somebody who never finds out there was a failure. The job now lets the error through to the handler that was waiting for it all along.