Why monitoring is the first thing I install
One Saturday made this non-negotiable.
A client's production went down on a weekend. It wasn't a bug, it wasn't traffic, and it wasn't a bad deploy. A TLS certificate had quietly expired. Every server was healthy, every service was running — and every customer got a browser security warning instead of a checkout page.
The renewal had always been manual, and the person who used to do it had moved on. Nobody was watching that one number, so nobody knew until revenue stopped.
So I built the thing that makes it impossible to repeat, and now every environment I take on gets it in the first week — certificate expiry watched continuously, 30 days of warning before anything breaks, and renewal automated wherever the DNS provider allows it. Same story for disks filling up, queues backing up, and backups that silently stopped running three months ago.
Almost every outage I've been called into was something ordinary that nobody was looking at.