The situation
Nothing is monitored until it has failed once. Backups exist but have never been restored. The runbook is in someone’s head, and that someone is on leave.
Our approach
First we agree what “working” means for each service in terms a business owner recognises, not CPU graphs. Then we instrument against that, alerting on the symptoms customers would notice, with thresholds that do not train people to ignore the alerts.
Alongside it, the unglamorous work: tested restores on a schedule, documented runbooks, patch windows, and a review after every incident that produces one concrete change.
What changes
You find out before your customers do. Recovery becomes a procedure anyone on the rota can follow.