Currently building out a disaster recovery runbook for environments spanning aws and gcp, and honestly feeling overwhelmed by the complexity.
We split infrastructure across both clouds for redundancy, mostly using terraform with some manual configurations and various scripts scattered around. The theory sounds solid for resilience, but our actual runbook has turned into about ten confluence pages, multiple playbooks across different repos, and a somewhat optimistic assumption that we could rebuild production in another region or cloud if something major went wrong.
Our security and compliance team keeps pushing for concrete RTO numbers and demonstrable proof we can restore complete environments quickly, including all the supporting pieces like iam configs, policies, secrets management, kafka topics, database subnet groups, dns routing, and everything else. We run some chaos engineering exercises but they typically focus on single region failures or individual managed service issues, not full provider outages or complex tenant problems. Not sure how realistic it is to target genuine multi cloud failover with hard evidence rather than theoretical plans.
For those with DR strategies you feel confident about, what does your actual runbook look like day to day, who maintains ownership, and how much is automated versus manual documentation? Open to any practical advice
Source: r/sre · by /u/Meloduriahicin4264