We have RTO and RPO targets for several AWS workloads, but I’m skeptical of recovery objectives that have never been measured under realistic conditions. A backup job succeeding does not prove that a service can be restored, dependencies can be reconnected, and users can complete the workflows that matter.
We’re looking for a testing approach that is stronger than a tabletop exercise but safer and less disruptive than failing over an entire production stack on a regular basis. That could mean isolated restore tests, targeted service recovery, regional failover drills, DNS validation, or synthetic checks against restored application paths.
For AWS-heavy teams, how often are you testing recovery, and what evidence do you use to determine that an RTO or RPO is actually achievable?
Source: r/Terraform · by /u/Efficiet_Wecather934