Professional Cloud ArchitectDesign and plan a cloud solution architectureHard
A company is designing a disaster recovery strategy for its critical application hosted on Compute Engine in a single region. The Recovery Time Objective (RTO) is 1 hour, and the Recovery Point Objective (RPO) is 15 minutes. The application data is stored on Persistent Disks attached to the VMs. Which disaster recovery approach should be recommended?
- AAsynchronous replication of Persistent Disks to a standby instance in a different region.
- BRegular backups to Cloud Storage and automated deployment of new VMs from custom images.
- CScheduled daily snapshots of Persistent Disks and manual restoration in a different zone.
- DContinuous replication of Persistent Disks to a standby instance in a different zone, combined with regional managed instance groups.
Show answer & explanationAnswer & explanation
Correct answer: D. Continuous replication of Persistent Disks to a standby instance in a different zone, combined with regional managed instance groups.
Continuous replication (e.g., using Persistent Disk asynchronous replication features or third-party tools) to a standby instance in a different zone meets the 15-minute RPO. Regional managed instance groups ensure automatic failover and quick recovery (RTO) within the region, even if a zone fails.
Why the other options are wrong
- A. Asynchronous replication to a different region provides good RPO, but the RTO might be longer due to cross-region failover and manual steps. The question specifies a single region primary deployment.
- B. Regular backups to Cloud Storage might exceed the 15-minute RPO, and automated deployment from images, while good for RTO, doesn't address the RPO of the data on Persistent Disks directly without continuous replication.
- C. Daily snapshots exceed the 15-minute RPO. Manual restoration will likely exceed the 1-hour RTO.
RTO and RPO
Recovery Time Objective (RTO) is the maximum acceptable delay before an application is available after a disaster. Recovery Point Objective (RPO) is the maximum acceptable amount of data loss after a disaster.
- RTO: How quickly you recover (time)
- RPO: How much data you can lose (data volume/time)
- Lower RTO/RPO implies more complex and costly solutions
- Critical for disaster recovery planning
Memory trick: RTO is 'Time Out', RPO is 'Point Of no return' for data.