Professional Cloud ArchitectDesign and plan a cloud solution architectureHard

A global company requires a disaster recovery (DR) strategy for its critical application hosted on Google Cloud. The application must be restored and fully operational within 4 hours (RTO) and should not lose more than 15 minutes of data (RPO). The application uses Compute Engine instances, regional persistent disks, and a Cloud SQL database. Which DR strategy would meet these requirements most cost-effectively?

  1. ABackup and restore to a different region.
  2. BSnapshot persistent disks and export Cloud SQL backups to Cloud Storage in a different region.
  3. CActive-active deployment across multiple regions.
  4. DActive-passive deployment in a different region.
Show answer & explanation

Correct answer: D. Active-passive deployment in a different region.

An active-passive (warm standby) deployment in a different region involves maintaining a scaled-down, ready-to-go environment. This can meet an RTO of 4 hours by scaling up and redirecting traffic. For an RPO of 15 minutes, Cloud SQL's high availability or cross-region replicas and persistent disk replication (or frequent snapshots) can ensure minimal data loss. This is generally more cost-effective than active-active while still meeting the RTO/RPO.

Why the other options are wrong

  • A. Backup and restore typically has a much higher RTO (many hours to days) and RPO (depends on backup frequency) than required, making it unsuitable for a 4-hour RTO.
  • B. While snapshots and backups are components of DR, simply having them in another region doesn't define a strategy that guarantees a 4-hour RTO. The time to provision new VMs, restore data, and bring up the application can easily exceed 4 hours without a pre-configured environment.
  • C. Active-active deployment provides very low RTO/RPO (seconds/zero) but is significantly more complex and expensive due to maintaining full infrastructure in multiple regions simultaneously, exceeding the 'most cost-effectively' requirement.

RTO & RPO

Recovery Time Objective (RTO) is the maximum acceptable downtime after a disaster. Recovery Point Objective (RPO) is the maximum acceptable data loss after a disaster. These define DR strategy requirements.

  • RTO: How quickly you must recover.
  • RPO: How much data you can afford to lose.
  • Lower RTO/RPO typically means higher cost and complexity.

Memory trick: RTO is 'time to get up', RPO is 'data lost, oh no!'

More Design and plan a cloud solution architecture questions