Professional Cloud ArchitectAnalyze and optimize technical and business processesHard

A global manufacturing company wants to implement a robust disaster recovery (DR) strategy for its critical ERP system, which runs on Compute Engine instances with a Cloud SQL database. The company requires a Recovery Time Objective (RTO) of less than 4 hours and a Recovery Point Objective (RPO) of less than 15 minutes. The DR solution must be cost-effective and minimize operational overhead. What is the MOST appropriate DR strategy for this scenario?

  1. AHot Standby: Active-active deployment across two regions with real-time data replication and load balancing.
  2. BWarm Standby: Maintain a scaled-down but running replica of the entire environment in a different region, ready for quick scale-up.
  3. CBackup and Restore: Daily backups to Cloud Storage, manual restoration in a different region upon disaster.
  4. DPilot Light: Regularly replicate the Cloud SQL database to a standby instance in a different region and deploy Compute Engine instances on demand.
Show answer & explanation

Correct answer: B. Warm Standby: Maintain a scaled-down but running replica of the entire environment in a different region, ready for quick scale-up.

A Warm Standby strategy typically involves a scaled-down but running environment in a secondary region, allowing for RTOs in hours and RPOs in minutes. This aligns well with the less than 4 hours RTO and less than 15 minutes RPO, while being more cost-effective than Hot Standby and offering faster recovery than Pilot Light or Backup and Restore.

Why the other options are wrong

  • A. Hot Standby offers the lowest RTO (near zero) and RPO (near zero) but is the most expensive and complex to implement, exceeding the 'cost-effective' and 'minimize operational overhead' requirements given the RTO/RPO targets.
  • C. Backup and Restore typically results in RTOs of many hours or days and RPOs of hours or days, not meeting the required sub-4-hour RTO or sub-15-minute RPO.
  • D. Pilot Light can achieve RPOs in minutes but often has RTOs in the range of several hours to a day due to the need to provision and start Compute Engine instances from scratch, which might barely meet or slightly exceed the 4-hour RTO.

Disaster Recovery (DR) Strategies

Disaster Recovery (DR) strategies define how an organization recovers its IT infrastructure and data after a disaster, balancing Recovery Time Objective (RTO) and Recovery Point Objective (RPO) with cost and complexity.

  • RTO is the maximum acceptable delay before services are restored.
  • RPO is the maximum acceptable amount of data loss.
  • Common strategies include Backup and Restore, Pilot Light, Warm Standby, and Hot Standby.

Memory trick: Backup, Pilot, Warm, Hot: DR's Path.

More Analyze and optimize technical and business processes questions