A global manufacturing company wants to implement a robust disaster recovery (DR) strategy for its critical ERP system, which runs on Compute Engine instances with a Cloud SQL database. The company requires a Recovery Time Objective (RTO) of less than 4 hours and a Recovery Point Objective (RPO) of less than 15 minutes. The DR solution must be cost-effective and minimize operational overhead. What is the MOST appropriate DR strategy for this scenario?
- AHot Standby: Active-active deployment across two regions with real-time data replication and load balancing.
- BWarm Standby: Maintain a scaled-down but running replica of the entire environment in a different region, ready for quick scale-up.
- CBackup and Restore: Daily backups to Cloud Storage, manual restoration in a different region upon disaster.
- DPilot Light: Regularly replicate the Cloud SQL database to a standby instance in a different region and deploy Compute Engine instances on demand.
Show answer & explanationAnswer & explanation
Correct answer: B. Warm Standby: Maintain a scaled-down but running replica of the entire environment in a different region, ready for quick scale-up.
A Warm Standby strategy typically involves a scaled-down but running environment in a secondary region, allowing for RTOs in hours and RPOs in minutes. This aligns well with the less than 4 hours RTO and less than 15 minutes RPO, while being more cost-effective than Hot Standby and offering faster recovery than Pilot Light or Backup and Restore.
Why the other options are wrong
- A. Hot Standby offers the lowest RTO (near zero) and RPO (near zero) but is the most expensive and complex to implement, exceeding the 'cost-effective' and 'minimize operational overhead' requirements given the RTO/RPO targets.
- C. Backup and Restore typically results in RTOs of many hours or days and RPOs of hours or days, not meeting the required sub-4-hour RTO or sub-15-minute RPO.
- D. Pilot Light can achieve RPOs in minutes but often has RTOs in the range of several hours to a day due to the need to provision and start Compute Engine instances from scratch, which might barely meet or slightly exceed the 4-hour RTO.
Disaster Recovery (DR) Strategies
Disaster Recovery (DR) strategies define how an organization recovers its IT infrastructure and data after a disaster, balancing Recovery Time Objective (RTO) and Recovery Point Objective (RPO) with cost and complexity.
- RTO is the maximum acceptable delay before services are restored.
- RPO is the maximum acceptable amount of data loss.
- Common strategies include Backup and Restore, Pilot Light, Warm Standby, and Hot Standby.
Memory trick: Backup, Pilot, Warm, Hot: DR's Path.