A global manufacturing company has an on-premises enterprise resource planning (ERP) system that is critical to its operations. The system runs on a relational database and requires high availability with an RPO of minutes and an RTO of less than an hour. The company wants to implement a disaster recovery (DR) solution in AWS, leveraging cloud capabilities while minimizing costs for the standby environment. Which DR strategy and AWS services should be recommended?
- AMulti-site Active/Active: Extend the on-premises environment to AWS with AWS Direct Connect, synchronize data in real-time to an Amazon Aurora Global Database, and run application instances simultaneously on-premises and in AWS.
- BBackup and Restore: Use AWS Backup to regularly back up the on-premises database and application data to Amazon S3. In a disaster, restore to Amazon EC2 instances and Amazon RDS.
- CWarm Standby: Replicate the on-premises database to an Amazon RDS Multi-AZ instance and deploy core application components on Amazon EC2 instances within an Auto Scaling group in a separate AWS Region. Keep EC2 instances running at minimum capacity.
- DPilot Light: Replicate the on-premises database to an Amazon RDS instance in a secondary AWS Region. Maintain a minimal set of core application EC2 instances in the secondary region, and scale up as needed during a disaster.
Show answer & explanationAnswer & explanation
Correct answer: D. Pilot Light: Replicate the on-premises database to an Amazon RDS instance in a secondary AWS Region. Maintain a minimal set of core application EC2 instances in the secondary region, and scale up as needed during a disaster.
A Pilot Light strategy aligns with the RPO of minutes and RTO of less than an hour, while being cost-effective. Replicating the database to Amazon RDS in a secondary region ensures data is current. Maintaining minimal EC2 instances allows for rapid scaling during a disaster, fulfilling the RTO without the higher cost of a warm standby.
Why the other options are wrong
- A. Multi-site Active/Active is the most expensive and complex DR strategy, providing the lowest RPO/RTO (seconds/minutes), but it is overkill for the stated RTO of 'less than an hour' and does not minimize costs for the standby environment.
- B. Backup and Restore typically has a higher RTO (hours to days), which does not meet the 'less than an hour' requirement. It's the most cost-effective but sacrifices RTO.
- C. Warm Standby would meet the RTO/RPO, but keeping EC2 instances running at minimum capacity in a separate region is more expensive than Pilot Light, which only keeps essential services running and scales up on demand.
Pilot Light Disaster Recovery
A DR strategy where a minimal set of core infrastructure is always running in the recovery region, and data is continuously replicated. Full capacity is provisioned during a disaster.
- Balances cost and recovery time.
- Data continuously replicated (low RPO).
- Applications scaled up on demand (moderate RTO).
Memory trick: Backup, Pilot, Warm, Hot: Know your recovery temperature.