AWS Certified SysOps Administrator – AssociateReliability and Business ContinuityMedium

A company operates a critical web application that serves customers globally. The application is hosted on Amazon EC2 instances behind an Application Load Balancer (ALB) in a single AWS Region. The company wants to improve the application's fault tolerance and reduce recovery time in the event of a regional outage, aiming for a Recovery Time Objective (RTO) of less than 4 hours and a Recovery Point Objective (RPO) of less than 1 hour. Which disaster recovery strategy should the SysOps Administrator recommend?

  1. ABackup and Restore.
  2. BWarm Standby.
  3. CPilot Light.
  4. DMulti-site Active/Active.
Show answer & explanation

Correct answer: B. Warm Standby.

A Warm Standby strategy involves having a scaled-down but fully functional copy of your application running in a separate region. This allows for a relatively quick failover (RTO < 4 hours) and continuous data replication (RPO < 1 hour), making it suitable for the specified RTO/RPO targets.

Why the other options are wrong

  • A. Backup and Restore typically has a high RTO (many hours to days) as it requires restoring data and launching infrastructure from scratch, not meeting the < 4 hour RTO.
  • C. Pilot Light involves minimal infrastructure running, requiring more effort to scale up during a disaster, thus a higher RTO than Warm Standby, potentially exceeding 4 hours.
  • D. Multi-site Active/Active provides the lowest RTO/RPO but is the most complex and expensive, typically used for RTO/RPO in minutes or seconds, which is overkill for the given requirements and budget consideration.

Warm Standby DR Strategy

Warm Standby is a disaster recovery strategy where a scaled-down but fully functional copy of your application is continuously running in a separate region, ready for rapid failover.

  • Lower RTO/RPO than Backup & Restore or Pilot Light.
  • Higher cost and complexity than Pilot Light.
  • Data is continuously replicated to the standby region.
  • Infrastructure is provisioned and running, but often at reduced capacity.

Memory trick: Faster recovery costs more, slower costs less.

More Reliability and Business Continuity questions