AWS Certified Solutions Architect – Associate (SAA-C03)Design Resilient ArchitecturesHard

A global ride-sharing company needs to ensure that its application, which relies heavily on a relational database, can quickly recover from a regional outage with minimal data loss. The company has a primary Amazon RDS for PostgreSQL instance in `us-east-1`. They need a cost-effective disaster recovery solution during normal operations, but one that can be scaled up quickly to a fully operational state in a secondary region. The RPO should be in minutes, and the RTO in tens of minutes. Which disaster recovery strategy is most suitable?

  1. AMulti-site Active/Active strategy with bidirectional replication between RDS instances.
  2. BWarm Standby strategy using a fully provisioned, scaled-down RDS instance in the secondary region.
  3. CBackup and Restore strategy with automated daily backups to S3 and cross-region copy.
  4. DPilot Light strategy using RDS cross-region read replicas.
Show answer & explanation

Correct answer: D. Pilot Light strategy using RDS cross-region read replicas.

A Pilot Light strategy is suitable for achieving RPO in minutes and RTO in tens of minutes while being cost-effective. For RDS, this means maintaining a cross-region read replica (the 'pilot light') that continuously replicates data. In a disaster, this read replica can be quickly promoted to a standalone primary instance, and other application components can be scaled up around it.

Why the other options are wrong

  • A. Multi-site Active/Active is the most expensive option as it requires running full production environments in multiple regions, contradicting the 'cost-effective during normal operations' requirement, although it offers the lowest RTO/RPO.
  • B. Warm Standby would involve a scaled-down *writeable* instance or a more complete environment, which is typically more expensive than a read replica (Pilot Light) and might offer a slightly better RTO but at a higher cost than specified for 'cost-effective'.
  • C. Backup and Restore typically results in RTOs of hours or more, as it involves provisioning a new database and restoring data, which does not meet the 'tens of minutes' RTO requirement.

RDS Pilot Light DR

A disaster recovery strategy for RDS where a cross-region read replica acts as a 'pilot light,' continuously replicating data at low cost, ready for quick promotion to a primary instance.

  • Achieves RPO in minutes (replication lag dependent).
  • Achieves RTO in tens of minutes (promotion and scaling).
  • Cost-effective during normal operations by only running a read replica.

Memory trick: Your RDS read replica is a Pilot Light, quietly humming in the next region, ready to flare up.

More Design Resilient Architectures questions