A financial services company uses Azure Kubernetes Service (AKS) for its core banking application. The SRE team needs to implement a disaster recovery strategy that ensures minimal data loss and rapid recovery in the event of a regional outage. The Recovery Point Objective (RPO) is 15 minutes, and the Recovery Time Objective (RTO) is 2 hours. Which disaster recovery strategy BEST meets these requirements?
- AActive-Active deployment across multiple Azure regions with real-time data synchronization.
- BWarm Standby deployment with synchronous data replication to a secondary region.
- CPilot Light deployment with asynchronous data replication to a secondary region.
- DBackup and Restore with daily backups to geo-redundant storage.
Show answer & explanationAnswer & explanation
Correct answer: C. Pilot Light deployment with asynchronous data replication to a secondary region.
An RPO of 15 minutes suggests asynchronous replication or near real-time data synchronization. An RTO of 2 hours means the secondary site needs to be quickly brought online. Pilot Light, with its core infrastructure running and data replicated, allows for faster recovery than cold backup and restore. Its asynchronous replication can meet the 15-minute RPO. Warm Standby with synchronous replication would typically be for a much lower RPO (near zero) and is more complex. Active-Active would achieve even lower RPO/RTO but is overkill and more expensive than required, and synchronous replication across regions might introduce unacceptable latency.
Why the other options are wrong
- A. Active-Active provides the lowest RPO and RTO, but it is the most complex and expensive strategy and exceeds the specified RTO/RPO, making it an over-engineered solution for the given requirements, especially with 'real-time data synchronization' which might imply synchronous replication across regions.
- B. Synchronous replication across regions generally implies very low RPO (near zero) but often introduces performance overhead. While it meets RPO, the 'Warm Standby' might be unnecessarily complex/expensive for a 2-hour RTO, and 'synchronous' across regions can be challenging.
- D. Daily backups would result in an RPO of 24 hours, far exceeding the 15-minute requirement.
DR Strategy: Pilot Light
A disaster recovery strategy where a minimal version of the application is continuously running in a secondary region, with data replicated asynchronously. In a disaster, the full application stack is provisioned and scaled up.
- Achieves RPOs of minutes to hours.
- Achieves RTOs of minutes to hours.
- Lower cost than Warm/Hot Standby but higher than Backup/Restore.
- Requires some pre-provisioned infrastructure in the DR region.
Memory trick: RPO/RTO: 'Cold, Pilot, Warm, Hot – pick your speed!'