Professional Cloud ArchitectAnalyze and optimize technical and business processesHard

A global manufacturing company wants to implement a robust disaster recovery (DR) strategy for its critical applications hosted on Compute Engine instances across multiple Google Cloud regions. The applications require a Recovery Time Objective (RTO) of less than 15 minutes and a Recovery Point Objective (RPO) of less than 5 minutes. Data is stored on Persistent Disks attached to the Compute Engine instances. Which DR approach should they adopt?

  1. ADaily snapshots of Persistent Disks and manual restoration in a secondary region.
  2. BRegular backups to Cloud Storage and automated deployment of new instances from images.
  3. CSynchronous replication of Persistent Disks across regions with custom orchestration.
  4. DAsynchronous replication of Persistent Disks to a secondary region with Managed Instance Groups.
Show answer & explanation

Correct answer: D. Asynchronous replication of Persistent Disks to a secondary region with Managed Instance Groups.

Asynchronous replication of Persistent Disks to a secondary region, combined with Managed Instance Groups, allows for automated failover with an RPO of minutes (often sub-minute depending on replication frequency) and an RTO of minutes for application recovery, meeting the specified objectives without the performance overhead of synchronous replication across regions.

Why the other options are wrong

  • A. Daily snapshots result in an RPO of up to 24 hours (or more), far exceeding the 5-minute RPO requirement, and manual restoration will exceed the 15-minute RTO.
  • B. Regular backups to Cloud Storage can meet RPO depending on frequency, but restoring from Cloud Storage and deploying new instances, even with automation, is unlikely to meet an RTO of less than 15 minutes due to data transfer times and provisioning overhead.
  • C. Synchronous replication across regions typically incurs significant latency penalties, making it unsuitable for most applications requiring low-latency operations, and is complex to implement and manage at scale with custom orchestration.

Disaster Recovery (DR) Objectives

RTO (Recovery Time Objective) is the maximum acceptable downtime after a disaster. RPO (Recovery Point Objective) is the maximum acceptable data loss after a disaster.

  • Lower RTO/RPO often means higher cost and complexity.
  • Asynchronous replication is common for regional DR.
  • Snapshots, backups, and replication are key DR tools.

Memory trick: RTO, RPO, plan it fast, data's safe, disaster's past!

More Analyze and optimize technical and business processes questions