Certified Cloud Security Professional (CCSP)Cloud Security OperationsHard

A cloud service provider (CSP) offers an 'always-on' service level agreement (SLA) with 99.999% availability for a critical application. To meet this stringent requirement, the CSP's operations team implements a strategy that includes active-active deployments across multiple geographically dispersed regions, automated failover mechanisms, and continuous health checks. This strategy primarily aligns with which aspect of Business Continuity and Disaster Recovery (BCDR)?

  1. AMean Time To Repair (MTTR).
  2. BRecovery Point Objective (RPO).
  3. CFault Tolerance.
  4. DRecovery Time Objective (RTO).
Show answer & explanation

Correct answer: C. Fault Tolerance.

Fault tolerance refers to a system's ability to continue operating without interruption despite the failure of one or more components. Active-active deployments, automated failover, and continuous health checks are all mechanisms designed to achieve high levels of fault tolerance, ensuring continuous availability even in the face of outages, directly supporting a 99.999% SLA.

Why the other options are wrong

  • A. MTTR measures how long it takes to repair a failed system, but the strategy is focused on avoiding repair time by maintaining continuous operation.
  • B. RPO defines the maximum acceptable data loss, which is related to data replication but not the primary focus of ensuring continuous operation during failures.
  • D. RTO defines the maximum acceptable downtime, but the described strategy is about preventing downtime altogether rather than just recovering quickly.

Fault Tolerance

The ability of a system to continue performing its intended function without interruption despite the failure of one or more of its components.

  • Achieved through redundancy, replication, and automatic failover.
  • Aims for continuous availability and minimal downtime.
  • Distinct from disaster recovery, which focuses on recovery after a major outage.

Memory trick: Fault tolerance means the system 'Tolerates' failure and keeps 'Flowing'.

More Cloud Security Operations questions