Microsoft Certified: DevOps Engineer ExpertImplement a Site Reliability Engineering (SRE) strategyHard

A DevOps team is managing an Azure-based critical payment processing system. Their Service Level Objective (SLO) for transaction success rate is 99.9%. They have observed that the system typically operates at 99.95% success. The team wants to set up alerts that notify them ONLY when there is a risk of breaching the SLO, rather than when the SLO has already been breached. Which type of alerting strategy should they implement?

  1. AError budget burn rate alerting.
  2. BThreshold-based alerting on current success rate.
  3. CAvailability-based alerting.
  4. DLatency-based alerting.
Show answer & explanation

Correct answer: A. Error budget burn rate alerting.

Error budget burn rate alerting is designed to predict SLO breaches. Instead of waiting for the actual SLO (99.9%) to be violated, it triggers an alert when errors are consuming the allowed 'error budget' (the difference between 100% and the SLO, in this case, 0.1%) at an unsustainable rate. This provides proactive notification, allowing the team to intervene before the SLO is formally breached.

Why the other options are wrong

  • B. Threshold-based alerting on the current success rate would only trigger when the SLO (99.9%) is already breached or very close, which is reactive, not proactive.
  • C. Availability-based alerting focuses on uptime, not directly on the transaction success rate, and doesn't inherently provide proactive warning of a performance degradation.
  • D. Latency-based alerting monitors response times, which is a different metric than transaction success rate, and while related, doesn't directly address the proactive warning for success rate SLOs.

Error Budget Burn Rate Alerting

A proactive alerting strategy that triggers when the rate of error consumption exceeds a predefined threshold, indicating that the service is on track to exhaust its error budget and potentially breach its SLO.

  • Provides early warning before an SLO is breached.
  • Calculates how fast the 'allowed' errors are being used up.
  • Helps teams prioritize and take corrective action proactively.

Memory trick: Don't just watch the fence, watch how fast the budget burns!

More Implement a Site Reliability Engineering (SRE) strategy questions