Microsoft Certified: DevOps Engineer ExpertImplement a Site Reliability Engineering (SRE) strategyHard
A DevOps team is managing an Azure-based critical payment processing system. Their Service Level Objective (SLO) for transaction success rate is 99.9%. They have observed that the system typically operates at 99.95% success. The team wants to set up alerts that notify them ONLY when there is a risk of breaching the SLO, rather than when the SLO has already been breached. Which type of alerting strategy should they implement?
- AError budget burn rate alerting.
- BThreshold-based alerting on current success rate.
- CAvailability-based alerting.
- DLatency-based alerting.
Show answer & explanationAnswer & explanation
Correct answer: A. Error budget burn rate alerting.
Error budget burn rate alerting is designed to predict SLO breaches. Instead of waiting for the actual SLO (99.9%) to be violated, it triggers an alert when errors are consuming the allowed 'error budget' (the difference between 100% and the SLO, in this case, 0.1%) at an unsustainable rate. This provides proactive notification, allowing the team to intervene before the SLO is formally breached.
Why the other options are wrong
- B. Threshold-based alerting on the current success rate would only trigger when the SLO (99.9%) is already breached or very close, which is reactive, not proactive.
- C. Availability-based alerting focuses on uptime, not directly on the transaction success rate, and doesn't inherently provide proactive warning of a performance degradation.
- D. Latency-based alerting monitors response times, which is a different metric than transaction success rate, and while related, doesn't directly address the proactive warning for success rate SLOs.
Error Budget Burn Rate Alerting
A proactive alerting strategy that triggers when the rate of error consumption exceeds a predefined threshold, indicating that the service is on track to exhaust its error budget and potentially breach its SLO.
- Provides early warning before an SLO is breached.
- Calculates how fast the 'allowed' errors are being used up.
- Helps teams prioritize and take corrective action proactively.
Memory trick: Don't just watch the fence, watch how fast the budget burns!