Microsoft Certified: DevOps Engineer ExpertImplement a Site Reliability Engineering (SRE) strategyMedium

A team is managing an Azure-based critical payment processing system. Their Service Level Objective (SLO) for processing latency is that 99.9% of payment transactions must complete within 500 milliseconds. During a recent incident, the system experienced a sustained period where 1% of transactions took longer than 1 second, and 0.5% failed outright. To improve the system's reliability and prevent similar incidents, the SRE team decides to implement a proactive failure detection strategy. Which of the following should be their PRIMARY focus for this strategy?

  1. AImplementing alerts when CPU utilization exceeds 90% for 15 minutes.
  2. BSetting up alerts based on the 99th percentile of payment processing latency exceeding 500ms.
  3. CAnalyzing application logs for specific error codes indicative of database connection issues.
  4. DEstablishing synthetic transaction monitoring that simulates payment processing every minute.
Show answer & explanation

Correct answer: B. Setting up alerts based on the 99th percentile of payment processing latency exceeding 500ms.

The SLO is defined by the 99.9th percentile of latency. Monitoring the 99th percentile (Option C) directly tracks the health in relation to the SLO, providing an early warning before the 99.9th percentile is breached. While the SLO is 99.9%, alerting on 99% exceeding 500ms provides a good proactive signal without waiting for the SLO to be fully breached. Option B (synthetic monitoring) is good but less direct and real-time than actual transaction metrics. Options A and D are important but are lower-level infrastructure or diagnostic signals, not direct indicators of user-perceived performance against the SLO.

Why the other options are wrong

  • A. CPU utilization is an infrastructure metric and might not directly correlate with payment processing latency from a user perspective. High CPU doesn't always mean slow transactions, and slow transactions don't always mean high CPU.
  • C. Log analysis for error codes is reactive and diagnostic, indicating a problem has already occurred. While useful for root cause analysis, it's not a proactive 'detection strategy' for performance degradation against an SLO.
  • D. Synthetic monitoring is valuable for basic availability and performance, but it represents simulated user activity, not the actual live transaction performance. It's less ideal for detecting nuanced performance degradation within a specific percentile of real transactions.

Proactive SLO-aligned Alerting

Setting up monitoring and alerts that track Service Level Indicators (SLIs) in a way that provides early warning of potential SLO breaches, often by alerting on a slightly less stringent threshold or a lower percentile than the SLO itself.

  • Aims to detect degradation before users are significantly impacted.
  • Often involves monitoring percentiles (e.g., 99th percentile for a 99.9% SLO).
  • Allows for intervention before an error budget is consumed too quickly.
  • Requires careful selection of SLIs and thresholds.

Memory trick: Proactive SLO: 'Alert early, save the budget.'

More Implement a Site Reliability Engineering (SRE) strategy questions