A team is designing an SRE strategy for a new critical service. They have defined a Service Level Objective (SLO) for request latency: 99% of requests must complete within 200ms. To ensure this SLO is met and to provide early warning of potential issues, they need to set up an appropriate alert threshold. What is the MOST effective alert threshold strategy for this SLO?
- AAlert when 99% of requests exceed 200ms for 1 minute.
- BAlert when 95% of requests exceed 200ms for 5 minutes.
- CAlert when the 99th percentile of request latency exceeds 200ms for 1 minute.
- DAlert when the average request latency exceeds 200ms for 5 minutes.
Show answer & explanationAnswer & explanation
Correct answer: C. Alert when the 99th percentile of request latency exceeds 200ms for 1 minute.
The SLO is defined as the 99th percentile of requests completing within 200ms. Therefore, the most direct and effective alert threshold should directly monitor this specific percentile. Alerting on the 99th percentile exceeding 200ms directly correlates to breaching the SLO. A 1-minute duration provides a balance between responsiveness and avoiding transient spikes. Averaging latency (Option A) can hide issues affecting a significant portion of users, and alerting on a different percentile (Option C) doesn't directly track the defined SLO.
Why the other options are wrong
- A. This phrasing is inverted and unclear; it should be '99% of requests are _less than_ 200ms' or '1% of requests are _greater than_ 200ms'. If it meant '1% of requests exceed 200ms', then it's closer, but the standard way to express this is via percentile.
- B. Alerting on the 95th percentile does not directly align with the 99th percentile SLO, meaning issues affecting the SLO could be missed or alerts could be triggered too early/late.
- D. Averages can mask problems affecting a subset of users, as a few very fast requests can offset many slow ones. This does not directly track the 99th percentile SLO.
Percentile-based SLO and Alerting
Defining Service Level Objectives (SLOs) and corresponding alerts based on percentiles (e.g., 99th percentile) ensures that a certain percentage of user requests meet a performance target, providing a better measure of user experience than simple averages.
- Percentiles are robust against outliers.
- Directly reflects user experience for a segment of users.
- Alerts should mirror the SLO's percentile definition.
- Commonly used for latency and throughput SLOs.
Memory trick: SLO percentile: 'Alert on the P, not the A.'