A DevOps team is adopting an SRE strategy. They are faced with a legacy application that frequently experiences critical failures, consuming significant on-call team time. The current mean time to recovery (MTTR) for these failures is unacceptably high. Which SRE practice should they prioritize to MOST effectively reduce the MTTR for this legacy application?
- AImplementing comprehensive end-to-end monitoring and alerting.
- BEstablishing clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- CConducting regular Post-Incident Reviews (PIRs) and implementing action items.
- DAutomating routine operational tasks and deployments.
Show answer & explanationAnswer & explanation
Correct answer: C. Conducting regular Post-Incident Reviews (PIRs) and implementing action items.
The problem states 'critical failures' and 'unacceptably high MTTR'. While all options contribute to SRE, Post-Incident Reviews (PIRs) are specifically designed to analyze incidents, identify root causes, and derive actionable improvements to prevent recurrence and reduce recovery times. By systematically addressing the lessons learned from each incident, PIRs directly lead to improvements that lower MTTR.
Why the other options are wrong
- A. Monitoring and alerting help with Mean Time To Detect (MTTD), but not directly with MTTR once an incident is detected, unless the alerts include actionable runbooks.
- B. SLOs and SLIs define targets but don't inherently reduce MTTR; they help identify when MTTR is too high.
- D. Automating tasks can reduce MTTR if those tasks are part of the recovery process, but without understanding the root causes through PIRs, automation might address symptoms rather than underlying issues, or automate inefficient processes.
Post-Incident Review (PIR)
A blameless process conducted after a significant incident to understand its causes, effects, and to identify actionable improvements that can prevent recurrence and reduce future Mean Time To Recovery (MTTR).
- Focuses on learning, not blaming.
- Identifies root causes, contributing factors, and systemic weaknesses.
- Generates actionable improvement items.
- Directly contributes to lower MTTR and increased system resilience.
Memory trick: Reduce MTTR: 'Review, Learn, Improve, Repeat!'