Microsoft Certified: DevOps Engineer ExpertImplement a Site Reliability Engineering (SRE) strategyEasy
A DevOps team is responsible for a critical e-commerce application. They are performing a Post-Incident Review (PIR) after a major outage caused by a misconfiguration. During the PIR, their primary goal is to identify ALL contributing factors, not just the immediate cause, and to develop actionable improvements to prevent recurrence. Which SRE principle is the team primarily applying during this PIR?
- AMinimizing toil
- BEmbracing risk
- CMonitoring everything
- DLearning from failures
Show answer & explanationAnswer & explanation
Correct answer: D. Learning from failures
The core purpose of a Post-Incident Review (PIR) in SRE is to thoroughly investigate incidents, understand all contributing factors (not just the proximate cause), and derive actionable lessons to prevent similar issues in the future. This directly embodies the SRE principle of 'Learning from failures' to continuously improve system reliability.
Why the other options are wrong
- A. Minimizing toil is about automating repetitive manual tasks, which is a broader SRE goal but not the direct purpose of a PIR.
- B. Embracing risk is about balancing reliability with innovation using error budgets, not the primary focus of a PIR.
- C. Monitoring everything is a good practice for detection, but a PIR focuses on understanding *why* an incident occurred and *how* to prevent it, which goes beyond just monitoring.
Learning from Failures (SRE)
A core SRE principle emphasizing thorough, blameless post-incident reviews to understand systemic weaknesses and implement improvements, rather than assigning blame.
- PIRs are blameless.
- Focuses on identifying systemic issues and contributing factors.
- Aims to generate actionable items for improvement.
Memory trick: After the fire, learn the lessons, avoid the blame game.