Microsoft Certified: DevOps Engineer ExpertImplement a Site Reliability Engineering (SRE) strategyEasy
A team manages a critical microservices-based application on Azure. They are implementing an SRE strategy. The application's service level objective (SLO) for availability is 99.95%, measured over a 30-day rolling window. During the last 30 days, the application experienced a total of 25 minutes of downtime. What is the current error budget status for this application?
- AThe application is within its error budget, but close to the limit.
- BThe error budget cannot be determined without knowing the total number of requests.
- CThe application has a significant error budget remaining.
- DThe application has exceeded its error budget and requires immediate intervention.
Show answer & explanationAnswer & explanation
Correct answer: C. The application has a significant error budget remaining.
The SLO of 99.95% for availability over 30 days means the allowable downtime is 0.05% of the total time. 30 days is 30 * 24 * 60 = 43200 minutes. 0.05% of 43200 minutes is 21.6 minutes. Since the actual downtime was 25 minutes, the application is slightly over its error budget.
Why the other options are wrong
- A. This is incorrect. The application is slightly over the error budget, not within it.
- B. This is incorrect. Availability-based error budgets are calculated based on uptime/downtime, not request count.
- D. This is incorrect. The calculation shows the application is slightly over, but not 'significantly exceeded' requiring 'immediate intervention' in a critical sense.
Error Budget Calculation (Availability)
The maximum allowable time a service can be unavailable or perform poorly while still meeting its Service Level Objective (SLO).
- Calculated as (100% - SLO%) * total time in period.
- Exceeding the error budget indicates a failure to meet the SLO.
- Helps balance reliability and innovation.
Memory trick: Error Budget: 'SLO, time, then compare.'