Microsoft Certified: DevOps Engineer ExpertImplement a Site Reliability Engineering (SRE) strategyMedium

A DevOps team is responsible for a critical e-commerce application. They notice that during peak sales events, the application experiences intermittent slowdowns and a slight increase in error rates, despite scaling resources. These issues are difficult to reproduce in lower environments. To effectively identify the root cause and improve the application's resilience, which health monitoring strategy should they prioritize implementing?

  1. ANetwork performance monitoring to identify latency issues between microservices.
  2. BApplication Performance Monitoring (APM) with distributed tracing and custom metrics.
  3. CInfrastructure monitoring using Azure Monitor to track VM CPU and memory usage.
  4. DLog aggregation and analysis using Azure Log Analytics and Kusto Query Language.
Show answer & explanation

Correct answer: B. Application Performance Monitoring (APM) with distributed tracing and custom metrics.

Intermittent slowdowns and increased error rates that are hard to reproduce often point to complex interactions within the application, especially in a microservices architecture. APM tools with distributed tracing are specifically designed to visualize requests end-to-end, identify bottlenecks across services, and pinpoint where latency or errors originate, even in production environments. While other monitoring types are useful, APM directly addresses the 'application's resilience' and 'root cause' of such complex, intermittent issues.

Why the other options are wrong

  • A. Network monitoring is important for network-related issues, but the description points to application-level behavior (slowdowns, error rates) that transcend simple network latency and require deeper application insights.
  • C. Infrastructure monitoring is foundational but might not pinpoint the application-level code or inter-service issues causing intermittent slowdowns, especially if resources appear adequate.
  • D. Log analysis is valuable for debugging but might not provide the holistic, real-time transaction view that distributed tracing offers for identifying performance bottlenecks across services.

Distributed Tracing (APM)

A technique used in Application Performance Monitoring (APM) to track requests as they flow through multiple services and components in a distributed system, providing an end-to-end view of transaction execution.

  • Helps identify latency bottlenecks and error origins.
  • Crucial for microservices architectures.
  • Provides a call graph or waterfall diagram of a request's journey.
  • Often implemented with OpenTelemetry or similar standards.

Memory trick: Intermittent bugs: 'Trace the path, see the flow.'

More Implement a Site Reliability Engineering (SRE) strategy questions