Kubernetes and Cloud Native Associate (KCNA)Cloud Native ObservabilityHard

A platform team is observing high cardinality metrics in their Prometheus setup, leading to increased storage consumption and slower query performance. Which best practice or tool can help manage this issue by reducing the number of unique label combinations while retaining critical information?

  1. AUsing recording rules to pre-aggregate metrics
  2. BDisabling all metric labels
  3. CSwitching to Pushgateway for all metrics
  4. DIncreasing Prometheus server resources
Show answer & explanation

Correct answer: A. Using recording rules to pre-aggregate metrics

High cardinality (many unique label combinations) is a common issue in Prometheus. Recording rules allow pre-aggregating metrics at a lower cardinality (e.g., aggregating by service name instead of individual pod names), reducing the number of time series stored and improving query performance. This is a best practice for managing cardinality.

Why the other options are wrong

  • B. Disabling all labels would make metrics unidentifiable and unusable for troubleshooting, which is not a viable solution.
  • C. Pushgateway is for short-lived jobs, not a solution for managing cardinality of existing metrics.
  • D. Increasing resources only postpones the problem; it doesn't address the underlying cardinality issue.

Prometheus Recording Rules

Rules defined in Prometheus that allow pre-calculating frequently needed or computationally expensive expressions and storing their result as new time series.

  • Reduces query load by pre-calculating results.
  • Can be used to reduce metric cardinality by aggregating labels.
  • Improves dashboard load times and alert evaluation.

Memory trick: Recording Rules Reduce Redundancy and make Queries Quicker.

More Cloud Native Observability questions