Kubernetes and Cloud Native Associate (KCNA)Cloud Native ObservabilityHard

A platform team is observing that their Prometheus server is frequently hitting its resource limits, primarily due to the high volume of metrics being scraped from a large fleet of ephemeral Kubernetes pods. They want to reduce the load on Prometheus while still retaining critical monitoring capabilities. Which strategy would be most effective in addressing this issue?

  1. AImplement relabeling rules to drop high-cardinality labels before ingestion.
  2. BIncrease the scrape interval for all metrics to 5 minutes.
  3. CDisable all alerting rules to free up CPU cycles.
  4. DMigrate all metrics to push-based collection using Pushgateway.
Show answer & explanation

Correct answer: A. Implement relabeling rules to drop high-cardinality labels before ingestion.

High cardinality (too many unique label combinations) is a primary cause of resource issues in Prometheus, leading to high memory usage and slow query times. Implementing relabeling rules to drop unnecessary high-cardinality labels (e.g., unique request IDs, ephemeral pod IDs that don't need to be indexed) before metrics are ingested into Prometheus is a highly effective strategy to reduce its resource footprint. Increasing scrape intervals reduces data resolution, Pushgateway is for ephemeral jobs, and disabling alerts sacrifices critical functionality.

Why the other options are wrong

  • B. Increasing scrape intervals reduces data resolution and might miss fast-changing issues.
  • C. Disabling alerting rules would free CPU but would eliminate a core monitoring capability, making it an unacceptable solution.
  • D. Migrating to Pushgateway for all metrics is generally not recommended for long-running services and can introduce other operational complexities.

Prometheus High Cardinality

A condition in Prometheus where a metric has an excessive number of unique label combinations, leading to high memory usage, increased disk I/O, and degraded query performance.

  • Caused by labels that change frequently or have many unique values.
  • Can lead to resource exhaustion and instability in Prometheus.
  • Mitigated by relabeling, aggregation, or using external storage solutions like Thanos.

Memory trick: Too many labels make Prometheus slow, so clean them up!

More Cloud Native Observability questions