A large e-commerce company operates multiple Kubernetes clusters across different geographical regions. They need to aggregate and analyze metrics from all these clusters into a single, unified view to enable global monitoring and correlation. However, they are concerned about the high cardinality of certain metrics (e.g., user IDs, session IDs) from their services, which could overwhelm their Prometheus instances in each cluster. Which Prometheus-ecosystem component is designed to handle high-cardinality metrics and long-term storage across multiple clusters efficiently?
- APrometheus Pushgateway
- BPrometheus Alertmanager
- CThanos
- DNode Exporter
Show answer & explanationAnswer & explanation
Correct answer: C. Thanos
Thanos is specifically designed to extend Prometheus for global query view, high availability, and long-term storage, making it ideal for multi-cluster scenarios with high-cardinality metrics. It achieves this through components like Thanos Sidecar, Querier, and Compactor, which can deduplicate, downsample, and store metrics efficiently in object storage. Pushgateway is for ephemeral jobs, Alertmanager handles alerts, and Node Exporter collects host-level metrics.
Why the other options are wrong
- A. Pushgateway is for short-lived jobs to push metrics, not for global aggregation or high cardinality.
- B. Alertmanager handles alerts generated by Prometheus, not metric storage or high cardinality.
- D. Node Exporter collects host-level metrics for a single node, not a solution for multi-cluster high cardinality.
Thanos
An open-source project that extends Prometheus to enable global query views, high availability, and long-term historical data storage across multiple Prometheus instances, typically using object storage.
- Solves multi-cluster monitoring challenges for Prometheus.
- Provides a single pane of glass for all Prometheus metrics.
- Handles long-term storage and downsampling of metrics.
Memory trick: Thanos collects all the Prometheus gems for a global view.