A data engineering team manages a complex data pipeline involving multiple AWS services, including AWS Glue, Amazon Kinesis, and Amazon Redshift. They need to implement real-time monitoring and anomaly detection for key operational metrics (e.g., Kinesis PutRecord failures, Glue job error rates, Redshift query latencies). When an anomaly is detected, the system should automatically trigger a custom remediation action (e.g., a Lambda function to restart a Glue job or scale a Kinesis stream) and notify on-call engineers. Which combination of AWS services provides the MOST automated and integrated solution?
- AAWS CloudWatch Anomaly Detection for metrics, CloudWatch Alarms to trigger, and Amazon EventBridge to invoke a Lambda function for remediation and send notifications.
- BAWS CloudWatch Metrics for data collection, CloudWatch Alarms for anomaly detection, and Amazon SNS for notifications.
- CAmazon QuickSight for dashboarding, AWS Budgets for cost alerts, and Amazon Chime for team notifications.
- DAWS CloudTrail for event logging, CloudWatch Logs for analysis, and AWS Lambda for custom alerts and remediation.
Show answer & explanationAnswer & explanation
Correct answer: A. AWS CloudWatch Anomaly Detection for metrics, CloudWatch Alarms to trigger, and Amazon EventBridge to invoke a Lambda function for remediation and send notifications.
This option provides the most automated and integrated solution. AWS CloudWatch Anomaly Detection automatically learns normal metric behavior and detects deviations. CloudWatch Alarms can then be configured on these anomaly detection models. EventBridge is the ideal service to trigger a Lambda function (for custom remediation actions) and send notifications (via an SNS topic, which EventBridge can also target) when an alarm state is reached.
Why the other options are wrong
- B. This covers basic monitoring and alerting but lacks automated anomaly detection and a direct mechanism for triggering custom remediation actions, relying only on standard metric thresholds.
- C. QuickSight is for business intelligence, Budgets for cost, and Chime for communication. These services are not designed for real-time operational anomaly detection and automated remediation within data pipelines.
- D. CloudTrail is for auditing API calls, and CloudWatch Logs are for log analysis. While useful for troubleshooting, they are not the primary services for real-time metric-based anomaly detection or automated remediation.
Proactive Pipeline Automation (CloudWatch + EventBridge)
A pattern for automating operational responses in AWS data pipelines, where CloudWatch monitors metrics (including anomaly detection), Alarms trigger on deviations, and EventBridge orchestrates custom remediation (e.g., Lambda) and notification workflows.
- CloudWatch Anomaly Detection learns normal metric patterns.
- CloudWatch Alarms can trigger on anomalies or thresholds.
- EventBridge acts as a central event bus for routing alarm events.
- Lambda functions enable custom, programmatic remediation actions.
Memory trick: Anomaly Detected by CloudWatch, Alarm Triggers, EventBridge Routes to Lambda for Fix.