AWS Certified Data Engineer – AssociateData Operations and MonitoringHard

A data engineering team operates a critical data pipeline that ingests data from various sources into Amazon S3, processes it with AWS Glue, and then loads it into Amazon Redshift. The pipeline runs hourly. The team has identified that the Redshift cluster is consistently over-provisioned during off-peak hours, leading to significant unnecessary costs. During peak hours, however, the cluster performs optimally. They need to implement a solution to automatically scale the Redshift cluster capacity up and down based on the actual workload, specifically for the data loading phase, to optimize costs without impacting peak performance. Which Redshift feature should they leverage?

  1. AImplement a custom Lambda function to manually resize the cluster based on CloudWatch metrics.
  2. BUtilize Amazon Redshift Spectrum for querying data directly in S3.
  3. CEnable Amazon Redshift Concurrency Scaling.
  4. DConfigure Amazon Redshift Managed Scaling.
Show answer & explanation

Correct answer: D. Configure Amazon Redshift Managed Scaling.

Amazon Redshift Managed Scaling automatically adjusts the number of nodes in a Redshift cluster to optimize performance and cost. It adds or removes nodes based on workload demand, ensuring that the cluster is adequately provisioned during peak loads and scaled down during off-peak hours, directly addressing the over-provisioning issue and cost optimization requirement.

Why the other options are wrong

  • A. Implementing a custom Lambda for manual resizing is an operational burden and less efficient than Redshift's native managed scaling feature, which is designed for this exact purpose.
  • B. Redshift Spectrum allows querying S3 data without loading it into Redshift, which can help with some cost optimization, but it doesn't address the scaling of the Redshift cluster itself for data ingestion workloads.
  • C. Redshift Concurrency Scaling adds temporary cluster capacity for read-only queries to handle concurrent users, but it does not scale the main cluster's compute for data loading or reduce base cluster size during off-peak periods.

Redshift Managed Scaling

Amazon Redshift Managed Scaling automatically adjusts the number of nodes in a Redshift cluster up or down based on workload demand, optimizing performance and cost by right-sizing the cluster capacity.

  • Automates cluster resizing (add/remove nodes).
  • Optimizes for both peak and off-peak workloads.
  • Reduces costs by preventing over-provisioning.

Memory trick: Redshift Managed Scaling is the 'Smart Size' for your wallet and speed.

More Data Operations and Monitoring questions