AWS Certified Data Engineer – AssociateData Operations and MonitoringMedium
A data engineering team manages a daily ETL pipeline that extracts data from a transactional database, transforms it using AWS Glue, and loads it into an Amazon S3 data lake. Over time, the volume of data has grown significantly, causing the Glue job to exceed its allocated runtime window. The team needs to reduce the job's execution time. They have already optimized the Glue script for efficiency. Which cost optimization strategy should they consider NEXT to improve runtime performance for the growing data volume?
- AImplement incremental data loading using AWS Glue Job bookmarks.
- BSwitch from Amazon S3 Standard to S3 Intelligent-Tiering for storage cost savings.
- CMigrate the data lake from S3 to Amazon Redshift for better query performance.
- DIncrease the number of Data Processing Units (DPUs) for the Glue job.
Show answer & explanationAnswer & explanation
Correct answer: A. Implement incremental data loading using AWS Glue Job bookmarks.
Since the script is already optimized, the next most effective strategy for reducing runtime with growing data volumes is to process less data. AWS Glue Job bookmarks enable incremental processing by tracking previously processed data, ensuring that only new or changed data is processed in subsequent runs, significantly reducing the workload and runtime.
Why the other options are wrong
- B. S3 Intelligent-Tiering is a storage cost optimization strategy for S3 itself, not a performance optimization for the Glue ETL job's runtime.
- C. Migrating to Redshift changes the data warehouse, but doesn't directly address the Glue ETL job's runtime for processing data into a data lake on S3. It's a different architectural decision.
- D. Increasing DPUs scales out resources, which helps with performance but also directly increases cost. While it can reduce runtime, it's not a cost *optimization* strategy in itself if the problem can be solved by processing less data.
AWS Glue Job Bookmarks
A feature in AWS Glue that helps process incremental data by tracking the state of previously processed data, allowing ETL jobs to process only new or changed data on subsequent runs.
- Saves processing state information in a persistent store (DynamoDB).
- Reduces processing time and DPU costs.
- Supports various data sources (S3, JDBC).
- Must be enabled and configured in the Glue job properties.
Memory trick: Bookmarks mark the spot, so old data gets forgot. New data only, fast the job will trot.