AWS Certified Data Engineer – AssociateData Operations and MonitoringMedium
A data analytics company uses AWS Glue to build and manage a data lake on Amazon S3. They have multiple Glue ETL jobs that transform raw data into a refined, queryable format. The company is experiencing high operational costs related to Glue job executions, particularly for jobs that process large datasets daily. The team has already optimized the Spark code within the Glue jobs. They now need to identify and implement a strategy to further reduce Glue ETL costs by minimizing the amount of data processed by each job run, without impacting data freshness or completeness. Which approach should they prioritize?
- AImplement AWS Glue Job Bookmarks to process only new or changed data.
- BTransition from Glue ETL to AWS Lambda functions for smaller data transformations.
- CSchedule Glue jobs less frequently to reduce the total number of runs.
- DIncrease the number of DPUs for Glue jobs to complete processing faster and reduce duration-based costs.
Show answer & explanationAnswer & explanation
Correct answer: A. Implement AWS Glue Job Bookmarks to process only new or changed data.
AWS Glue Job Bookmarks track previously processed data, allowing Glue jobs to process only new or changed data from the source in subsequent runs. This significantly reduces the amount of data scanned and processed, directly leading to lower execution times and thus lower costs, while maintaining data freshness and completeness.
Why the other options are wrong
- B. Lambda is suitable for small, event-driven tasks, but transitioning complex ETL for terabytes of data from Glue to Lambda is generally not feasible or cost-effective.
- C. Scheduling less frequently would reduce costs but would directly impact data freshness, which the requirement states should not be impacted.
- D. Increasing DPUs might reduce job duration but also increases the per-hour cost, potentially leading to higher overall costs, and doesn't address processing less data.
Glue Job Bookmarks
AWS Glue Job Bookmarks enable incremental processing by tracking previously processed data, ensuring that subsequent job runs only process new or changed data from the source, reducing processing time and cost.
- Tracks processed data for incremental runs.
- Reduces data scanned and processed.
- Lowers Glue ETL costs and execution time.
Memory trick: Bookmarks make Glue smart, only processing the new parts.