AWS Certified Data Engineer – AssociateData Operations and MonitoringHard

A data engineering team operates an AWS Glue ETL job that processes large datasets daily. The job is configured to use a specific number of Data Processing Units (DPUs). Recently, the team noticed that the job's execution time has significantly increased, and CloudWatch metrics show high CPU utilization and memory pressure on the Glue worker nodes. They want to improve the job's performance and reduce its runtime without incurring excessive costs. What is the MOST effective first step to optimize the Glue job's performance in this scenario?

  1. AIncrease the number of DPUs allocated to the Glue job.
  2. BSwitch to a larger Glue worker type (e.g., G.2X to G.4X).
  3. CRefactor the Glue script to use more efficient Spark transformations and predicates.
  4. DEnable AWS Glue Job bookmarks to avoid reprocessing old data.
Show answer & explanation

Correct answer: C. Refactor the Glue script to use more efficient Spark transformations and predicates.

Before scaling up resources (DPUs or worker types), the most cost-effective and fundamental first step is to optimize the code itself. Efficient Spark transformations, pushing down predicates (filters) to the data source, and optimizing data reads can drastically reduce the computational load, memory footprint, and I/O, often leading to better performance without increasing DPU costs.

Why the other options are wrong

  • A. Increasing DPUs adds more processing power but also increases cost. It's a scaling-out solution, not necessarily an optimization of the underlying workload, and should be considered after code optimization.
  • B. Switching to a larger worker type provides more memory and CPU per worker, which is a scaling-up solution. Similar to increasing DPUs, it adds cost and should be considered after code-level optimizations.
  • D. Job bookmarks prevent reprocessing old data, which saves time on subsequent runs but doesn't improve the efficiency of processing the *new* data or address the high CPU/memory for the current workload.

Glue Job Optimization (Code-Level)

Techniques applied within an AWS Glue ETL script (e.g., using PySpark/Scala) to improve execution efficiency, reduce resource consumption, and decrease runtime, primarily by optimizing data transformations and I/O operations.

  • Use `pushdown_predicate` for S3/JDBC sources.
  • Choose efficient Spark transformations (e.g., `filter` before `join`).
  • Cache frequently used DataFrames.
  • Optimize partitioning strategies for writes.

Memory trick: Optimize the Code, then Scale the Load; save your cash, make your job fast.

More Data Operations and Monitoring questions