Professional Data EngineerBuilding and operationalizing data processing systemsMedium

A data team developed a new Dataflow pipeline that processes gigabytes of data hourly. After initial deployment, they observe that the pipeline consistently falls behind, leading to increasing backlogs in Pub/Sub. The Dataflow job metrics show high CPU utilization and low memory utilization on the workers. What is the most effective action to improve the pipeline's throughput and catch up with the incoming data?

  1. AIncrease the worker machine type to a larger size with more vCPUs.
  2. BDecrease the maximum number of workers to reduce overhead.
  3. CIncrease the maximum number of workers for Dataflow autoscaling.
  4. DSwitch the Dataflow job from streaming to batch processing.
Show answer & explanation

Correct answer: A. Increase the worker machine type to a larger size with more vCPUs.

High CPU utilization and low memory indicate that the workers are CPU-bound. Increasing the worker machine type to one with more vCPUs will provide more processing power per worker, improving throughput.

Why the other options are wrong

  • B. Decreasing workers would further reduce throughput and exacerbate the backlog.
  • C. While increasing workers can help, if individual workers are CPU-bound, simply adding more may not be the most efficient solution without addressing the per-worker bottleneck.
  • D. Switching to batch processing would not address the real-time processing requirement and would likely make the backlog worse due to delayed processing.

Dataflow Worker Sizing

Adjusting the machine type and number of workers for a Dataflow job to optimize performance and cost based on resource utilization metrics.

  • CPU-bound jobs benefit from more vCPUs.
  • Memory-bound jobs benefit from more RAM.
  • Autoscaling adjusts the number of workers, but machine type is fixed.

Memory trick: Match the worker's muscle to the job's demands.

More Building and operationalizing data processing systems questions