Professional Data EngineerBuilding and operationalizing data processing systemsHard

A financial institution needs to process daily batch files containing millions of customer transactions. Each file is approximately 50 GB and arrives in Cloud Storage around midnight. The processing involves complex transformations, data validation, and aggregation before loading the results into BigQuery. The entire process must complete within a 4-hour window, and the solution needs to be cost-effective, paying only for the resources consumed during processing. Which Dataflow runner configuration should they prioritize to meet these requirements?

  1. AUsing a small number of workers with high CPU and memory.
  2. BEnabling Dataflow Prime with right-fitting and resource hints.
  3. CMaximizing worker instance size for faster processing.
  4. DImplementing custom autoscaling logic within the pipeline code.
Show answer & explanation

Correct answer: B. Enabling Dataflow Prime with right-fitting and resource hints.

Dataflow Prime, with its right-fitting and resource hints features, is designed to optimize resource allocation and cost for batch jobs by intelligently selecting the optimal mix of CPU, memory, and I/O resources, ensuring efficient processing within the time window while minimizing costs.

Why the other options are wrong

  • A. A small number of high-spec workers might not be able to process 50 GB files with millions of transactions fast enough within the 4-hour window.
  • C. Maximizing worker size might be faster but could be very expensive if not precisely optimized, potentially leading to over-provisioning.
  • D. Implementing custom autoscaling logic is complex and unnecessary, as Dataflow's built-in autoscaling and Dataflow Prime features already provide intelligent resource management.

Dataflow Prime Right-Fitting

A Dataflow Prime feature that intelligently analyzes pipeline characteristics and data patterns to automatically select the optimal worker type and resource configuration (CPU, memory, storage) for each stage of the pipeline, optimizing for both performance and cost.

  • Automated resource optimization
  • Reduces manual tuning efforts
  • Aims for optimal cost-performance balance

Memory trick: Prime right-fits the puzzle for perfect costs.

More Building and operationalizing data processing systems questions