Professional Data EngineerBuilding and operationalizing data processing systemsMedium

A global IoT company collects sensor data from millions of devices worldwide. This data is ingested into Cloud Pub/Sub and then processed by a Dataflow streaming pipeline. The data volume can fluctuate significantly throughout the day, with peak loads reaching 5x the average. The Dataflow pipeline needs to dynamically adjust its worker resources to handle these fluctuations without manual intervention, ensuring consistent processing latency and minimizing operational costs. Which Dataflow feature is crucial for achieving this dynamic resource allocation?

  1. AWorker pre-warming
  2. BFixed number of workers
  3. CManual scaling
  4. DAutoscaling
Show answer & explanation

Correct answer: D. Autoscaling

Dataflow's autoscaling feature automatically adjusts the number of workers based on the workload, ensuring that the pipeline can handle fluctuating data volumes efficiently, maintain consistent latency, and optimize costs by only utilizing necessary resources.

Why the other options are wrong

  • A. Worker pre-warming is not a standard Dataflow feature for dynamic resource allocation; it's more relevant to other compute services.
  • B. A fixed number of workers would lead to either over-provisioning during low traffic or under-provisioning during peak traffic.
  • C. Manual scaling requires constant monitoring and intervention, which is not suitable for dynamic, fluctuating workloads.

Dataflow Autoscaling

A Dataflow feature that automatically adjusts the number of worker instances in a pipeline based on the current workload, optimizing for throughput, latency, and cost.

  • Dynamically adds/removes workers
  • Handles fluctuating data volumes
  • Optimizes resource utilization and cost

Memory trick: Autoscaling expands and contracts like a data accordion.

More Building and operationalizing data processing systems questions