Professional Data EngineerBuilding and operationalizing data processing systemsEasy

A global IoT company collects sensor data from millions of devices worldwide. This data is ingested into Cloud Pub/Sub and then processed by a Dataflow streaming job. The data volume can fluctuate significantly throughout the day, with peak loads reaching ten times the average. The company needs to ensure the Dataflow job can handle these spikes without manual intervention, maintaining consistent processing latency. Which Dataflow feature is crucial for this requirement?

  1. AFixed Windows
  2. BSide Inputs
  3. CManual Scaling
  4. DAutoscaling
Show answer & explanation

Correct answer: D. Autoscaling

Dataflow's autoscaling dynamically adjusts the number of workers based on the incoming data volume and processing load, ensuring the pipeline can handle fluctuating spikes without manual intervention.

Why the other options are wrong

  • A. Fixed windows define how data is grouped by time, not how the processing resources adapt to load.
  • B. Side inputs are used to provide additional data to a PTransform, not for resource management or scaling.
  • C. Manual scaling requires human intervention and cannot react dynamically to fluctuating loads.

Dataflow Autoscaling

A Dataflow feature that dynamically adjusts the number of workers in a pipeline based on current processing needs, ensuring optimal resource utilization and performance.

  • Automatically scales up during peak loads and down during lulls.
  • Minimizes operational overhead.
  • Helps maintain consistent latency and throughput.

Memory trick: Autoscaling is like a self-adjusting engine.

More Building and operationalizing data processing systems questions