Professional Data EngineerBuilding and operationalizing data processing systemsMedium
A data team developed a new Dataflow pipeline that processes gigabytes of data hourly. After deploying to production, they observe that the pipeline frequently stalls and reports 'Out of memory' errors on certain worker nodes, particularly during peak load. The pipeline code has been reviewed and optimized for memory efficiency as much as possible. Which Dataflow feature should they adjust to address these memory issues and ensure stable operation?
- AMax Workers
- BMachine Type
- CDisk Size
- DAutoscaling Algorithm
Show answer & explanationAnswer & explanation
Correct answer: B. Machine Type
If a Dataflow pipeline is experiencing 'Out of memory' errors, it indicates that the worker machines do not have sufficient RAM to process the current workload. Adjusting the `Machine Type` to one with more memory (e.g., from `n1-standard-1` to `n1-standard-4` or a memory-optimized type) will provide the necessary resources to prevent these errors.
Why the other options are wrong
- A. Increasing `Max Workers` allows more parallelism but doesn't solve memory issues on individual workers; it might even exacerbate them if each new worker also runs out of memory.
- C. Increasing `Disk Size` helps with I/O-bound or disk-intensive operations, but 'Out of memory' errors specifically point to RAM limitations.
- D. Adjusting the `Autoscaling Algorithm` affects how workers are added or removed, but it won't resolve fundamental memory deficiencies on existing workers.
Dataflow Worker Sizing
Dataflow workers are Compute Engine VMs that execute pipeline steps. Their `Machine Type` defines CPU and memory resources, which are critical for performance and preventing resource exhaustion.
- Machine Type determines CPU and RAM per worker.
- Insufficient RAM leads to 'Out of memory' errors.
- Choosing the right machine type is crucial for stability and cost.
Memory trick: Dataflow's power: Match the machine's mind to the data's might.