Professional Data EngineerBuilding and operationalizing data processing systemsMedium
A retail company collects customer clickstream data from its website and mobile applications. This data needs to be transformed to enrich user sessions with demographic information from a CRM system, filter out bot traffic, and aggregate events for product recommendation engines. The company requires a serverless, horizontally scalable solution that can handle varying data volumes efficiently. Which Google Cloud service should be used for this transformation?
- ADataflow
- BCloud Functions
- CDataproc
- DCompute Engine
Show answer & explanationAnswer & explanation
Correct answer: A. Dataflow
Dataflow is a fully-managed, serverless service for executing Apache Beam pipelines, designed for both batch and stream processing. It provides auto-scaling and high throughput, making it ideal for complex data transformations like enrichment, filtering, and aggregation on varying data volumes.
Why the other options are wrong
- B. Cloud Functions are suitable for small, event-driven tasks, but not for large-scale, continuous data transformations and aggregations typically found in a data pipeline.
- C. Dataproc is a managed Apache Hadoop and Spark service, which provides server clusters and is not fully serverless, incurring more operational overhead than Dataflow for this use case.
- D. Compute Engine provides IaaS VMs, requiring manual management and scaling, which doesn't meet the 'serverless' requirement for data transformation.
Dataflow
A fully-managed service for executing Apache Beam pipelines for both batch and stream data processing, offering serverless auto-scaling.
- Serverless and auto-scaling
- Supports batch and stream processing
- Based on Apache Beam
Memory trick: Dataflow crafts data with serverless ease, a transforming river.