Professional Data EngineerBuilding and operationalizing data processing systemsMedium

A retail company collects customer clickstream data from its website and mobile applications. This data needs to be transformed to enrich user sessions with demographic information from a CRM system, filter out bot traffic, and aggregate events for product recommendation engines. The company requires a serverless, horizontally scalable solution that can handle varying data volumes efficiently. Which Google Cloud service should be used for this transformation?

  1. ADataflow
  2. BCloud Functions
  3. CDataproc
  4. DCompute Engine
Show answer & explanation

Correct answer: A. Dataflow

Dataflow is a fully-managed, serverless service for executing Apache Beam pipelines, designed for both batch and stream processing. It provides auto-scaling and high throughput, making it ideal for complex data transformations like enrichment, filtering, and aggregation on varying data volumes.

Why the other options are wrong

  • B. Cloud Functions are suitable for small, event-driven tasks, but not for large-scale, continuous data transformations and aggregations typically found in a data pipeline.
  • C. Dataproc is a managed Apache Hadoop and Spark service, which provides server clusters and is not fully serverless, incurring more operational overhead than Dataflow for this use case.
  • D. Compute Engine provides IaaS VMs, requiring manual management and scaling, which doesn't meet the 'serverless' requirement for data transformation.

Dataflow

A fully-managed service for executing Apache Beam pipelines for both batch and stream data processing, offering serverless auto-scaling.

  • Serverless and auto-scaling
  • Supports batch and stream processing
  • Based on Apache Beam

Memory trick: Dataflow crafts data with serverless ease, a transforming river.

More Building and operationalizing data processing systems questions