Professional Data EngineerBuilding and operationalizing data processing systemsEasy

A data analytics team is designing a new batch processing pipeline to analyze historical customer purchase data. The data, currently stored in CSV files in Cloud Storage, needs to be loaded into BigQuery for complex analytical queries. The team wants to ensure data quality and perform transformations like data type conversions and column renaming before loading, without writing extensive custom code. Which Google Cloud service should they use to achieve this efficiently?

  1. ADataflow
  2. BCloud Pub/Sub
  3. CCloud Data Fusion
  4. DCloud SQL
Show answer & explanation

Correct answer: C. Cloud Data Fusion

Cloud Data Fusion is a fully managed, cloud-native data integration service that helps users efficiently build and manage ETL/ELT data pipelines using a visual interface, ideal for batch transformations and loading into BigQuery.

Why the other options are wrong

  • A. Dataflow is a powerful service for both batch and streaming data processing, but Cloud Data Fusion offers a more managed, low-code approach for common ETL scenarios.
  • B. Cloud Pub/Sub is a messaging service, primarily used for real-time data ingestion and event-driven architectures, not batch ETL.
  • D. Cloud SQL is a relational database service, not an ETL tool for transforming and loading data into BigQuery.

Cloud Data Fusion

A fully managed, cloud-native data integration service built on open-source CDAP, enabling ETL/ELT pipelines with a visual interface.

  • Visual, code-free data pipeline creation
  • Supports various data sources and sinks
  • Ideal for batch ETL/ELT and data quality tasks

Memory trick: Visual flows make data fuse seamlessly.

More Building and operationalizing data processing systems questions