Professional Data EngineerBuilding and operationalizing data processing systemsEasy
A data engineering team is designing a new batch processing pipeline on Google Cloud to analyze historical sales data from multiple regional databases. The total data volume is in the hundreds of terabytes, and new data arrives daily, requiring a full refresh of the analytical dataset. The team needs a highly scalable, fully managed service that can perform complex SQL transformations efficiently and integrate well with BigQuery for final analysis. Which Google Cloud service should they choose for the data transformation step?
- ACloud Data Fusion
- BBigQuery
- CCloud SQL
- DCloud Spanner
Show answer & explanationAnswer & explanation
Correct answer: A. Cloud Data Fusion
Cloud Data Fusion is a fully managed, cloud-native data integration service that helps users efficiently build and manage ETL/ELT data pipelines. It is well-suited for large-scale batch processing and complex transformations, especially when integrating data from multiple sources and preparing it for BigQuery.
Why the other options are wrong
- B. BigQuery is primarily a data warehousing and analytics service, not a dedicated ETL/ELT tool for orchestrating complex transformations from diverse sources.
- C. Cloud SQL is a relational database service, not a data integration or transformation platform for hundreds of terabytes of batch data.
- D. Cloud Spanner is a globally distributed relational database, designed for transactional workloads, not for batch ETL/ELT of historical sales data.
Cloud Data Fusion
A fully managed, cloud-native data integration service built on open-source CDAP that enables users to efficiently build and manage ETL/ELT data pipelines.
- Graphical interface for pipeline development
- Supports various data sources and sinks
- Scalable for large-scale batch and streaming data
Memory trick: Fusion brings all your data together for BigQuery analysis.