Professional Cloud ArchitectAnalyze and optimize technical and business processesHard

A global financial institution is migrating its on-premises data warehouse to BigQuery. They have strict data governance policies, including requirements for data lineage, masking sensitive data, and auditing all data access. They also need to ensure that data transformations are repeatable and version-controlled. Which Google Cloud service should be primarily used to meet these requirements?

  1. ACloud Storage for raw data staging and archival.
  2. BCloud Data Fusion for data integration and governance.
  3. CCloud Dataflow for data ingestion and transformation.
  4. DBigQuery for data storage and SQL-based transformations.
Show answer & explanation

Correct answer: B. Cloud Data Fusion for data integration and governance.

Cloud Data Fusion is an excellent choice for this scenario as it provides a fully managed, code-free data integration service built on open-source CDAP. It offers native capabilities for data lineage, data masking, auditing, and visual pipeline development, making transformations repeatable and version-controllable, directly addressing the strict governance needs.

Why the other options are wrong

  • A. Cloud Storage is used for raw data staging but does not offer the data transformation, lineage, masking, or auditing features needed for governance.
  • C. Cloud Dataflow is powerful for data ingestion and transformation but requires more code-centric development and doesn't natively provide the comprehensive data governance features like lineage tracking and masking in a visual, managed way as Data Fusion does.
  • D. BigQuery is the destination data warehouse and can perform SQL transformations, but it doesn't provide the end-to-end data integration, lineage, masking, and auditing capabilities required for the entire pipeline.

Cloud Data Fusion

A fully managed, cloud-native data integration service built on open-source CDAP, providing a visual, code-free environment for building and managing ETL/ELT pipelines with built-in data governance features.

  • Visual pipeline development (GUI-based).
  • Built-in data lineage tracking.
  • Data masking and security features.
  • Auditing capabilities for data access and transformations.
  • Supports various data sources and sinks.
  • Enables repeatable and version-controlled data pipelines.

Memory trick: Fusion ensures your data flows with full governance and clarity.

More Analyze and optimize technical and business processes questions