Professional Data EngineerManaging and securing dataMedium
A research institution is building a data lake on Google Cloud using Cloud Storage. They frequently ingest various unstructured and semi-structured datasets from different research partners. To ensure proper data governance and discoverability, they need a centralized metadata catalog that automatically extracts schema and technical metadata, allows for business metadata tagging, and supports data lineage tracking across their datasets. Which Google Cloud service is best suited for this requirement?
- ABigQuery
- BCloud SQL
- CCloud Spanner
- DCloud Data Catalog
Show answer & explanationAnswer & explanation
Correct answer: D. Cloud Data Catalog
Cloud Data Catalog is a fully managed metadata management service that automatically discovers and catalogs data assets, extracts technical metadata, allows for business metadata tagging, and provides data lineage functionality, directly meeting the institution's needs.
Why the other options are wrong
- A. BigQuery is a data warehouse for analytics, not primarily a metadata catalog for diverse data assets.
- B. Cloud SQL is a managed relational database service, not a metadata catalog.
- C. Cloud Spanner is a globally distributed relational database, not a metadata catalog.
Cloud Data Catalog
A fully managed, scalable metadata management service that enables data discovery, governance, and understanding across an organization's data assets.
- Automatically extracts technical metadata.
- Allows tagging with business metadata.
- Supports data lineage tracking.
- Integrates with various Google Cloud data services.
Memory trick: Catalog your data, discover its secrets, track its lineage, and govern it well.