Professional Data EngineerManaging and securing dataEasy
A research institution is building a data lake on Google Cloud using Cloud Storage. They frequently ingest new datasets from various external sources, and data scientists need to quickly discover, understand, and access these datasets. The institution requires a centralized metadata management solution that automatically catalogs data, supports custom metadata, and helps with data governance. Which Google Cloud service should they integrate?
- ACloud Spanner
- BBigQuery
- CCloud SQL
- DCloud Data Catalog
Show answer & explanationAnswer & explanation
Correct answer: D. Cloud Data Catalog
Cloud Data Catalog is designed for centralized metadata management, data discovery, and data governance across various Google Cloud data sources, supporting automatic metadata extraction and custom metadata.
Why the other options are wrong
- A. Cloud Spanner is a globally distributed relational database, not a metadata catalog.
- B. BigQuery is a data warehouse for analytics, not a metadata catalog for data discovery across a data lake.
- C. Cloud SQL is a managed relational database service, not a metadata catalog.
Cloud Data Catalog
Cloud Data Catalog is a fully managed, scalable metadata management service that enables organizations to discover, understand, and manage all their data assets across Google Cloud.
- Centralized metadata store.
- Supports automatic metadata extraction.
- Allows custom metadata (tags).
Memory trick: Catalog Connects Cloud Data Carefully.