Professional Data EngineerOperationalizing machine learning modelsMedium

A financial institution is developing a machine learning model to detect fraudulent transactions. Due to strict regulatory compliance requirements, all data used for training and inference must be stored within a specific geographical region and adhere to data residency policies. Which Google Cloud service is most appropriate to ensure that the datasets for this ML model are consistently stored and processed within the required region?

  1. ACloud SQL with Any-region instances
  2. BCloud Storage with Regional buckets
  3. CBigQuery with Multi-region datasets
  4. DCloud Spanner with Multi-regional instances
Show answer & explanation

Correct answer: B. Cloud Storage with Regional buckets

For strict data residency requirements, Cloud Storage Regional buckets ensure that data remains within a single, specified geographical region. Multi-regional options for other services distribute data across multiple regions, which would violate strict single-region residency policies.

Why the other options are wrong

  • A. Cloud SQL instances are regional, but the option 'Any-region instances' is not a standard configuration and doesn't guarantee the specific single-region residency required.
  • C. BigQuery multi-region datasets store data in multiple physical locations within a continent, which violates strict single-region residency.
  • D. Cloud Spanner multi-regional instances distribute data across multiple regions, which violates strict single-region residency.

Data Residency

The geographical location where data is stored and processed, often mandated by legal or regulatory requirements.

  • Crucial for compliance in regulated industries.
  • Requires careful selection of cloud resource locations.
  • Different from data locality or data sovereignty, though related.

Memory trick: Keep the data 'resident' in its own 'region' like a local.

More Operationalizing machine learning models questions