Professional Cloud Security EngineerEnsuring data protectionMedium
A healthcare provider stores patient records in BigQuery. Due to strict regulatory compliance, they need to ensure that specific columns containing highly sensitive health information (PHI) are never visible in plain text to analysts, even if they have full BigQuery data viewer access, but still allow aggregate queries. Which BigQuery security feature should be implemented?
- AAuthorized views with column exclusion
- BRow-level security policies
- CColumn-level security with data masking
- DDataset access controls with data obfuscation
Show answer & explanationAnswer & explanation
Correct answer: C. Column-level security with data masking
Column-level security with data masking allows specific columns to be masked (e.g., partially hidden or tokenized) for users who do not have the appropriate permissions, while still allowing queries on the masked data for aggregate analysis. This meets the requirement of never showing plain text PHI but still allowing aggregate queries.
Why the other options are wrong
- A. Authorized views can restrict access to columns, but 'column exclusion' would mean the column is entirely absent, preventing aggregate queries on it. Masking provides a transformed value.
- B. Row-level security restricts which rows a user can see, not specific columns within those rows.
- D. Dataset access controls manage access to the entire dataset, not granular column-level visibility, and 'data obfuscation' is a general term, not a specific BigQuery feature for this scenario.
BigQuery Column-level Security with Data Masking
A BigQuery feature that allows you to define policies to mask (transform) data in specific columns based on user permissions, ensuring sensitive information is never exposed in plain text while still enabling analytical operations.
- Applies to columns, not rows.
- Masks data based on user roles/permissions.
- Allows aggregate queries on masked data.
Memory trick: To see or not to see, BigQuery's mask is the key!