A financial services company is storing customer transaction data in Amazon S3. Due to regulatory requirements (e.g., GDPR, CCPA), they must be able to identify and redact or delete specific customer data upon request, even within archived datasets. The data is currently stored in S3 Glacier Deep Archive for cost efficiency. Which data classification and management strategy allows for efficient identification and targeted deletion of specific customer records while maintaining cost-effectiveness for long-term storage?
- AEncrypt each customer record individually using a unique KMS data key, store the encrypted records in S3 Glacier Deep Archive, and manage key access for deletion.
- BStore data in a structured format (e.g., Apache Parquet or ORC) in S3, use an external data catalog (e.g., AWS Glue Data Catalog) for metadata, and use Amazon Athena queries for identification.
- CStore data as individual customer files in S3 and use S3 Object Tags to link to customer IDs, then apply S3 Glacier Deep Archive lifecycle rules.
- DUse Amazon Macie to discover sensitive data, then manually process and delete identified customer data from S3 Glacier Deep Archive archives.
Show answer & explanationAnswer & explanation
Correct answer: B. Store data in a structured format (e.g., Apache Parquet or ORC) in S3, use an external data catalog (e.g., AWS Glue Data Catalog) for metadata, and use Amazon Athena queries for identification.
Storing data in a structured, queryable format like Parquet or ORC, combined with an external data catalog and Amazon Athena, allows for efficient identification and targeted deletion of specific rows (customer records) even within large, archived datasets without needing to restore entire archives. This is crucial for GDPR/CCPA data subject rights.
Why the other options are wrong
- A. While individual encryption might seem appealing for security, managing millions or billions of unique KMS data keys for individual records is complex and expensive. Decryption and deletion would still require restoring the entire object containing many records.
- C. Storing individual customer files might be too granular and lead to an explosion of objects, increasing management overhead and potentially cost. Deleting individual objects from Glacier Deep Archive also incurs retrieval costs and time delays.
- D. Amazon Macie is for discovery, not for efficient, targeted deletion. Manually processing and deleting from Glacier Deep Archive is extremely slow, expensive (due to retrieval costs), and impractical for frequent data subject requests.
Queryable Archived Data for GDPR
A strategy for storing large volumes of data in cost-effective archival storage (like S3 Glacier Deep Archive) while maintaining the ability to efficiently identify and delete specific customer records to comply with regulations like GDPR or CCPA.
- Involves storing data in structured, self-describing formats (e.g., Parquet, ORC).
- Leverages data catalogs (AWS Glue) and query services (Amazon Athena).
- Allows for targeted deletion of records without full archive restoration, minimizing cost and time.
Memory trick: Query Structured Archives, Delete with Athena.