AWS Certified Data Engineer – AssociateData Storage and ManagementMedium
A data engineering team is designing a data lake using Amazon S3. The data consists of large files, and they want to improve query performance and reduce the amount of data scanned by analytical engines like Amazon Athena. They decide to organize their data in S3 using a directory structure that reflects the data's characteristics. Which strategy should they employ?
- AEncrypt all S3 objects using server-side encryption with S3-managed keys (SSE-S3).
- BStore all data in a single flat directory.
- CPartition the data by relevant columns such as `year`, `month`, and `day`.
- DUse S3 bucket versioning for all data files.
Show answer & explanationAnswer & explanation
Correct answer: C. Partition the data by relevant columns such as `year`, `month`, and `day`.
Partitioning data in S3 by relevant columns (e.g., year, month, day) creates a hierarchical directory structure. Analytical engines like Athena can then use these partitions to prune data, scanning only the necessary subsets of data, which significantly improves query performance and reduces costs.
Why the other options are wrong
- A. Encrypting S3 objects is a security measure and does not directly improve query performance or reduce the amount of data scanned by analytical engines. It might even introduce a slight overhead.
- B. Storing all data in a single flat directory would lead to full table scans for most queries, severely hindering performance and increasing costs.
- D. S3 bucket versioning helps with data recovery and protection against accidental deletion, but it does not improve query performance or reduce data scanned by analytical engines.
S3 Data Partitioning
Organizing data in Amazon S3 by creating a hierarchical directory structure based on column values, enabling query engines to filter data more efficiently.
- Improves query performance in services like Athena
- Reduces the amount of data scanned, lowering costs
- Commonly partitions by time (year, month, day) or other categorical data
- Data is stored in folders like `s3://bucket/table/year=YYYY/month=MM/day=DD/`
Memory trick: Divide and conquer your data with partitions to make queries faster and cheaper.