AWS Certified Data Engineer – AssociateData Storage and ManagementMedium
A data engineer is designing a data lake using Amazon S3. The data consists of large files (several GB each) that are appended daily, but historical data is rarely updated. The team needs to optimize for query performance and cost efficiency when using services like Amazon Athena. Which partitioning strategy should the data engineer recommend?
- APartition by file size (e.g., /data/small/, /data/medium/, /data/large/)
- BNo partitioning, store all files in a single bucket prefix
- CPartition by year and month (e.g., /data/year=YYYY/month=MM/)
- DPartition by data type (e.g., /data/csv/, /data/json/, /data/parquet/)
Show answer & explanationAnswer & explanation
Correct answer: C. Partition by year and month (e.g., /data/year=YYYY/month=MM/)
Partitioning by year and month aligns with time-series data access patterns, allowing Athena to scan only relevant subsets of data, significantly improving query performance and reducing costs by minimizing data scanned.
Why the other options are wrong
- A. Partitioning by file size doesn't align with common query patterns and won't effectively reduce data scanned by Athena.
- B. No partitioning would force Athena to scan the entire dataset for every query, leading to high costs and poor performance, especially with petabytes of data.
- D. Partitioning by data type is useful for schema management but doesn't optimize for typical time-based analytical queries.
S3 Data Partitioning
S3 data partitioning organizes data in a hierarchical structure based on key values, which allows query engines like Athena to prune data and scan only relevant subsets, improving performance and reducing cost.
- Organizes data into logical groups
- Reduces data scanned by query engines
- Improves query performance and reduces cost
- Commonly uses time-based keys (year, month, day)
Memory trick: Time-based partitions make queries fly, saving money as data goes by.