AWS Certified Data Engineer – AssociateData Storage and ManagementMedium
A data engineer is designing a data lake on Amazon S3 for an e-commerce company. The data consists of customer transaction records stored as CSV files. To optimize query performance and reduce storage costs, the engineer decides to convert these CSV files into a columnar format. Which columnar data format offers the best balance of compression, query performance, and schema evolution capabilities for this use case?
- AJSON
- BXML
- CAvro
- DParquet
Show answer & explanationAnswer & explanation
Correct answer: D. Parquet
Apache Parquet is a columnar storage format optimized for analytical queries. It provides efficient data compression, improves query performance by allowing column pruning, and supports schema evolution, making it ideal for data lake scenarios.
Why the other options are wrong
- A. JSON is a row-oriented, text-based format, not columnar, and less efficient for analytical queries.
- B. XML is a hierarchical text-based format, inefficient for large-scale analytical processing in a data lake.
- C. Avro is a row-oriented binary format, good for data serialization and schema evolution, but not optimized for columnar analytical performance.
Apache Parquet
Apache Parquet is a free and open-source columnar storage format designed for efficient data storage and retrieval. It is optimized for analytical workloads, offering high compression ratios, improved query performance, and support for complex nested data structures and schema evolution.
- Columnar storage format
- Optimized for analytical queries
- High compression and efficient encoding
- Supports schema evolution
Memory trick: Parquet's Performance Powers Queries.