AWS Certified Data Engineer – AssociateData Storage and ManagementMedium

A data engineer is designing a data lake using Amazon S3. The data consists of large files (hundreds of MB to several GB each) that are frequently accessed by analytical queries. To optimize query performance and reduce scan times, the engineer decides to use a columnar storage format. Which columnar format is commonly used with S3-based data lakes and is known for its efficiency with analytical workloads?

  1. AXML
  2. BParquet
  3. CCSV
  4. DJSON
Show answer & explanation

Correct answer: B. Parquet

Parquet is a widely adopted columnar storage format in data lakes, particularly with S3. It is highly optimized for analytical queries because it stores data column by column, allowing query engines to read only the necessary columns, significantly reducing I/O and improving performance compared to row-based formats like JSON, CSV, or XML.

Why the other options are wrong

  • A. XML is a hierarchical, row-based format, highly verbose, and very inefficient for analytical workloads in a data lake.
  • C. CSV is a simple row-based text format, not columnar, and inefficient for large-scale analytical queries due to full row scans.
  • D. JSON is a row-based, semi-structured format, not columnar, and less efficient for analytical queries requiring specific columns.

Apache Parquet

Apache Parquet is a free and open-source columnar storage format designed for efficient data storage and retrieval, especially for big data processing frameworks.

  • Columnar storage format.
  • Optimized for analytical queries.
  • Supports complex nested data structures.
  • Offers efficient compression and encoding schemes.
  • Widely used with Apache Spark, Hive, Presto, and AWS Athena.

Memory trick: Columnar's the key, for speed you see, Parquet sets data free.

More Data Storage and Management questions