AWS Certified Data Engineer – AssociateData Storage and ManagementMedium

A data architect is designing a data lake on Amazon S3 for an e-commerce company. The data includes customer orders, product catalogs, and website clickstreams. To optimize query performance and reduce scanning costs for analytical queries, the architect wants to store the data in a columnar format that supports predicate pushdown and efficient compression. Which file format should be recommended?

  1. ACSV
  2. BApache ORC
  3. CXML
  4. DJSON
Show answer & explanation

Correct answer: B. Apache ORC

Apache ORC (Optimized Row Columnar) is a columnar storage format that provides significant performance benefits for analytic queries by allowing engines like Athena, Redshift Spectrum, and Spark to read only the necessary columns. It also supports efficient compression and predicate pushdown, which reduces the amount of data scanned.

Why the other options are wrong

  • A. CSV is a row-based format, not optimized for analytical queries needing columnar access or efficient compression.
  • C. XML is a hierarchical, text-based format, very verbose and not suitable for efficient analytical querying or storage optimization.
  • D. JSON is a semi-structured, row-based format, not ideal for columnar analytics or efficient compression for large datasets.

Apache ORC

An open-source, columnar data storage format optimized for big data processing engines like Apache Hive, Spark, and Presto. It provides efficient data compression and query performance.

  • Columnar storage minimizes I/O for analytical queries.
  • Supports various compression codecs.
  • Includes strong type support and predicate pushdown for query optimization.

Memory trick: Columnar formats cut query costs.

More Data Storage and Management questions