AWS Certified Data Engineer – AssociateData Storage and ManagementHard

A data engineering team is working with a large dataset (hundreds of terabytes) stored in Amazon S3. They need to reduce the storage footprint and improve query performance for analytical queries run by Amazon Athena. The data is accessed frequently, and the team wants to use a compression format that offers a good balance between high compression ratio and fast decompression, without significantly increasing CPU overhead during query execution. Which compression format should they choose?

  1. ASnappy
  2. BBZIP2
  3. CGZIP
  4. DLZMA
Show answer & explanation

Correct answer: A. Snappy

Snappy is a compression format that prioritizes speed over maximum compression ratio. It offers very fast compression and decompression, making it ideal for analytical workloads where CPU overhead during query execution needs to be minimized. While GZIP offers a better compression ratio, its slower decompression can negatively impact query performance for frequently accessed data. BZIP2 and LZMA offer even higher compression but are significantly slower.

Why the other options are wrong

  • B. BZIP2 offers a very high compression ratio but is significantly slower for both compression and decompression, making it less suitable for performance-critical analytical queries.
  • C. GZIP offers a good compression ratio but has slower decompression speeds compared to Snappy, which can impact query performance on frequently accessed data.
  • D. LZMA (used in XZ) offers the highest compression ratios but is the slowest for decompression, making it unsuitable for interactive analytical queries.

Snappy Compression

Snappy is a fast compression/decompression library developed by Google, designed for high-speed data processing rather than maximum compression ratio.

  • Optimized for speed (fast compression and decompression).
  • Lower CPU overhead during processing.
  • Good compression ratio, though not the highest.
  • Widely used in big data ecosystems (e.g., Hadoop, Spark, Parquet).
  • Ideal for analytical workloads where query performance is critical.

Memory trick: Snappy's swift, for queries to lift, GZIP's tight, but slows the light.

More Data Storage and Management questions