A data engineering team is building a new data lake on Amazon S3. They plan to use AWS Glue Data Catalog for metadata management and AWS Athena for ad-hoc querying. Data is ingested daily in CSV format. To optimize query performance and reduce Athena costs, the team needs to convert the data into an columnar, compressed format and partition it effectively. Which approach offers the BEST balance of cost-effectiveness, performance, and operational overhead?
- AUse an AWS Lambda function triggered by S3 events to convert CSV to Parquet and partition data.
- BUse Amazon Kinesis Firehose to ingest data directly into Parquet format with partitioning enabled.
- CSchedule an AWS Glue ETL job to convert CSV to Parquet and partition data.
- DManually convert and partition data using Amazon EC2 instances on a daily basis.
Show answer & explanationAnswer & explanation
Correct answer: C. Schedule an AWS Glue ETL job to convert CSV to Parquet and partition data.
An AWS Glue ETL job is ideal for this scenario. It provides a serverless, managed Spark environment perfect for large-scale data transformations like converting CSV to Parquet and partitioning. It integrates natively with S3 and Glue Data Catalog, offering excellent performance, cost-effectiveness (pay-as-you-go), and minimal operational overhead compared to managing EC2 instances or custom Lambda logic for complex transformations.
Why the other options are wrong
- A. Lambda is suitable for small, event-driven tasks but can be inefficient and complex for large-scale data transformations, especially with Spark-like features required for optimal Parquet conversion and partitioning.
- B. Kinesis Firehose can convert to Parquet and partition, but it's designed for real-time streaming ingestion. For daily batch ingestion of existing CSV files, it's not the primary tool and might add unnecessary complexity and cost compared to Glue ETL.
- D. Manually managing EC2 instances for daily data conversion is high in operational overhead, requires server management, and is generally less cost-effective than serverless alternatives for this type of workload.
Glue ETL for Data Lake Refinement
Utilizing AWS Glue ETL jobs to transform raw data in a data lake into optimized formats (e.g., Parquet, ORC) and structures (e.g., partitioning) for improved query performance and reduced costs.
- Serverless Apache Spark environment.
- Integrates with Glue Data Catalog for schema management.
- Supports various data sources and targets.
- Ideal for batch processing and large-scale transformations.
Memory trick: Glue's ETL is the Gold Standard for transforming your data lake's CSVs to Parquet, saving bucks and boosting speed.