AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium
A financial institution needs to process large volumes of historical transaction data stored in Amazon S3. This data, in Parquet format, needs to be joined with customer master data from a relational database, transformed, and then loaded into a data warehouse for business intelligence reporting. The processing must be serverless, cost-effective, and able to scale dynamically based on data volume. Which AWS service is most suitable for this data transformation task?
- AAWS Glue ETL
- BAmazon EMR
- CAWS Lambda
- DAmazon Athena
Show answer & explanationAnswer & explanation
Correct answer: A. AWS Glue ETL
AWS Glue ETL is a fully managed, serverless ETL service designed for preparing and loading data for analytics. It can read various data formats from S3, connect to relational databases, perform complex transformations using Spark, and write to a data warehouse. Its serverless nature ensures cost-effectiveness and dynamic scaling.
Why the other options are wrong
- B. Amazon EMR is a managed cluster platform, not serverless, and requires managing EC2 instances, which is less cost-effective for intermittent large-scale jobs than Glue.
- C. AWS Lambda is suitable for small, short-running functions, not for large-scale, long-running data transformations involving terabytes of data and complex joins.
- D. Amazon Athena is an interactive query service for S3, not an ETL service for complex transformations and loading into a data warehouse.
AWS Glue ETL
A serverless data integration service that makes it easy to discover, prepare, move, and combine data for analytics, machine learning, and application development.
- Fully managed, serverless, and scales dynamically.
- Uses Apache Spark for powerful data transformations.
- Includes Glue Data Catalog for metadata management.
Memory trick: Glue: The serverless sticky solution for your data transformations.