AWS Certified Data Engineer – AssociateData Ingestion and TransformationHard

A data analytics team needs to build a robust and fault-tolerant pipeline to process customer interaction data from various sources (web, mobile, CRM). The data is in various formats (JSON, CSV) and needs to be standardized, de-duplicated, and enriched with geographical information. The transformed data must be stored in a data warehouse (Amazon Redshift) for reporting. The team requires full control over the Spark environment for custom libraries and complex logic, and the ability to scale resources dynamically based on data volume. Which AWS service should they use for the transformation phase?

  1. AAWS Glue ETL
  2. BAmazon EMR with Apache Spark
  3. CAmazon Redshift Spectrum
  4. DAmazon Kinesis Data Analytics for Apache Flink
Show answer & explanation

Correct answer: B. Amazon EMR with Apache Spark

Amazon EMR with Apache Spark provides a managed cluster environment that offers fine-grained control over the Spark framework, allowing the use of custom libraries and complex logic. It can dynamically scale resources, making it suitable for varying data volumes, and is ideal for complex, batch-oriented ETL tasks that require advanced customization before loading into a data warehouse like Redshift.

Why the other options are wrong

  • A. AWS Glue ETL is serverless and excellent for many ETL tasks, but might not provide the 'full control over the Spark environment for custom libraries and complex logic' that a dedicated EMR cluster offers, especially for very specific or advanced Spark configurations.
  • C. Amazon Redshift Spectrum allows querying data directly in S3 using Redshift, but it's a query service, not a transformation engine for de-duplication, enrichment, and standardization into a data warehouse.
  • D. Kinesis Data Analytics for Apache Flink is for real-time stream processing, not for batch processing of historical customer interaction data into a data warehouse.

Amazon EMR with Apache Spark

A managed cluster platform that simplifies running big data frameworks, such as Apache Spark, Hadoop, Presto, and Hive, on AWS to process and analyze vast amounts of data.

  • Provides a managed Apache Spark (or other framework) cluster.
  • Offers full control over the cluster and framework configuration.
  • Scalable resources, can be provisioned and de-provisioned as needed.
  • Ideal for complex, custom, and batch-oriented big data processing.

Memory trick: EMR is your 'Personal Lab' for Big Data experiments.

More Data Ingestion and Transformation questions