AWS Certified Data Engineer – AssociateData Ingestion and TransformationMedium

A global media company needs to process large volumes of video metadata (XML files) generated hourly from various content creation studios worldwide. These files are typically 10-50 MB each, and the company requires a serverless, scalable, and cost-effective solution to parse these XML files, extract specific attributes (e.g., title, duration, creation date), and store them in a structured format in Amazon DynamoDB for quick querying. Which combination of AWS services would BEST meet these requirements?

  1. AAWS Glue ETL (Spark), Amazon S3, Amazon DynamoDB
  2. BAmazon Kinesis Data Firehose, AWS Lambda, Amazon DynamoDB
  3. CAWS Transfer Family, AWS Lambda, Amazon DynamoDB
  4. DAWS Step Functions, AWS Lambda, Amazon S3
Show answer & explanation

Correct answer: C. AWS Transfer Family, AWS Lambda, Amazon DynamoDB

AWS Transfer Family (specifically SFTP/FTP/FTPS) is ideal for receiving files from external partners. Upon file arrival in S3, an S3 event can trigger an AWS Lambda function to parse the XML and extract metadata, which can then be stored in Amazon DynamoDB for structured access. This combination is serverless, scalable, and cost-effective for event-driven processing of individual files.

Why the other options are wrong

  • A. AWS Glue ETL (Spark) is powerful for large-scale batch processing but might be overkill and less cost-effective for hourly processing of individual 10-50 MB files, especially when simple parsing is needed. It's not primarily an ingestion endpoint.
  • B. Amazon Kinesis Data Firehose is for streaming data ingestion. While it can deliver to S3, it's not the primary mechanism for receiving files via FTP/SFTP, and Lambda would still be needed for parsing, making Transfer Family a better fit for the ingestion part.
  • D. AWS Step Functions orchestrates workflows but doesn't provide the file ingestion endpoint directly, nor does it inherently handle the parsing and storage. S3 is for storage, but the direct ingestion and processing logic is missing here.

Serverless File Ingestion & Processing

A pattern for receiving files from external sources, triggering an event-driven function to process their content, and storing the extracted data in a structured database, all without managing servers.

  • Uses services like AWS Transfer Family for secure file reception.
  • Leverages S3 events to trigger processing functions (e.g., Lambda).
  • Ideal for event-driven, small-to-medium file processing.

Memory trick: Transfer Family brings the XML, Lambda parses the data, DynamoDB stores it fast.

More Data Ingestion and Transformation questions