AWS Certified Solutions Architect – ProfessionalDesign Solutions for Organizational ComplexityMedium

A global media company needs to process vast amounts of unstructured data (images, videos, documents) generated from various sources worldwide. This data needs to be ingested, transformed, and analyzed to extract insights. The solution must be highly scalable, cost-effective, and provide robust data governance capabilities across multiple AWS accounts and regions. Which AWS services and architecture pattern should a solutions architect recommend?

  1. AUse Amazon Kinesis Data Firehose for ingestion, Amazon EC2 instances for transformation, and Amazon Redshift for analysis.
  2. BIngest data into Amazon S3, use AWS Glue for ETL and cataloging, store processed data back in S3, and query with Amazon Athena or Amazon Redshift Spectrum.
  3. CUtilize AWS Transfer Family for data ingestion, Amazon DynamoDB for storage, and Amazon OpenSearch Service for analysis.
  4. DStore data in Amazon EBS volumes attached to Amazon EC2, process with Apache Spark on EC2, and visualize with Amazon QuickSight.
Show answer & explanation

Correct answer: B. Ingest data into Amazon S3, use AWS Glue for ETL and cataloging, store processed data back in S3, and query with Amazon Athena or Amazon Redshift Spectrum.

Amazon S3 is highly scalable and cost-effective for storing vast amounts of unstructured data. AWS Glue provides serverless ETL and a Data Catalog for data governance. Querying directly from S3 using Athena or Redshift Spectrum is cost-effective and avoids data movement, enabling robust analytics.

Why the other options are wrong

  • A. Kinesis Data Firehose is good for streaming, but EC2 instances for transformation can be expensive and require more operational overhead than serverless Glue. Redshift is typically for structured data warehousing, not ideal for direct analysis of unstructured data without prior structuring.
  • C. AWS Transfer Family is for file transfers (SFTP/FTPS/FTP), not bulk ingestion of unstructured data from various sources. DynamoDB is a NoSQL key-value store, not ideal for storing vast unstructured files. OpenSearch Service is for search and log analytics, not primary data lake analysis of unstructured data.
  • D. EBS is block storage, not suitable for vast amounts of unstructured data across multiple accounts/regions. Apache Spark on EC2 requires significant operational management compared to serverless options like Glue.

AWS Data Lake

A centralized repository that allows you to store all your structured and unstructured data at any scale, making it accessible for analytics.

  • Typically built on Amazon S3.
  • Uses services like AWS Glue for ETL and cataloging.
  • Analyzed with services like Athena, Redshift Spectrum, EMR.

Memory trick: S3 lake, Glue transforms, Athena queries.

More Design Solutions for Organizational Complexity questions