AWS Certified Solutions Architect – ProfessionalDesign Solutions for Organizational ComplexityMedium

A global telecommunications company has a vast amount of historical call detail record (CDR) data stored in an on-premises Hadoop cluster. This data is critical for regulatory compliance and advanced analytics, but it is accessed infrequently by data scientists and auditors. The company wants to migrate this petabyte-scale data to AWS to reduce on-premises infrastructure costs and enable more flexible, serverless analytics. The solution must be cost-effective for long-term storage and support ad-hoc querying. Which AWS solution should a Solutions Architect recommend?

  1. ATransfer data to Amazon S3 and build a data lake using AWS Glue and Amazon Athena.
  2. BMigrate data to Amazon DynamoDB and use DynamoDB Accelerator (DAX) for querying.
  3. CMigrate data to Amazon RDS for PostgreSQL and use Amazon QuickSight for analytics.
  4. DRe-host the Hadoop cluster on Amazon EC2 instances and use Amazon EMR for processing.
Show answer & explanation

Correct answer: A. Transfer data to Amazon S3 and build a data lake using AWS Glue and Amazon Athena.

Migrating Hadoop data to Amazon S3 creates a cost-effective data lake for petabyte-scale, infrequently accessed data. AWS Glue can be used for cataloging and transforming the data, while Amazon Athena provides serverless, ad-hoc querying capabilities directly on S3 data, perfectly aligning with the requirements for cost-effectiveness, long-term storage, and flexible analytics without managing servers.

Why the other options are wrong

  • B. Amazon DynamoDB is a NoSQL database designed for high-performance, real-time applications with structured and semi-structured data, not for petabyte-scale, infrequently accessed historical data from a Hadoop cluster. DAX is a cache for DynamoDB, not a general analytics solution.
  • C. Amazon RDS for PostgreSQL is a relational database, not suitable for petabyte-scale, infrequently accessed, semi-structured/unstructured data like CDRs from Hadoop. QuickSight is a BI tool, but the underlying data storage and querying is incorrect.
  • D. Re-hosting Hadoop on EC2 and using Amazon EMR for processing means managing EC2 instances and EMR clusters, which is not fully serverless and can be less cost-effective for infrequently accessed data compared to an S3-based data lake with serverless query engines.

AWS Data Lake on S3

A centralized, curated, and secured repository that stores all your data, both structured and unstructured, at any scale, typically built on Amazon S3.

  • Scalable and cost-effective for petabyte-scale data storage.
  • Supports various data formats (structured, semi-structured, unstructured).
  • Integrates with serverless analytics services like Athena, Glue, and Redshift Spectrum.

Memory trick: From Hadoop's past to S3's lake, serverless queries we make.

More Design Solutions for Organizational Complexity questions