AWS Certified Solutions Architect – ProfessionalDesign Solutions for Organizational ComplexityMedium
A global telecommunications company has a vast amount of historical call detail record (CDR) data stored in an on-premises Hadoop cluster. This data is critical for regulatory compliance and advanced analytics, but it is accessed infrequently by data scientists and auditors. The company wants to migrate this petabyte-scale data to AWS to reduce on-premises infrastructure costs and enable more flexible, serverless analytics. The solution must be cost-effective for long-term storage and support ad-hoc querying. Which AWS solution should a Solutions Architect recommend?
- ATransfer data to Amazon S3 and build a data lake using AWS Glue and Amazon Athena.
- BMigrate data to Amazon DynamoDB and use DynamoDB Accelerator (DAX) for querying.
- CMigrate data to Amazon RDS for PostgreSQL and use Amazon QuickSight for analytics.
- DRe-host the Hadoop cluster on Amazon EC2 instances and use Amazon EMR for processing.
Show answer & explanationAnswer & explanation
Correct answer: A. Transfer data to Amazon S3 and build a data lake using AWS Glue and Amazon Athena.
Migrating Hadoop data to Amazon S3 creates a cost-effective data lake for petabyte-scale, infrequently accessed data. AWS Glue can be used for cataloging and transforming the data, while Amazon Athena provides serverless, ad-hoc querying capabilities directly on S3 data, perfectly aligning with the requirements for cost-effectiveness, long-term storage, and flexible analytics without managing servers.
Why the other options are wrong
- B. Amazon DynamoDB is a NoSQL database designed for high-performance, real-time applications with structured and semi-structured data, not for petabyte-scale, infrequently accessed historical data from a Hadoop cluster. DAX is a cache for DynamoDB, not a general analytics solution.
- C. Amazon RDS for PostgreSQL is a relational database, not suitable for petabyte-scale, infrequently accessed, semi-structured/unstructured data like CDRs from Hadoop. QuickSight is a BI tool, but the underlying data storage and querying is incorrect.
- D. Re-hosting Hadoop on EC2 and using Amazon EMR for processing means managing EC2 instances and EMR clusters, which is not fully serverless and can be less cost-effective for infrequently accessed data compared to an S3-based data lake with serverless query engines.
AWS Data Lake on S3
A centralized, curated, and secured repository that stores all your data, both structured and unstructured, at any scale, typically built on Amazon S3.
- Scalable and cost-effective for petabyte-scale data storage.
- Supports various data formats (structured, semi-structured, unstructured).
- Integrates with serverless analytics services like Athena, Glue, and Redshift Spectrum.
Memory trick: From Hadoop's past to S3's lake, serverless queries we make.