A global media company has a large archive of video content stored in an on-premises data center. They need to analyze this data for content categorization, metadata extraction, and compliance auditing using machine learning services in AWS. The data volume is in the petabytes, and the analysis process is compute-intensive. The company also wants to ensure that the data transfer to AWS is secure, reliable, and optimized for large datasets, minimizing network costs. Which solution should an architect recommend?
- AStream data in real-time to Amazon Kinesis Data Streams for processing with AWS Lambda.
- BTransfer data to Amazon S3 using a Site-to-Site VPN and store in Amazon S3 Glacier Deep Archive for cost savings.
- CUtilize AWS Snowball Edge devices for initial bulk data transfer, then use AWS Direct Connect for incremental updates to Amazon S3, and process with Amazon EMR.
- DMigrate data to Amazon S3 using AWS DataSync over the public internet and process with AWS Glue.
Show answer & explanationAnswer & explanation
Correct answer: C. Utilize AWS Snowball Edge devices for initial bulk data transfer, then use AWS Direct Connect for incremental updates to Amazon S3, and process with Amazon EMR.
AWS Snowball Edge is ideal for initial bulk migration of petabytes of data, overcoming network bandwidth limitations. AWS Direct Connect provides a dedicated, reliable, and secure link for ongoing incremental updates, optimizing network costs. Amazon S3 is the scalable and durable storage for the video content, and Amazon EMR is highly suitable for compute-intensive big data processing with machine learning, fulfilling all requirements.
Why the other options are wrong
- A. Streaming petabytes of archival video content in real-time to Kinesis Data Streams is not suitable for a large existing archive and is designed for real-time data, not bulk historical data.
- B. Site-to-Site VPN over the public internet would be too slow and unreliable for petabytes of data. S3 Glacier Deep Archive is for long-term archival with slow retrieval, not ideal for data that needs immediate access for compute-intensive analysis.
- D. Transferring petabytes over the public internet with DataSync might be slow and incur high egress costs for on-premises data. AWS Glue is for ETL, while EMR is typically better for large-scale, compute-intensive ML processing.
Hybrid Data Migration and Analytics with AWS
A strategy combining AWS Snowball Edge for initial bulk data transfer, AWS Direct Connect for ongoing incremental updates, Amazon S3 for scalable storage, and Amazon EMR for big data analytics and machine learning.
- Snowball Edge accelerates migration of petabytes of data without internet bandwidth constraints.
- Direct Connect provides a private, consistent, and cost-effective link for continuous data flow.
- Amazon S3 offers highly durable, scalable, and cost-effective object storage.
- Amazon EMR is a managed cluster platform for running big data frameworks like Apache Spark and Hadoop, ideal for ML analytics.
Memory trick: Snowball starts the move, Direct Connect keeps it flowing, S3 stores it, EMR analyzes it.