AWS Certified Machine Learning – SpecialtyData EngineeringMedium

A financial institution is building a fraud detection system. They collect transaction data from various sources, including real-time payment gateways and batch processing systems. Due to regulatory compliance and the sensitive nature of financial data, all data must be encrypted both in transit and at rest, and access must be strictly controlled. Which AWS data storage solution is most suitable for storing this highly sensitive, transactional data, ensuring both security and scalability for ML model training?

  1. AAmazon Redshift
  2. BAmazon DynamoDB
  3. CAmazon RDS for PostgreSQL
  4. DAmazon S3
Show answer & explanation

Correct answer: D. Amazon S3

Amazon S3 offers high scalability, durability, and robust security features required for sensitive data. It supports encryption at rest (SSE-S3, SSE-KMS, SSE-C) and in transit (SSL/TLS). S3 is a cost-effective solution for storing large volumes of raw and processed data used for ML training, and it integrates well with other AWS services for data processing and model training.

Why the other options are wrong

  • A. Amazon Redshift is a petabyte-scale data warehouse, optimized for analytical queries on structured data. While secure, it's primarily for aggregated analysis, not the raw, diverse data storage needs of ML training datasets.
  • B. Amazon DynamoDB is a NoSQL database, excellent for high-performance key-value and document workloads, but typically used for operational data, not for large-scale, cost-effective storage of raw ML training datasets.
  • C. Amazon RDS is a relational database service, suitable for structured transactional data, but it's not optimized for the massive scale and cost-effectiveness of raw data storage for ML training compared to S3.

Amazon S3 for ML Data Lakes

Amazon S3 (Simple Storage Service) is an object storage service offering industry-leading scalability, data availability, security, and performance. It's often used as the foundation for data lakes for machine learning.

  • Stores any type of object (data lake foundation).
  • High durability (11 nines) and availability.
  • Encryption at rest (SSE-S3, SSE-KMS, SSE-C) and in transit (SSL/TLS).
  • Highly scalable, cost-effective for large datasets.
  • Integrates with numerous AWS ML and data processing services.

Memory trick: Securely store your ML treasures in the scalable S3 cloud.

More Data Engineering questions