AWS Certified Solutions Architect – ProfessionalDesign for New SolutionsHard

A research institution is building a new data platform to ingest genomic sequencing data, which involves petabytes of large, unstructured files (e.g., FASTQ, BAM files). This data needs to be stored cost-effectively, accessed frequently for high-performance computing (HPC) analysis with low latency, and capable of being shared securely with collaborators globally. The solution must also support data lifecycle management. Which storage solution should the architect recommend?

  1. AAmazon S3 Intelligent-Tiering for storage, Amazon FSx for Lustre for HPC access, and Amazon S3 Replication for global sharing.
  2. BAmazon S3 Standard for storage, Amazon EFS for HPC access, and AWS Transfer Family for secure sharing.
  3. CAmazon S3 Standard-IA for storage, Amazon FSx for Windows File Server for HPC access, and Amazon CloudFront for global sharing.
  4. DAmazon S3 Glacier Deep Archive for cost-effective storage, Amazon EBS for HPC access, and AWS Direct Connect for secure sharing.
Show answer & explanation

Correct answer: A. Amazon S3 Intelligent-Tiering for storage, Amazon FSx for Lustre for HPC access, and Amazon S3 Replication for global sharing.

This solution provides optimal performance for HPC with FSx for Lustre, cost-effective and intelligent storage with S3 Intelligent-Tiering for large unstructured files, and secure global sharing via S3 Replication. This combination addresses all requirements for petabyte-scale, high-performance, cost-effective, and globally sharable genomic data.

Why the other options are wrong

  • B. EFS is a shared file system but not designed for the extreme performance demands of HPC. S3 Standard is not the most cost-effective for petabytes with varying access patterns.
  • C. S3 Standard-IA is for infrequently accessed data, not necessarily the most cost-effective for petabytes with varying access. FSx for Windows File Server is not optimized for Linux-based HPC workloads. CloudFront is a CDN, not for secure file sharing of raw data.
  • D. S3 Glacier Deep Archive is for archival data with high retrieval times, not for frequent HPC access with low latency. EBS is block storage for EC2 instances, not a scalable shared file system for petabytes of data.

HPC Storage on AWS

AWS solutions for storing and accessing large datasets for High-Performance Computing (HPC) workloads, focusing on throughput, latency, cost-effectiveness, and scalability.

  • Uses S3 for object storage and data lakes.
  • FSx for Lustre provides high-performance file systems for HPC.
  • Intelligent-Tiering optimizes S3 costs.
  • S3 Replication for global data sharing.

Memory trick: S3 Intelligent Lustre for Global Genomic Gold.

More Design for New Solutions questions