A research institution is building a new data platform to ingest genomic sequencing data, which consists of large files (up to several terabytes each) and associated metadata. The platform needs to store petabytes of this data, make it available for scientific analysis (often involving parallel processing across many compute nodes), and ensure long-term archival for regulatory compliance. Data access patterns are primarily sequential reads for analysis, with infrequent random access. The solution must be highly available, durable, and cost-optimized for storage. Which AWS storage solution is MOST suitable?
- AAmazon S3 Glacier Deep Archive for long-term storage and Amazon S3 Standard for active analysis.
- BAmazon FSx for Lustre for high-performance computing workloads and Amazon S3 for archival.
- CAmazon EBS Provisioned IOPS volumes attached to Amazon EC2 instances.
- DAmazon EFS for shared file storage across multiple compute nodes.
Show answer & explanationAnswer & explanation
Correct answer: B. Amazon FSx for Lustre for high-performance computing workloads and Amazon S3 for archival.
Amazon FSx for Lustre is a high-performance file system optimized for HPC workloads, providing very low-latency access and high throughput for petabyte-scale datasets, which is ideal for genomic analysis involving parallel processing and sequential reads of large files. It integrates directly with Amazon S3, allowing data to be cost-effectively stored in S3 and then imported into FSx for Lustre for analysis, and exported back for durability and archival. This combination meets the requirements for performance, scalability, durability, and cost-optimization for both active analysis and long-term archival.
Why the other options are wrong
- A. S3 Glacier Deep Archive is for long-term archival with retrieval times in hours, not suitable for active scientific analysis. While S3 Standard can store data, it doesn't provide a high-performance POSIX-compliant file system layer optimized for HPC directly.
- C. EBS volumes are block storage, good for single EC2 instances but not for shared access across many compute nodes or cost-effective petabyte-scale archival.
- D. EFS provides shared file storage but may not offer the same ultra-high throughput and low latency as FSx for Lustre for HPC workloads involving multi-terabyte files and parallel processing.
HPC Storage on AWS
Storage solutions designed for high-performance computing (HPC) workloads that require extremely high throughput, low latency, and parallel access to large datasets, often integrated with object storage for cost-effective archival.
- FSx for Lustre is optimized for HPC file systems.
- S3 provides durable, scalable, and cost-effective object storage.
- Integration allows data to be moved between high-performance file systems and object storage.
Memory trick: FSx for Lustre speeds analysis, S3 archives the genome's vast data.