AWS Certified Solutions Architect – ProfessionalDesign for New SolutionsMedium

A research institution is building a new data platform to ingest genomic sequencing data, which consists of extremely large files (up to several terabytes each) that need to be accessed by high-performance computing (HPC) clusters for intensive analysis. The storage solution must provide very high throughput and low latency for concurrent reads from multiple compute nodes, support petabyte-scale storage, and be cost-effective for long-term retention. Which AWS storage service should be recommended?

  1. AAmazon S3 Standard
  2. BAmazon EBS Provisioned IOPS SSD (io2 Block Express)
  3. CAmazon EFS with Provisioned Throughput
  4. DAmazon FSx for Lustre
Show answer & explanation

Correct answer: D. Amazon FSx for Lustre

Amazon FSx for Lustre is specifically designed for HPC workloads, offering high-performance, parallel file system access with very high throughput and low latency, optimized for large files and concurrent access from compute clusters. It integrates well with S3 for long-term retention and cost-effectiveness.

Why the other options are wrong

  • A. Amazon S3 Standard is excellent for cost-effective object storage and scalability, but it does not provide the very high throughput and low-latency POSIX file system access required by HPC clusters for concurrent reads.
  • B. EBS io2 Block Express provides very high IOPS and throughput for single EC2 instances but is not a shared file system suitable for concurrent access from multiple HPC compute nodes or for petabyte-scale, multi-terabyte files.
  • C. Amazon EFS provides a scalable, shared file system but generally offers lower throughput and higher latency compared to FSx for Lustre, making it less suitable for the 'very high throughput and low latency' demands of HPC for multi-terabyte files.

HPC Storage on AWS

AWS storage solutions optimized for High-Performance Computing (HPC) workloads, characterized by requirements for very high throughput, low latency, parallel access from multiple compute nodes, and petabyte-scale storage for large files.

  • Requires very high throughput and low latency.
  • Supports parallel access from many compute nodes.
  • Handles extremely large files (TB scale).
  • Scales to petabytes and integrates with S3 for cost-effectiveness.

Memory trick: Lustre is the lightning-fast file system for genomic giants.

More Design for New Solutions questions