Microsoft Certified: Azure Solutions Architect ExpertDesign data storage solutionsMedium
A research institution is building a genomics analysis platform. The platform will store petabytes of raw genomic sequencing data (FASTQ files) which are large, unstructured, and accessed infrequently after initial processing. However, when accessed, high throughput is required for batch processing. The data must be highly durable and cost-optimized for long-term storage. Which Azure data storage solution is most suitable?
- AAzure Data Lake Storage Gen2
- BAzure SQL Database Hyperscale
- CAzure Premium File Shares
- DAzure Cosmos DB
Show answer & explanationAnswer & explanation
Correct answer: A. Azure Data Lake Storage Gen2
Azure Data Lake Storage Gen2 is designed for petabyte-scale analytics, handling large, unstructured files with high throughput for batch processing, and integrates well with analytical engines. It is highly durable and cost-effective for long-term storage of such data.
Why the other options are wrong
- B. Azure SQL Database Hyperscale is a relational database for transactional workloads, not suitable for storing petabytes of raw, unstructured genomic files.
- C. Azure Premium File Shares are for SMB/NFS file access and are not designed for petabyte-scale unstructured data analytics with high throughput for batch processing.
- D. Azure Cosmos DB is a NoSQL database for operational workloads, not optimized for petabyte-scale unstructured file storage and batch analytics.
Azure Data Lake Storage Gen2
A set of capabilities built on Azure Blob Storage dedicated to big data analytics. It combines the scalability and cost-effectiveness of object storage with a hierarchical file system.
- Petabyte-scale storage for big data analytics
- Hierarchical namespace for POSIX-compliant file system semantics
- Optimized for high-throughput batch processing
Memory trick: Data Lake Gen2: The 'Lake' for all your 'Gen'omic data.