Microsoft Certified: Azure Solutions Architect ExpertDesign data storage solutionsHard

A global pharmaceutical company is conducting a large-scale genomics research project. The project involves storing petabytes of raw genomic sequencing data, which are typically very large files (hundreds of GBs to TBs each). This data needs to be accessible for complex analytics workloads using Apache Spark and Hadoop ecosystems. The solution must provide hierarchical namespace capabilities and strong consistency. Which Azure storage solution should be recommended?

  1. AAzure Cosmos DB
  2. BAzure Blob Storage (General-purpose v2)
  3. CAzure Data Lake Storage Gen2
  4. DAzure SQL Database Hyperscale
Show answer & explanation

Correct answer: C. Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 is specifically designed for big data analytics workloads, offering petabyte-scale storage for large files. Its hierarchical namespace is crucial for efficient data organization and access for Spark/Hadoop, and it provides strong consistency with atomicity for file operations.

Why the other options are wrong

  • A. Azure Cosmos DB is a globally distributed NoSQL database, suitable for structured/semi-structured transactional data, not for raw, unstructured large files for data lake analytics.
  • B. While Azure Blob Storage can store large files, it lacks a true hierarchical namespace and is not optimized for the specific demands of big data analytics engines like Spark/Hadoop, which benefit from file system semantics.
  • D. Azure SQL Database Hyperscale is a relational database optimized for transactional workloads and petabyte-scale relational data, not for storing raw, unstructured large files for big data analytics.

Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 (ADLS Gen2) is a set of capabilities built on Azure Blob Storage, dedicated to big data analytics. It offers a hierarchical file system (HDFS compatible), petabyte-scale data storage, and optimized performance for analytics engines like Apache Spark and Hadoop.

  • Built on Azure Blob Storage, optimized for analytics
  • Hierarchical namespace for folder/file organization
  • HDFS compatibility for existing big data tools
  • Petabyte-scale storage for large files
  • Atomic file operations and strong consistency

Memory trick: Genomic treasure in Data Lake Gen2, analytics flow.

More Design data storage solutions questions