Microsoft Certified: Azure Solutions Architect ExpertDesign data storage solutionsMedium

A telecommunications company needs to store petabytes of call detail records (CDRs) for billing, fraud detection, and historical analysis. These records are generated continuously throughout the day and need to be available for querying by various analytics engines, including Apache Spark and Hive, which rely on HDFS compatibility. The solution must be highly scalable and cost-effective for large volumes of semi-structured data. Which Azure data storage solution is most suitable?

  1. AAzure Managed Instance for Apache Cassandra
  2. BAzure Synapse Analytics Dedicated SQL Pool
  3. CAzure Data Lake Storage Gen2
  4. DAzure SQL Database Hyperscale
Show answer & explanation

Correct answer: C. Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 is designed for storing petabytes of data for big data analytics. Its HDFS compatibility allows direct integration with Spark and Hive, and its hierarchical namespace and cost-effectiveness make it ideal for large volumes of semi-structured CDRs for various analytical workloads.

Why the other options are wrong

  • A. Azure Managed Instance for Apache Cassandra is a NoSQL wide-column store, not typically used for petabyte-scale semi-structured files that require HDFS compatibility for Spark/Hive.
  • B. Azure Synapse Analytics Dedicated SQL Pool is a data warehousing solution for structured data and complex SQL queries, not primarily for storing raw, semi-structured files with HDFS compatibility for data lake analytics.
  • D. Azure SQL Database Hyperscale is a relational database for transactional workloads, not optimized for petabyte-scale semi-structured files and HDFS compatibility for big data analytics engines.

Azure Data Lake Storage Gen2 (HDFS Compatibility)

Azure Data Lake Storage Gen2 (ADLS Gen2) builds on Azure Blob Storage and is optimized for big data analytics. Its HDFS compatibility means it can be directly accessed by tools and frameworks that use the Hadoop Distributed File System (HDFS) API, such as Apache Spark, Hadoop, and Hive.

  • Petabyte-scale storage for big data
  • HDFS-compatible endpoint for analytics engines
  • Hierarchical namespace for file system semantics
  • Cost-effective with Blob Storage tiers
  • Ideal for semi-structured and unstructured data for analytics

Memory trick: CDRs in Data Lake Gen2, HDFS powers analytics.

More Design data storage solutions questions