Microsoft Azure Data FundamentalsDescribe how to work with non-relational data on AzureMedium

A data analytics team needs to store petabytes of raw sensor data from IoT devices for long-term analysis. The data will be ingested in real-time, often consisting of small, individual files. The team requires a cost-effective storage solution that supports a hierarchical namespace for organization and is optimized for analytical workloads, including integration with Spark and Hadoop. Which Azure service is the most appropriate choice?

  1. AAzure Blob Storage (Standard)
  2. BAzure Table Storage
  3. CAzure Data Lake Storage Gen2
  4. DAzure Cosmos DB
Show answer & explanation

Correct answer: C. Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 is built on Azure Blob Storage but adds a hierarchical namespace and is optimized for big data analytics workloads. It provides file system semantics, file-level security, and is highly scalable and cost-effective for petabytes of data, making it ideal for IoT sensor data analysis with Spark and Hadoop integration.

Why the other options are wrong

  • A. Azure Blob Storage (Standard) lacks the hierarchical namespace and optimizations for analytical workloads like ADLS Gen2.
  • B. Azure Table Storage is a key-value store, not suitable for file-based big data analytics or a hierarchical namespace.
  • D. Azure Cosmos DB is a transactional NoSQL database, not designed for petabytes of raw files for big data analytics.

Azure Data Lake Storage Gen2

A highly scalable and cost-effective data lake solution built on Azure Blob Storage, optimized for big data analytics.

  • Provides a hierarchical namespace for file system semantics.
  • Offers file-level security and POSIX-compliant ACLs.
  • Optimized for Spark, Hadoop, and other big data analytics engines.
  • Supports petabytes of data with high throughput.

Memory trick: For 'big lakes' of data and 'analytics', think Data Lake Storage.

More Describe how to work with non-relational data on Azure questions