Microsoft Azure Data FundamentalsDescribe how to work with non-relational data on AzureMedium

A data analytics team needs to store petabytes of raw sensor data from IoT devices for long-term retention and future big data processing using tools like Apache Spark. The solution must support a hierarchical namespace for efficient organization and access control. Which Azure non-relational data service is best suited for this purpose?

  1. AAzure Blob Storage (Standard)
  2. BAzure Table Storage
  3. CAzure Data Lake Storage Gen2
  4. DAzure Cosmos DB
Show answer & explanation

Correct answer: C. Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 is built on Azure Blob Storage but adds a hierarchical namespace and file system semantics, making it ideal for large-scale analytics workloads with petabytes of data from IoT devices and integration with big data tools like Apache Spark.

Why the other options are wrong

  • A. Azure Blob Storage (Standard) lacks the hierarchical namespace and optimized file system semantics needed for efficient big data analytics.
  • B. Azure Table Storage is a key-value store, not suitable for storing large files of raw sensor data for big data processing.
  • D. Azure Cosmos DB is a transactional database, not primarily designed for petabyte-scale raw file storage and big data analytics.

Azure Data Lake Storage Gen2

A set of capabilities built on Azure Blob Storage, optimized for big data analytics workloads, offering a hierarchical file system and petabyte-scale storage.

  • Provides a hierarchical namespace for organizing data.
  • Optimized for high-throughput, low-latency access for analytics.
  • Compatible with Hadoop Distributed File System (HDFS) interfaces.

Memory trick: Data Lake: deep storage for big data processing.

More Describe how to work with non-relational data on Azure questions