Microsoft Azure Data FundamentalsDescribe how to work with non-relational data on AzureMedium

A data engineering team needs to store petabytes of raw sensor data from IoT devices for long-term analysis. The data will be ingested continuously and processed using big data analytics frameworks like Apache Spark. A hierarchical namespace is required for efficient organization and access control. Which Azure non-relational data service is most suitable?

  1. AAzure Queue Storage
  2. BAzure Cosmos DB
  3. CAzure Data Lake Storage Gen2
  4. DAzure Table Storage
Show answer & explanation

Correct answer: C. Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 is built on Azure Blob Storage and provides a hierarchical file system, making it ideal for big data analytics workloads. It supports petabyte-scale data, continuous ingestion, and integrates well with frameworks like Apache Spark.

Why the other options are wrong

  • A. Azure Queue Storage is for message queuing, not for long-term storage of petabytes of raw data.
  • B. Azure Cosmos DB is a transactional NoSQL database, not optimized for petabyte-scale raw data storage for batch analytics.
  • D. Azure Table Storage is a simple key-value store, not suitable for petabyte-scale, hierarchical file system needs of big data analytics.

Azure Data Lake Storage Gen2

Azure Data Lake Storage Gen2 is a set of capabilities built on Azure Blob Storage dedicated to big data analytics. It combines the scalability of object storage with the semantics of a file system.

  • Optimized for big data analytics workloads.
  • Provides a hierarchical namespace for file system-like access.
  • Petabyte-scale storage capacity.
  • Integrates with Apache Spark, Hadoop, and other analytics engines.

Memory trick: Data Lake: A vast, organized reservoir for all your data analytics.

More Describe how to work with non-relational data on Azure questions