Microsoft Azure Data FundamentalsDescribe how to work with non-relational data on AzureMedium
A data analytics team needs to store petabytes of raw sensor data from IoT devices for long-term retention and future big data processing using tools like Apache Spark. The solution must support a hierarchical namespace for efficient organization and access control. Which Azure non-relational data service is best suited for this purpose?
- AAzure Blob Storage (Standard)
- BAzure Table Storage
- CAzure Data Lake Storage Gen2
- DAzure Cosmos DB
Show answer & explanationAnswer & explanation
Correct answer: C. Azure Data Lake Storage Gen2
Azure Data Lake Storage Gen2 is built on Azure Blob Storage but adds a hierarchical namespace and file system semantics, making it ideal for large-scale analytics workloads with petabytes of data from IoT devices and integration with big data tools like Apache Spark.
Why the other options are wrong
- A. Azure Blob Storage (Standard) lacks the hierarchical namespace and optimized file system semantics needed for efficient big data analytics.
- B. Azure Table Storage is a key-value store, not suitable for storing large files of raw sensor data for big data processing.
- D. Azure Cosmos DB is a transactional database, not primarily designed for petabyte-scale raw file storage and big data analytics.
Azure Data Lake Storage Gen2
A set of capabilities built on Azure Blob Storage, optimized for big data analytics workloads, offering a hierarchical file system and petabyte-scale storage.
- Provides a hierarchical namespace for organizing data.
- Optimized for high-throughput, low-latency access for analytics.
- Compatible with Hadoop Distributed File System (HDFS) interfaces.
Memory trick: Data Lake: deep storage for big data processing.