Microsoft Azure Data FundamentalsDescribe how to work with non-relational data on AzureMedium
A data analytics team needs to store petabytes of raw sensor data from IoT devices for long-term analysis. The data will be ingested continuously and processed in large batches using Apache Spark. The solution requires a hierarchical namespace for efficient data organization and security, and compatibility with the Hadoop ecosystem. Which Azure non-relational data service is most appropriate for this scenario?
- AAzure Table Storage
- BAzure Data Lake Storage Gen2
- CAzure Cosmos DB
- DAzure Blob Storage
Show answer & explanationAnswer & explanation
Correct answer: B. Azure Data Lake Storage Gen2
Azure Data Lake Storage Gen2 is built on Azure Blob Storage and provides a hierarchical namespace, making it highly optimized for big data analytics workloads like those involving Apache Spark and petabytes of IoT sensor data. Its HDFS compatibility allows seamless integration with the Hadoop ecosystem.
Why the other options are wrong
- A. Azure Table Storage is a key-value store and is completely unsuitable for storing petabytes of raw sensor data for big data analytics.
- C. Azure Cosmos DB is a NoSQL database, not a data lake solution, and is not designed for storing petabytes of raw data for batch processing with Apache Spark.
- D. Azure Blob Storage is suitable for large amounts of unstructured data, but it lacks the hierarchical file system semantics and optimizations for big data analytics workloads that Data Lake Storage Gen2 offers.
Azure Data Lake Storage Gen2
Azure Data Lake Storage Gen2 is a set of capabilities dedicated to big data analytics, built on Azure Blob Storage. It provides file system semantics, file-level security, and optimized scale for petabytes of data.
- Combines features of Azure Data Lake Storage Gen1 with Azure Blob Storage.
- Offers a hierarchical namespace for folder and file management.
- HDFS compatible, enabling integration with Apache Hadoop and Spark.
- Optimized for big data analytics workloads, high throughput.
Memory trick: Think 'Data Lake Gen2: Deep Dive for Big Data, HDFS for Hadoop'.