Microsoft Azure Data FundamentalsDescribe how to work with non-relational data on AzureMedium
A data engineering team needs to store petabytes of raw sensor data from IoT devices for long-term analysis. The data will be ingested continuously and processed using big data analytics frameworks like Apache Spark. A hierarchical namespace is required for efficient organization and access control. Which Azure non-relational data service is most suitable?
- AAzure Queue Storage
- BAzure Cosmos DB
- CAzure Data Lake Storage Gen2
- DAzure Table Storage
Show answer & explanationAnswer & explanation
Correct answer: C. Azure Data Lake Storage Gen2
Azure Data Lake Storage Gen2 is built on Azure Blob Storage and provides a hierarchical file system, making it ideal for big data analytics workloads. It supports petabyte-scale data, continuous ingestion, and integrates well with frameworks like Apache Spark.
Why the other options are wrong
- A. Azure Queue Storage is for message queuing, not for long-term storage of petabytes of raw data.
- B. Azure Cosmos DB is a transactional NoSQL database, not optimized for petabyte-scale raw data storage for batch analytics.
- D. Azure Table Storage is a simple key-value store, not suitable for petabyte-scale, hierarchical file system needs of big data analytics.
Azure Data Lake Storage Gen2
Azure Data Lake Storage Gen2 is a set of capabilities built on Azure Blob Storage dedicated to big data analytics. It combines the scalability of object storage with the semantics of a file system.
- Optimized for big data analytics workloads.
- Provides a hierarchical namespace for file system-like access.
- Petabyte-scale storage capacity.
- Integrates with Apache Spark, Hadoop, and other analytics engines.
Memory trick: Data Lake: A vast, organized reservoir for all your data analytics.