Microsoft Azure Data FundamentalsDescribe core data conceptsMedium
A data architect is designing a system to store large volumes of diverse data, including structured customer records, semi-structured log files, and unstructured sensor readings. The system needs to support schema-on-read capabilities and be highly scalable for future growth, without immediately enforcing a rigid structure. Which storage option is MOST appropriate?
- AData Lake
- BOnline Transaction Processing (OLTP) Database
- CRelational Database
- DData Warehouse
Show answer & explanationAnswer & explanation
Correct answer: A. Data Lake
A data lake is designed to store raw, diverse data (structured, semi-structured, unstructured) in its native format, offering schema-on-read flexibility and high scalability, which perfectly matches the requirements for storing varied data types without immediate rigid structure.
Why the other options are wrong
- B. OLTP databases are for transactional workloads, not for storing large volumes of diverse raw data for analytical purposes.
- C. Relational databases require a predefined schema and are not suitable for diverse, unstructured data.
- D. Data warehouses are structured for analytical queries on cleaned, transformed data, not raw, diverse data with schema-on-read.
Data Lake
A centralized repository that stores a vast amount of raw data in its native format, including structured, semi-structured, and unstructured data.
- Stores data without predefined schema (schema-on-read).
- Highly scalable and cost-effective for large volumes.
- Supports various data processing frameworks (e.g., Spark, Hadoop).
- Used for big data analytics, machine learning, and data exploration.
Memory trick: A Data Lake is a big, raw pool where all your data can swim freely.