Microsoft Azure Data FundamentalsDescribe core data conceptsMedium
A research institution is collecting data from various scientific experiments. The data includes structured results in CSV files, unstructured lab notes in PDF documents, and semi-structured metadata in JSON files. They need a centralized repository that can store all these diverse data types without requiring a predefined schema, and then make it available for future analysis by different tools. Which storage option is best suited for this requirement?
- ARelational Database
- BData Warehouse
- CData Lake
- DNoSQL Database
Show answer & explanationAnswer & explanation
Correct answer: C. Data Lake
The scenario describes storing diverse data types (structured, unstructured, semi-structured) without a predefined schema, and making it available for future analysis. This is the core definition and purpose of a data lake, which stores raw data in its native format.
Why the other options are wrong
- A. A Relational Database requires a predefined schema and is best for structured data, not diverse types without schema.
- B. A Data Warehouse stores structured, cleaned, and transformed data for analytical purposes, not raw, diverse formats.
- D. A NoSQL Database offers schema flexibility but usually focuses on specific data models (key-value, document, graph) and isn't typically a 'centralized repository' for all raw, diverse file types.
Data Lake
A centralized repository that allows you to store all your data, structured and unstructured, at any scale. It stores data in its native format without requiring a predefined schema.
- Stores raw data in its native format.
- Supports structured, semi-structured, and unstructured data.
- Designed for 'schema-on-read' and future analysis.
Memory trick: Lake of Data: Store everything raw, then decide how to cook it.