Microsoft Azure Data FundamentalsDescribe core data conceptsMedium
A research team needs to store large datasets containing various types of scientific experimental results, including text documents, images, and sensor readings. The data is often unstructured or semi-structured, and the team requires a flexible storage solution that can scale to petabytes. Which data storage option is most appropriate for this scenario?
- AKey-Value Store
- BRelational Database
- CData Lake
- DGraph Database
Show answer & explanationAnswer & explanation
Correct answer: C. Data Lake
A data lake is designed to store vast amounts of raw data in its native format, regardless of structure, and can scale massively, making it ideal for diverse scientific experimental results.
Why the other options are wrong
- A. Key-value stores are suitable for simple, schema-less data but less for diverse, complex scientific results.
- B. Relational databases require a predefined schema and are less flexible for unstructured data.
- D. Graph databases are specialized for representing and querying relationships, not general-purpose storage for diverse data types.
Data Lake
A centralized repository that stores vast amounts of raw data in its native format, including structured, semi-structured, and unstructured data.
- Stores data in original format
- Massively scalable
- Supports various data types and analytics
Memory trick: Imagine a 'lake' where all kinds of data 'flow' in and settle.