Microsoft Azure Data FundamentalsDescribe core data conceptsMedium

A research institution is collecting data from various scientific experiments. The data includes structured results in CSV files, unstructured lab notes in PDF documents, and semi-structured metadata in JSON files. They need a centralized repository that can store all these diverse data types without requiring a predefined schema, and then make it available for future analysis by different tools. Which storage option is best suited for this requirement?

  1. ARelational Database
  2. BData Warehouse
  3. CData Lake
  4. DNoSQL Database
Show answer & explanation

Correct answer: C. Data Lake

The scenario describes storing diverse data types (structured, unstructured, semi-structured) without a predefined schema, and making it available for future analysis. This is the core definition and purpose of a data lake, which stores raw data in its native format.

Why the other options are wrong

  • A. A Relational Database requires a predefined schema and is best for structured data, not diverse types without schema.
  • B. A Data Warehouse stores structured, cleaned, and transformed data for analytical purposes, not raw, diverse formats.
  • D. A NoSQL Database offers schema flexibility but usually focuses on specific data models (key-value, document, graph) and isn't typically a 'centralized repository' for all raw, diverse file types.

Data Lake

A centralized repository that allows you to store all your data, structured and unstructured, at any scale. It stores data in its native format without requiring a predefined schema.

  • Stores raw data in its native format.
  • Supports structured, semi-structured, and unstructured data.
  • Designed for 'schema-on-read' and future analysis.

Memory trick: Lake of Data: Store everything raw, then decide how to cook it.

More Describe core data concepts questions