Microsoft Azure Data FundamentalsDescribe core data conceptsMedium

A data architect is designing a system to store large volumes of diverse data, including structured customer records, semi-structured log files, and unstructured sensor readings. The system needs to support schema-on-read capabilities and be highly scalable for future growth, without immediately enforcing a rigid structure. Which storage option is MOST appropriate?

  1. AData Lake
  2. BOnline Transaction Processing (OLTP) Database
  3. CRelational Database
  4. DData Warehouse
Show answer & explanation

Correct answer: A. Data Lake

A data lake is designed to store raw, diverse data (structured, semi-structured, unstructured) in its native format, offering schema-on-read flexibility and high scalability, which perfectly matches the requirements for storing varied data types without immediate rigid structure.

Why the other options are wrong

  • B. OLTP databases are for transactional workloads, not for storing large volumes of diverse raw data for analytical purposes.
  • C. Relational databases require a predefined schema and are not suitable for diverse, unstructured data.
  • D. Data warehouses are structured for analytical queries on cleaned, transformed data, not raw, diverse data with schema-on-read.

Data Lake

A centralized repository that stores a vast amount of raw data in its native format, including structured, semi-structured, and unstructured data.

  • Stores data without predefined schema (schema-on-read).
  • Highly scalable and cost-effective for large volumes.
  • Supports various data processing frameworks (e.g., Spark, Hadoop).
  • Used for big data analytics, machine learning, and data exploration.

Memory trick: A Data Lake is a big, raw pool where all your data can swim freely.

More Describe core data concepts questions