CompTIA Data+ (DA0-002)Data MiningMedium

A data architect is designing a new data pipeline for an e-commerce platform. The platform generates a high volume of clickstream data, which needs to be processed quickly and made available for near real-time analytics. The raw data will be stored in a data lake, and subsequent transformations, aggregations, and data quality checks will be performed within the data lake environment using distributed processing frameworks. Which data integration paradigm best fits this approach?

  1. ABatch ETL
  2. BData Virtualization
  3. CELT
  4. DMaster Data Management (MDM)
Show answer & explanation

Correct answer: C. ELT

ELT (Extract, Load, Transform) is designed for scenarios involving high volumes of raw data that need to be loaded quickly into a data lake/warehouse, with transformations performed post-load leveraging the target system's processing power. This aligns perfectly with the requirements of clickstream data and near real-time analytics.

Why the other options are wrong

  • A. Batch ETL involves transformations before loading, which might introduce latency for near real-time processing of high-volume data.
  • B. Data Virtualization provides a unified view of disparate data without physical movement, not a primary method for ingesting and transforming raw data at scale.
  • D. Master Data Management (MDM) focuses on creating a single, consistent view of core business entities, not an integration paradigm for raw data ingestion and transformation.

ELT Paradigm

A data integration approach where data is first extracted from sources, loaded directly into a target data store (like a data lake), and then transformed within that same data store.

  • Optimized for cloud and big data environments.
  • Allows for faster initial data loading.
  • Leverages the scalability and processing power of the target system.

Memory trick: Integration choices shape data flow.

More Data Mining questions