CompTIA Data+ (DA0-002)Data MiningMedium
A data architect is designing a new data pipeline for an e-commerce platform. The platform generates a high volume of clickstream data, which needs to be processed quickly and made available for near real-time analytics. The raw data will be stored in a data lake, and subsequent transformations, aggregations, and data quality checks will be performed within the data lake environment using distributed processing frameworks. Which data integration paradigm best fits this approach?
- ABatch ETL
- BData Virtualization
- CELT
- DMaster Data Management (MDM)
Show answer & explanationAnswer & explanation
Correct answer: C. ELT
ELT (Extract, Load, Transform) is designed for scenarios involving high volumes of raw data that need to be loaded quickly into a data lake/warehouse, with transformations performed post-load leveraging the target system's processing power. This aligns perfectly with the requirements of clickstream data and near real-time analytics.
Why the other options are wrong
- A. Batch ETL involves transformations before loading, which might introduce latency for near real-time processing of high-volume data.
- B. Data Virtualization provides a unified view of disparate data without physical movement, not a primary method for ingesting and transforming raw data at scale.
- D. Master Data Management (MDM) focuses on creating a single, consistent view of core business entities, not an integration paradigm for raw data ingestion and transformation.
ELT Paradigm
A data integration approach where data is first extracted from sources, loaded directly into a target data store (like a data lake), and then transformed within that same data store.
- Optimized for cloud and big data environments.
- Allows for faster initial data loading.
- Leverages the scalability and processing power of the target system.
Memory trick: Integration choices shape data flow.