Professional Data EngineerOperationalizing machine learning modelsMedium

A data engineering team is building an automated ML pipeline for a credit risk assessment model. The pipeline needs to ensure that the training data used for each retraining job is consistent and traceable. Specifically, they need to know exactly which version of the raw data and which preprocessing steps were applied to generate the final training dataset. This is crucial for debugging, auditing, and reproducing model results. Which concept is being addressed here?

  1. AData drift detection
  2. BData lineage
  3. CFeature engineering
  4. DModel monitoring
Show answer & explanation

Correct answer: B. Data lineage

Data lineage refers to the ability to track the origin, transformations, and movement of data over time. In an ML pipeline, understanding data lineage is critical for traceability, auditing, and ensuring reproducibility of training datasets and model results.

Why the other options are wrong

  • A. Data drift detection identifies changes in data distribution, but doesn't provide the full historical path and transformations of the data.
  • C. Feature engineering is the process of creating new features from raw data, but it doesn't inherently track the history or transformations of the data.
  • D. Model monitoring tracks the performance of a deployed model, not the history of the training data.

Data Lineage

The end-to-end lifecycle of data, tracking its origin, transformations, and movement over time, providing an audit trail for data quality, compliance, and reproducibility.

  • Provides a historical record of data transformations.
  • Crucial for data governance and compliance.
  • Enables reproducibility of ML training datasets.

Memory trick: Data lineage is your GPS for data, showing every step from raw to ML-ready.

More Operationalizing machine learning models questions