Professional Data EngineerOperationalizing machine learning modelsEasy
A financial institution is building a credit scoring model. They need to ensure that the data used for training the model is consistent, accurate, complete, and timely. Inconsistent data could lead to biased predictions, and missing data could reduce model performance significantly. They want to implement checks throughout their data pipeline, from ingestion to feature engineering, to maintain high data quality for ML. What is a key practice they should adopt to ensure data quality for their ML model?
- AStore all raw data indefinitely for auditing purposes.
- BRegularly retrain the model on new data.
- CUse a complex model architecture to compensate for data issues.
- DImplement schema validation and anomaly detection at each pipeline stage.
Show answer & explanationAnswer & explanation
Correct answer: D. Implement schema validation and anomaly detection at each pipeline stage.
Implementing schema validation and anomaly detection at each stage of the data pipeline is a fundamental practice for ensuring data quality, as it catches inconsistencies, errors, and missing values early, preventing them from propagating to the ML model.
Why the other options are wrong
- A. Storing raw data indefinitely is a data retention policy, not a direct practice for ensuring data quality, although it can aid in debugging data quality issues.
- B. Regularly retraining the model is good for addressing data drift but doesn't proactively ensure the quality of the incoming data itself.
- C. Using a complex model architecture to compensate for data issues is a reactive approach and often leads to models that are brittle and difficult to debug; it's better to fix data quality at the source.
Data Quality for ML
The practice of ensuring that data used in machine learning pipelines is accurate, complete, consistent, timely, and valid to produce reliable model performance.
- Critical for model performance and fairness.
- Involves validation, cleansing, and monitoring.
- Should be addressed at every stage of the data pipeline.
Memory trick: Validate and check data at every ML gate.