AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisMedium
A healthcare startup is building a machine learning model to predict patient readmission rates. They are integrating patient data from multiple sources, and they notice that the 'diagnosis_code' field has various inconsistent formats (e.g., 'ICD-10-CM: J45.909', 'J45909', 'J45.909 Asthma'). This inconsistency will hinder model performance. Which data cleaning step should be prioritized to address this issue?
- AFeature scaling
- BStandardizing data formats
- CMissing value imputation
- DOutlier detection and removal
Show answer & explanationAnswer & explanation
Correct answer: B. Standardizing data formats
The problem explicitly states 'various inconsistent formats'. Standardizing data formats involves transforming data into a uniform, consistent structure, which is crucial before any further analysis or model training can accurately proceed with categorical data like diagnosis codes.
Why the other options are wrong
- A. Feature scaling is applied to numerical features to normalize their range, not to standardize categorical text formats.
- C. Missing value imputation addresses gaps in data, not variations in data format.
- D. Outlier detection is for extreme values in numerical data, not for inconsistent text formats.
Data Format Standardization
The process of transforming data from various inconsistent formats into a single, uniform, and consistent structure to ensure data quality and usability.
- Crucial for data integration from multiple sources.
- Ensures accurate comparisons and aggregations.
- Prevents errors in downstream analysis and model training.
Memory trick: Standardize: Make it all 'standard', like school uniforms.