AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisMedium

A healthcare startup is building a machine learning model to predict patient readmission rates. They are integrating patient data from multiple sources, and they notice that the 'diagnosis_code' field has various inconsistent formats (e.g., 'ICD-10-CM: J45.909', 'J45909', 'J45.909 Asthma'). This inconsistency will hinder model performance. Which data cleaning step should be prioritized to address this issue?

  1. AFeature scaling
  2. BStandardizing data formats
  3. CMissing value imputation
  4. DOutlier detection and removal
Show answer & explanation

Correct answer: B. Standardizing data formats

The problem explicitly states 'various inconsistent formats'. Standardizing data formats involves transforming data into a uniform, consistent structure, which is crucial before any further analysis or model training can accurately proceed with categorical data like diagnosis codes.

Why the other options are wrong

  • A. Feature scaling is applied to numerical features to normalize their range, not to standardize categorical text formats.
  • C. Missing value imputation addresses gaps in data, not variations in data format.
  • D. Outlier detection is for extreme values in numerical data, not for inconsistent text formats.

Data Format Standardization

The process of transforming data from various inconsistent formats into a single, uniform, and consistent structure to ensure data quality and usability.

  • Crucial for data integration from multiple sources.
  • Ensures accurate comparisons and aggregations.
  • Prevents errors in downstream analysis and model training.

Memory trick: Standardize: Make it all 'standard', like school uniforms.

More Exploratory Data Analysis questions