AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisMedium
A data engineering team is integrating data from various sources into a unified data lake. They encounter a 'CustomerID' field that is represented as an integer in one source (e.g., 12345) and as a string with leading zeros in another (e.g., '0012345'). Both represent the same unique identifier. To ensure data consistency and enable proper joins and analysis, which data cleaning operation should be applied?
- AOne-hot encoding
- BBinning
- CType conversion and string manipulation
- DFeature scaling
Show answer & explanationAnswer & explanation
Correct answer: C. Type conversion and string manipulation
To reconcile 'CustomerID' across sources, the data types must be consistent. This involves converting the string representation to an integer and potentially removing leading zeros, or vice-versa, to ensure both sources represent the same ID identically. This falls under type conversion and string manipulation.
Why the other options are wrong
- A. One-hot encoding converts categorical data into a numerical format for machine learning, not for ID consistency.
- B. Binning groups continuous numerical data into discrete bins, which is irrelevant for customer IDs.
- D. Feature scaling (normalization/standardization) adjusts numerical feature ranges, which is not applicable here.
Data Type and Format Consistency
Ensuring that data representing the same entity or concept across different sources or within a dataset has uniform data types, formats, and representations.
- Critical for accurate joins, comparisons, and analysis.
- Involves type conversion (e.g., string to int, date parsing).
- May require string manipulation (e.g., trimming, padding, regex).
- Resolves heterogeneity across disparate data sources.
Memory trick: Consistent data connects, converting types corrects.