CompTIA Data+ (DA0-002)Data MiningHard
A data engineer is tasked with building a robust data pipeline that sources customer data from multiple regional databases. Each database uses a slightly different identifier for customers (e.g., `customer_id`, `cust_num`, `client_id`). To create a unified customer view, the engineer needs to convert these various identifiers into a single, consistent `global_customer_id` format. This process also involves mapping legacy codes to new standardized codes. Which data transformation technique is being primarily applied here?
- AAggregation
- BStandardization and mapping
- CFiltering
- DPivoting
Show answer & explanationAnswer & explanation
Correct answer: B. Standardization and mapping
Standardization involves ensuring that data conforms to a unified format or convention. Mapping, in this context, is the process of translating values from one system or format to another (e.g., legacy IDs to new global IDs). Together, these techniques achieve the goal of a single, consistent identifier format.
Why the other options are wrong
- A. Aggregation involves summarizing data (e.g., summing, averaging), not transforming identifiers into a consistent format.
- C. Filtering selects subsets of data based on criteria, but does not change the format or map identifiers.
- D. Pivoting reorganizes data from rows to columns or vice versa, which is not relevant to unifying customer identifiers.
Data Mapping
The process of creating a correspondence between data elements from a source system to target data elements in another system, often involving conversion or transformation rules.
- Essential for data integration and migration.
- Defines how data fields relate and convert.
- Can include value-level transformations and data type conversions.
Memory trick: Transform data: shape it, combine it, or make it fit just right.