Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Medium
A retail company collects customer feedback data through various channels, resulting in semi-structured JSON files stored in Azure Data Lake Storage Gen2. Before this data can be analyzed, it needs to be parsed, flattened, and certain PII (Personally Identifiable Information) fields must be masked. The analytics engineering team prefers to use a code-first approach with robust data manipulation capabilities. Which Microsoft Fabric tool is best suited for these requirements?
- ADataflows Gen2
- BKusto Query Language (KQL) in a KQL queryset
- CSpark notebook (PySpark)
- DData Pipelines
Show answer & explanationAnswer & explanation
Correct answer: C. Spark notebook (PySpark)
Spark notebooks provide a powerful, code-first environment (PySpark) for complex data transformations like parsing semi-structured JSON, flattening data, and applying custom masking logic.
Why the other options are wrong
- A. While Dataflows Gen2 can handle JSON, complex parsing, flattening of deeply nested structures, and custom masking are often more efficiently and flexibly handled with code.
- B. KQL is for querying data in Kusto databases and is not the primary tool for ingesting and transforming semi-structured JSON files from ADLS Gen2 in this manner.
- D. Data Pipelines are primarily for orchestration and data movement, not for complex, code-driven data transformations like advanced JSON processing or PII masking.
Spark Notebooks
Interactive development environments within Microsoft Fabric for data processing using Apache Spark with languages like PySpark, Scala, or C#.
- Offers high flexibility for complex transformations.
- Scalable for large datasets.
- Supports various data formats and operations.
Memory trick: Spark code sculpts raw data into analyzed perfection.