Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Medium
A data engineer is working with a large dataset (several terabytes) of customer interaction logs stored in a Fabric Lakehouse. The team needs to perform complex data cleansing, enrichment with external lookup tables, and aggregation before loading the data into a data warehouse. This process requires custom logic and highly optimized performance. Which Fabric tool is the most appropriate for these complex transformations?
- AData Pipelines with multiple sequential Copy Data activities.
- BKQL Querysets for in-warehouse transformations.
- CSpark Notebooks with PySpark or Scala.
- DDataflows Gen2 with advanced Power Query transformations.
Show answer & explanationAnswer & explanation
Correct answer: C. Spark Notebooks with PySpark or Scala.
Spark Notebooks, using PySpark or Scala, provide the most powerful and flexible environment for performing complex, large-scale data transformations. They leverage Spark's distributed processing capabilities, allowing for custom code, integration with various libraries, and efficient handling of multi-terabyte datasets for cleansing, enrichment, and aggregation.
Why the other options are wrong
- A. Data Pipelines primarily orchestrate data movement; Copy Data activities are for transfer, not complex, custom transformations on this scale.
- B. KQL Querysets are for querying and analyzing data already in a KQL database, not for performing large-scale, complex ETL on data originating from a Lakehouse into a data warehouse.
- D. While Dataflows Gen2 are good for many transformations, they may struggle with the scale and complexity of custom logic required for 'several terabytes' and advanced enrichment/aggregation compared to Spark.
Spark for Large-Scale ETL
Spark Notebooks in Microsoft Fabric, utilizing PySpark or Scala, are the preferred tool for performing complex, large-scale ETL (Extract, Transform, Load) operations on multi-terabyte datasets, offering distributed processing, custom logic, and high performance.
- Handles multi-terabyte datasets efficiently.
- Supports custom code (PySpark, Scala).
- Leverages distributed computing for performance.
- Ideal for complex cleansing, enrichment, and aggregation.
Memory trick: Spark handles complex ETL at scale.