Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Medium

A data engineer is working with a large dataset (several terabytes) of customer interaction logs stored in a Fabric Lakehouse. The team needs to perform complex data cleansing, enrichment with external lookup tables, and aggregation before loading the data into a data warehouse. This process requires custom logic and highly optimized performance. Which Fabric tool is the most appropriate for these complex transformations?

  1. AData Pipelines with multiple sequential Copy Data activities.
  2. BKQL Querysets for in-warehouse transformations.
  3. CSpark Notebooks with PySpark or Scala.
  4. DDataflows Gen2 with advanced Power Query transformations.
Show answer & explanation

Correct answer: C. Spark Notebooks with PySpark or Scala.

Spark Notebooks, using PySpark or Scala, provide the most powerful and flexible environment for performing complex, large-scale data transformations. They leverage Spark's distributed processing capabilities, allowing for custom code, integration with various libraries, and efficient handling of multi-terabyte datasets for cleansing, enrichment, and aggregation.

Why the other options are wrong

  • A. Data Pipelines primarily orchestrate data movement; Copy Data activities are for transfer, not complex, custom transformations on this scale.
  • B. KQL Querysets are for querying and analyzing data already in a KQL database, not for performing large-scale, complex ETL on data originating from a Lakehouse into a data warehouse.
  • D. While Dataflows Gen2 are good for many transformations, they may struggle with the scale and complexity of custom logic required for 'several terabytes' and advanced enrichment/aggregation compared to Spark.

Spark for Large-Scale ETL

Spark Notebooks in Microsoft Fabric, utilizing PySpark or Scala, are the preferred tool for performing complex, large-scale ETL (Extract, Transform, Load) operations on multi-terabyte datasets, offering distributed processing, custom logic, and high performance.

  • Handles multi-terabyte datasets efficiently.
  • Supports custom code (PySpark, Scala).
  • Leverages distributed computing for performance.
  • Ideal for complex cleansing, enrichment, and aggregation.

Memory trick: Spark handles complex ETL at scale.

More Prepare and transform data (20-25%) questions