Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Medium

A data team needs to transform a large dataset (several terabytes) of customer interaction logs stored in a Lakehouse. The transformation involves complex regex parsing of log messages, sentiment analysis using a custom Python library, and then aggregating results by customer ID. The team requires a highly scalable solution with maximum flexibility for custom code. Which approach within Microsoft Fabric is the most appropriate?

  1. AData Pipelines with a Copy Data activity
  2. BSpark notebook using PySpark
  3. CDataflows Gen2 with custom functions
  4. DKQL Queryset with external data
Show answer & explanation

Correct answer: B. Spark notebook using PySpark

Spark notebooks with PySpark offer the unparalleled scalability of Apache Spark and the flexibility to integrate custom Python libraries for complex tasks like regex parsing and sentiment analysis on large datasets.

Why the other options are wrong

  • A. Data Pipelines are for orchestration and data movement, not for performing complex, code-driven transformations like sentiment analysis.
  • C. While Dataflows Gen2 supports custom functions, it's generally not designed for highly complex, large-scale code-driven transformations involving external libraries and regex on terabytes of data.
  • D. KQL is a query language; it's not suitable for executing custom Python libraries or complex regex parsing as described for large-scale transformations.

PySpark for Advanced ETL

Leveraging PySpark in Spark notebooks for highly scalable and customizable Extract, Transform, Load (ETL) operations, including complex parsing and external library integration.

  • Scales to petabytes of data.
  • Supports Python, Scala, R, SQL.
  • Ideal for machine learning and complex analytical transformations.

Memory trick: PySpark's code unleashes powerful insights from massive data.

More Prepare and transform data (20-25%) questions