Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Hard

A manufacturing company needs to analyze sensor data from IoT devices. This data arrives in nearly real-time in CSV format and is stored temporarily in a blob storage account. Before loading into a Lakehouse, the data needs to be aggregated hourly, and specific error codes (e.g., 'ERR-001', 'ERR-002') should be filtered out. The data engineers want to minimize manual coding as much as possible for this continuous process. Which Microsoft Fabric tool combination is the most efficient?

  1. AData Pipeline with a custom Spark notebook activity
  2. BData Pipeline with a Dataflows Gen2 activity
  3. CDataflows Gen2 with a Blob Storage connector
  4. DKQL Queryset directly on the blob storage
Show answer & explanation

Correct answer: B. Data Pipeline with a Dataflows Gen2 activity

A Data Pipeline can orchestrate the continuous ingestion, triggering a Dataflows Gen2 activity to perform the low-code transformations and hourly aggregations before loading into the Lakehouse. This combines orchestration with visual transformation.

Why the other options are wrong

  • A. While possible, using a custom Spark notebook activity for simple filtering and aggregation might be overkill if a low-code option exists, increasing development and maintenance effort.
  • C. Dataflows Gen2 can perform the transformations, but it typically runs on a schedule. A Data Pipeline is better for orchestrating a continuous process that might involve triggers from new file arrivals.
  • D. KQL Queryset is for querying data already in a Kusto database, not for ingesting CSV from blob storage and performing ETL operations like aggregation.

Data Pipeline + Dataflows Gen2

Combining Data Pipelines for orchestration and scheduling with Dataflows Gen2 for low-code data ingestion and transformation.

  • Pipelines manage flow, Dataflows handle ETL.
  • Enables continuous, event-driven processing.
  • Leverages low-code for transformations while ensuring robust orchestration.

Memory trick: The pipeline guides the dataflow, transforming it without a single line of code.

More Prepare and transform data (20-25%) questions