Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Hard
A manufacturing company needs to analyze sensor data from IoT devices. This data arrives in nearly real-time in CSV format and is stored temporarily in a blob storage account. Before loading into a Lakehouse, the data needs to be aggregated hourly, and specific error codes (e.g., 'ERR-001', 'ERR-002') should be filtered out. The data engineers want to minimize manual coding as much as possible for this continuous process. Which Microsoft Fabric tool combination is the most efficient?
- AData Pipeline with a custom Spark notebook activity
- BData Pipeline with a Dataflows Gen2 activity
- CDataflows Gen2 with a Blob Storage connector
- DKQL Queryset directly on the blob storage
Show answer & explanationAnswer & explanation
Correct answer: B. Data Pipeline with a Dataflows Gen2 activity
A Data Pipeline can orchestrate the continuous ingestion, triggering a Dataflows Gen2 activity to perform the low-code transformations and hourly aggregations before loading into the Lakehouse. This combines orchestration with visual transformation.
Why the other options are wrong
- A. While possible, using a custom Spark notebook activity for simple filtering and aggregation might be overkill if a low-code option exists, increasing development and maintenance effort.
- C. Dataflows Gen2 can perform the transformations, but it typically runs on a schedule. A Data Pipeline is better for orchestrating a continuous process that might involve triggers from new file arrivals.
- D. KQL Queryset is for querying data already in a Kusto database, not for ingesting CSV from blob storage and performing ETL operations like aggregation.
Data Pipeline + Dataflows Gen2
Combining Data Pipelines for orchestration and scheduling with Dataflows Gen2 for low-code data ingestion and transformation.
- Pipelines manage flow, Dataflows handle ETL.
- Enables continuous, event-driven processing.
- Leverages low-code for transformations while ensuring robust orchestration.
Memory trick: The pipeline guides the dataflow, transforming it without a single line of code.