Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Medium

A data engineer is configuring a Data Pipeline in Microsoft Fabric to ingest CSV files from an Azure Data Lake Storage Gen2 account into a Lakehouse. The source ADLS Gen2 account has multiple containers and folders, and the file paths often include dynamic elements like dates (e.g., `/rawdata/sales/2023/10/26/data.csv`). The engineer needs to ingest files from a specific date range or a particular subfolder based on pipeline execution parameters. Which Data Pipeline feature should be used to achieve this flexible and dynamic file path ingestion?

  1. AParameterizing the source dataset's file path and passing dynamic values from pipeline parameters.
  2. BUsing a wildcard file path (e.g., `/rawdata/sales/*/*/*.csv`) without parameters.
  3. CHardcoding the full file path in the Copy Data activity's source settings.
  4. DCreating a separate Dataflow Gen2 for each specific date range or subfolder.
Show answer & explanation

Correct answer: A. Parameterizing the source dataset's file path and passing dynamic values from pipeline parameters.

Parameterizing the source dataset's file path in a Data Pipeline allows for dynamic ingestion based on pipeline parameters. This enables the engineer to pass values like dates or folder names at runtime, making the pipeline highly flexible and reusable for different ingestion scenarios without modifying the dataset definition or creating multiple pipelines/dataflows.

Why the other options are wrong

  • B. While wildcards can match patterns, they don't allow for precise control over specific date ranges or subfolders based on runtime parameters, potentially ingesting unwanted data.
  • C. Hardcoding paths removes flexibility and requires manual updates for each new date or folder, which is inefficient for dynamic ingestion.
  • D. Creating separate Dataflows Gen2 for each variation is inefficient, leads to pipeline sprawl, and does not leverage the dynamic capabilities of Data Pipelines.

Data Pipeline Parameterized ADLS Gen2 Ingestion

The ability to dynamically specify source file paths in Azure Data Lake Storage Gen2 within a Microsoft Fabric Data Pipeline using parameters, allowing for flexible and reusable ingestion.

  • Uses pipeline parameters to pass dynamic values (e.g., dates, folders).
  • Dataset path is defined with parameters.
  • Enables ingestion from varying paths without pipeline modification.
  • Essential for incremental or date-based loads.

Memory trick: Parameters are like GPS coordinates, guiding your pipeline to the right data each time.

More Prepare and transform data (20-25%) questions