A data engineer needs to ingest log files from an Azure Data Lake Storage Gen2 account into a Fabric Lakehouse. The log files are organized in a hierarchical folder structure by year, month, and day (e.g., `logs/2023/01/01/log.json`). The requirement is to load all log files from a specific year and month into a single table, ensuring that new files added to that month's folder are automatically picked up in subsequent runs. Which Data Pipeline activity and configuration should be used?
- AA Copy Data activity with a parameterized source path for year/month, and 'Recursively' enabled.
- BA Copy Data activity with a wildcard path for the folder, and 'Last modified date' as the 'File filter'.
- CA Dataflows Gen2 with a Data Lake connector and manual folder selection.
- DA Web activity to list files, followed by a ForEach activity to copy each file individually.
Show answer & explanationAnswer & explanation
Correct answer: A. A Copy Data activity with a parameterized source path for year/month, and 'Recursively' enabled.
A Copy Data activity in a Data Pipeline, configured with a parameterized source path (e.g., `logs/@{formatDateTime(pipeline().parameters.YearMonth, 'yyyy/MM')}/`) and the 'Recursively' option enabled, allows dynamic selection of a specific month's folder while automatically picking up all files within it. For new files, the pipeline would rerun for the relevant month. Fabric's Copy Data activity is optimized for this kind of bulk ingestion from hierarchical storage.
Why the other options are wrong
- B. While wildcards work, 'Last modified date' as a file filter is for incremental loads based on file modification, not for efficiently picking up all files in a specific dynamic folder structure. The 'Recursively' option is key here.
- C. Dataflows Gen2 can ingest from Data Lake, but Data Pipelines provide better orchestration, scheduling, and parameterized path capabilities for this specific bulk ingestion scenario, especially when dealing with dynamic folder structures and ensuring new files are picked up.
- D. This approach is overly complex and less efficient for bulk copying compared to a single Copy Data activity designed for this purpose. It introduces unnecessary overhead for listing and iterating.
Data Pipeline Parameterized ADLS Gen2 Ingestion
Data Pipelines can ingest from Azure Data Lake Storage Gen2 using a Copy Data activity with parameterized source paths. By enabling 'Recursively' and using pipeline parameters for year/month, it efficiently loads all files from dynamic hierarchical folders.
- Uses Copy Data activity for bulk ingestion.
- Parameterized paths (`@{pipeline().parameters.YearMonth}`) for dynamic folder selection.
- Recursively option to include all subfolders and files.
- Efficient for hierarchical data in ADLS Gen2.
- Ideal for scheduled ingestion of new files in a given path.
Memory trick: Pipeline parameters guide data through folders to Lakehouse.