Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Easy
A data engineer needs to ingest log files generated by a web application into a Lakehouse. The log files are stored in an Azure Blob Storage container. Due to potential network issues or temporary service outages, the ingestion process might occasionally fail. The engineer wants the Data Pipeline to automatically attempt to re-run failed ingestion steps a few times before giving up. Which setting in the Data Pipeline's Copy Data activity should be configured?
- ARetry policy
- BData consistency verification
- CDegree of copy parallelism
- DFault tolerance settings
Show answer & explanationAnswer & explanation
Correct answer: A. Retry policy
The 'Retry policy' setting in a Data Pipeline's activity directly controls how many times an activity will be retried and the delay between retries if it fails. This is precisely what's needed for handling transient issues during data ingestion.
Why the other options are wrong
- B. Data consistency verification ensures data integrity after transfer but doesn't manage the retrying of failed transfers.
- C. Degree of copy parallelism affects how many parallel connections are used during data transfer, influencing performance, not retry behavior.
- D. Fault tolerance settings deal with handling corrupt or incompatible data during transfer (e.g., skipping incompatible rows), not retrying the entire activity on failure.
Data Pipeline Retry Policy
The Retry Policy in Data Pipeline activities defines the number of retries and the delay between them for failed activity executions.
- Handles transient failures (e.g., network issues, temporary service unavailability).
- Configurable per activity (e.g., Copy Data).
- Improves pipeline robustness and reliability.
Memory trick: A robust pipeline, like a persistent athlete, tries again after a stumble.