A global logistics company uses Data Pipelines in Microsoft Fabric to ingest tracking data from various regional operational systems into a central Lakehouse. Due to varying network conditions and source system availability, some ingestions occasionally fail or are delayed. The data engineering team needs a mechanism to automatically retry failed data copy activities with an increasing delay between attempts, and to limit the total number of retries to prevent indefinite loops. Which Data Pipeline configuration should they use?
- AConfigure the 'Retry policy' settings: 'Retry' count and 'Retry interval (seconds)'.
- BSet the 'Concurrency' property to a high value and 'Timeout' to 00:05:00.
- CUse a 'Set variable' activity to track attempts and an 'If condition' for failure.
- DImplement a 'Wait' activity followed by an 'Until' loop for custom retry logic.
Show answer & explanationAnswer & explanation
Correct answer: A. Configure the 'Retry policy' settings: 'Retry' count and 'Retry interval (seconds)'.
Data Pipelines in Microsoft Fabric (based on Azure Data Factory) provide built-in retry mechanisms for activities. The 'Retry policy' settings, specifically 'Retry' count and 'Retry interval (seconds)', allow engineers to define how many times an activity should be re-attempted and the delay between these attempts, directly addressing the requirement for automatic retries with increasing delay and a retry limit.
Why the other options are wrong
- B. Concurrency controls parallel execution, and Timeout limits execution duration; neither directly handles automatic retries with increasing delay.
- C. This approach involves complex manual implementation for a feature already provided natively by the pipeline activity's retry policy.
- D. While custom logic could be built, the built-in retry policy is more efficient and direct for this common scenario, and 'Until' loops are for iterating until a condition is met, not primarily for retries.
Data Pipeline Retry Policy
A built-in mechanism in Microsoft Fabric Data Pipelines (and Azure Data Factory) to automatically re-execute failed activities a specified number of times with a defined delay between attempts.
- Configurable per activity or pipeline.
- Includes 'Retry count' and 'Retry interval'.
- Essential for fault tolerance and robust data ingestion.
Memory trick: Don't just fail; retry with a policy, like a persistent mail carrier.