Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Easy
A data engineering team needs to ingest historical sales data from an on-premises Oracle database into a Microsoft Fabric Lakehouse. The data volume is large, and the ingestion process must be scheduled daily. Which Microsoft Fabric component is the most appropriate for this task?
- AEventstream
- BDataflows Gen2
- CData Pipelines (Copy Data activity)
- DSpark Notebooks
Show answer & explanationAnswer & explanation
Correct answer: C. Data Pipelines (Copy Data activity)
For scheduled, large-volume data ingestion from on-premises sources like Oracle databases into a Lakehouse, Data Pipelines with a Copy Data activity are the most suitable and efficient choice in Microsoft Fabric. They provide robust scheduling, monitoring, and connectivity options.
Why the other options are wrong
- A. Eventstream is designed for real-time, high-throughput event ingestion, not for scheduled batch ingestion of historical data from an on-premises database.
- B. Dataflows Gen2 is better suited for smaller to medium-sized data transformations and can be used for ingestion but might be less efficient for very large, scheduled on-premises data transfers.
- D. Spark Notebooks are primarily for data transformation and analytics, not for direct, scheduled data ingestion from external sources like an on-premises Oracle database.
Data Pipelines Copy Data
Data Pipelines' Copy Data activity in Microsoft Fabric enables scheduled, scalable data movement between various sources and destinations, including on-premises databases.
- Supports a wide range of connectors including on-premises via gateway.
- Ideal for scheduled batch data ingestion.
- Provides monitoring and retry capabilities.
Memory trick: Connect, Schedule, Move: The Fabric pipeline for data flow.