Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Easy
A financial services company needs to ingest daily transaction data from an on-premises Oracle database into a Fabric Lakehouse. The data volume is moderate (tens of gigabytes per day), and the process requires reliable scheduling and monitoring. The company has already set up an On-premises data gateway. Which Fabric component should be used to orchestrate this data movement?
- AData Pipelines with a Copy Data activity.
- BReal-time Analytics Eventstreams.
- CSpark Notebooks using JDBC to connect to Oracle.
- DDataflows Gen2 with an Oracle connector.
Show answer & explanationAnswer & explanation
Correct answer: A. Data Pipelines with a Copy Data activity.
Data Pipelines are ideal for orchestrating scheduled data movement activities, especially from on-premises sources using a gateway. The Copy Data activity is specifically designed for efficient data transfer between various data stores, including Oracle to Lakehouse.
Why the other options are wrong
- B. Eventstreams are for real-time ingestion of continuous data streams, not for daily batch ingestion from a transactional database.
- C. Spark Notebooks are for complex data processing and transformations, not primarily for orchestrating scheduled data ingestion from on-premises databases.
- D. While Dataflows Gen2 can connect to Oracle, Data Pipelines offer better orchestration, scheduling, and monitoring capabilities for recurring data movement tasks.
Data Pipelines for On-Premises Ingestion
Data Pipelines in Microsoft Fabric provide orchestration capabilities, including scheduling and monitoring, for ingesting data from various sources, particularly on-premises databases via an On-premises data gateway, using activities like 'Copy Data'.
- Orchestrates data movement and transformation.
- Supports scheduling and monitoring of activities.
- Uses Copy Data activity for efficient data transfer.
- Connects to on-premises sources via On-premises data gateway.
Memory trick: On-premises data moves through a pipeline to the Lakehouse.