Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Easy

A financial services company needs to ingest daily transaction data from an on-premises Oracle database into a Fabric Lakehouse. The data volume is moderate (tens of gigabytes per day), and the process requires reliable scheduling and monitoring. The company has already set up an On-premises data gateway. Which Fabric component should be used to orchestrate this data movement?

  1. AData Pipelines with a Copy Data activity.
  2. BReal-time Analytics Eventstreams.
  3. CSpark Notebooks using JDBC to connect to Oracle.
  4. DDataflows Gen2 with an Oracle connector.
Show answer & explanation

Correct answer: A. Data Pipelines with a Copy Data activity.

Data Pipelines are ideal for orchestrating scheduled data movement activities, especially from on-premises sources using a gateway. The Copy Data activity is specifically designed for efficient data transfer between various data stores, including Oracle to Lakehouse.

Why the other options are wrong

  • B. Eventstreams are for real-time ingestion of continuous data streams, not for daily batch ingestion from a transactional database.
  • C. Spark Notebooks are for complex data processing and transformations, not primarily for orchestrating scheduled data ingestion from on-premises databases.
  • D. While Dataflows Gen2 can connect to Oracle, Data Pipelines offer better orchestration, scheduling, and monitoring capabilities for recurring data movement tasks.

Data Pipelines for On-Premises Ingestion

Data Pipelines in Microsoft Fabric provide orchestration capabilities, including scheduling and monitoring, for ingesting data from various sources, particularly on-premises databases via an On-premises data gateway, using activities like 'Copy Data'.

  • Orchestrates data movement and transformation.
  • Supports scheduling and monitoring of activities.
  • Uses Copy Data activity for efficient data transfer.
  • Connects to on-premises sources via On-premises data gateway.

Memory trick: On-premises data moves through a pipeline to the Lakehouse.

More Prepare and transform data (20-25%) questions