Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Easy

A data engineering team needs to ingest historical sales data from an on-premises Oracle database into a Microsoft Fabric Lakehouse. The data volume is large, and the ingestion process must be scheduled daily. Which Microsoft Fabric component is the most appropriate for this task?

  1. AEventstream
  2. BDataflows Gen2
  3. CData Pipelines (Copy Data activity)
  4. DSpark Notebooks
Show answer & explanation

Correct answer: C. Data Pipelines (Copy Data activity)

For scheduled, large-volume data ingestion from on-premises sources like Oracle databases into a Lakehouse, Data Pipelines with a Copy Data activity are the most suitable and efficient choice in Microsoft Fabric. They provide robust scheduling, monitoring, and connectivity options.

Why the other options are wrong

  • A. Eventstream is designed for real-time, high-throughput event ingestion, not for scheduled batch ingestion of historical data from an on-premises database.
  • B. Dataflows Gen2 is better suited for smaller to medium-sized data transformations and can be used for ingestion but might be less efficient for very large, scheduled on-premises data transfers.
  • D. Spark Notebooks are primarily for data transformation and analytics, not for direct, scheduled data ingestion from external sources like an on-premises Oracle database.

Data Pipelines Copy Data

Data Pipelines' Copy Data activity in Microsoft Fabric enables scheduled, scalable data movement between various sources and destinations, including on-premises databases.

  • Supports a wide range of connectors including on-premises via gateway.
  • Ideal for scheduled batch data ingestion.
  • Provides monitoring and retry capabilities.

Memory trick: Connect, Schedule, Move: The Fabric pipeline for data flow.

More Prepare and transform data (20-25%) questions