Microsoft Certified: Fabric Analytics Engineer AssociatePlan and implement data analytics solutions (10-15%)Hard
A data engineer needs to ingest a large volume of CSV files from an external SFTP server into a Microsoft Fabric Lakehouse. The CSV files contain header rows and are comma-delimited. The ingestion process must be robust, handle potential malformed records by quarantining them, and allow for schema inference. Which Data Pipeline activity should be used for this scenario?
- ACopy data activity.
- BStored procedure activity.
- CDataflow Gen2 activity.
- DNotebook activity.
Show answer & explanationAnswer & explanation
Correct answer: A. Copy data activity.
The 'Copy data' activity in a Data Pipeline is highly versatile for ingesting data from various sources to a Lakehouse. It natively supports SFTP, CSV parsing with header rows, schema inference, and robust error handling options like fault tolerance to skip or redirect malformed rows, which effectively quarantines them.
Why the other options are wrong
- B. A stored procedure activity is used to execute SQL stored procedures, not for ingesting files from an SFTP server.
- C. Dataflow Gen2 is more suited for complex data transformations and cleansing, not typically for direct, robust file ingestion with specific error handling like quarantining malformed rows during initial copy.
- D. A Notebook activity would require custom code for SFTP connection, CSV parsing, schema inference, and error handling, which is more complex than using the pre-built 'Copy data' activity for this common pattern.
Data Pipeline Copy Data Activity
The 'Copy data' activity in Microsoft Fabric Data Pipelines facilitates efficient and robust data ingestion from diverse sources to destinations, supporting various file formats, schema inference, and fault tolerance.
- Supports a wide range of connectors (e.g., SFTP, ADLS Gen2).
- Handles common file formats like CSV, Parquet.
- Includes options for schema inference and fault tolerance (skip/redirect malformed rows).
Memory trick: Copy data, handle errors, infer schema, no terrors!