Microsoft Certified: Fabric Analytics Engineer AssociatePrepare and transform data (20-25%)Medium
A data engineering team is setting up an ingestion process for customer survey responses. The responses are collected via a third-party service that provides a daily export of data in Parquet format, delivered to a specific Azure Data Lake Storage Gen2 path. The team needs to ensure that historical data is preserved and that new daily files are appended to the existing Lakehouse table without re-processing the entire dataset each day. Which Spark DataFrame write mode should be used to achieve this incremental loading strategy?
- Aerrorifexists
- Bappend
- Coverwrite
- Dignore
Show answer & explanationAnswer & explanation
Correct answer: B. append
The 'append' write mode in Spark DataFrames is specifically designed to add new rows to an existing table, which is essential for incremental loading strategies where new daily data is added without overwriting historical records.
Why the other options are wrong
- A. The 'errorifexists' mode would throw an error if the table already exists, preventing any data from being written if it's not a fresh creation.
- C. The 'overwrite' mode would delete all existing data in the table and replace it with the new data, which is not suitable for preserving historical data.
- D. The 'ignore' mode would do nothing if the table already exists, meaning the new data would not be added at all.
Spark DataFrame Write Modes
Settings that control how a Spark DataFrame writes data to a destination when the table/path already exists.
- append: Adds new data to existing.
- overwrite: Replaces existing data.
- ignore: Does nothing if exists.
- errorifexists: Fails if exists.
Memory trick: Append mode adds to the story, never erasing history.