Microsoft Certified: Fabric Analytics Engineer AssociateExplore and analyze data (15-20%)Medium
A data engineer is working with a Spark Delta table named `product_reviews` in Microsoft Fabric. The table contains a `review_text` column and a `sentiment_score` column. The engineer needs to add a new column named `review_category` based on the `sentiment_score`. If the `sentiment_score` is greater than 0.75, the `review_category` should be 'Positive'. If it's less than 0.25, it should be 'Negative'. Otherwise, it should be 'Neutral'. Which PySpark function or method should the data engineer use to achieve this?
- Adf.selectExpr('*', 'IF(sentiment_score > 0.75, "Positive", IF(sentiment_score < 0.25, "Negative", "Neutral")) AS review_category')
- Bdf.apply(lambda row: 'Positive' if row['sentiment_score'] > 0.75 else ('Negative' if row['sentiment_score'] < 0.25 else 'Neutral'), axis=1)
- Cdf.withColumn('review_category', case(col('sentiment_score') > 0.75, 'Positive').case(col('sentiment_score') < 0.25, 'Negative').default('Neutral'))
- Ddf.withColumn('review_category', when(col('sentiment_score') > 0.75, 'Positive').when(col('sentiment_score') < 0.25, 'Negative').otherwise('Neutral'))
Show answer & explanationAnswer & explanation
Correct answer: D. df.withColumn('review_category', when(col('sentiment_score') > 0.75, 'Positive').when(col('sentiment_score') < 0.25, 'Negative').otherwise('Neutral'))
The PySpark `when().otherwise()` construct is the standard and most efficient way to implement conditional logic for creating new columns in a DataFrame based on multiple conditions. It allows for chained `when()` clauses and a final `otherwise()` for the default case.
Why the other options are wrong
- A. This uses Spark SQL syntax (`selectExpr` with `IF` statements) which is not PySpark, though it could achieve the result if the question asked for Spark SQL.
- B. Using `apply` with a lambda function on a DataFrame is generally inefficient for large datasets in PySpark as it processes row-by-row, defeating Spark's distributed nature.
- C. This syntax is incorrect for PySpark. `case` and `default` are not the functions used for this purpose.
PySpark when().otherwise()
A PySpark function used with `withColumn()` to apply conditional logic (if-else if-else) when creating or updating a DataFrame column.
- Chains `when()` conditions followed by an `otherwise()` for a default value.
- Requires importing `when` and `col` from `pyspark.sql.functions`.
- Efficient for conditional logic on large datasets compared to UDFs or `apply`.
Memory trick: When the light changes, otherwise you stop.