Microsoft Certified: Fabric Analytics Engineer AssociateExplore and analyze data (15-20%)Medium

A data engineer is working with a Spark Delta table named `product_reviews` in Microsoft Fabric. The table contains a `review_text` column and a `sentiment_score` column. The engineer needs to add a new column named `review_category` based on the `sentiment_score`. If the `sentiment_score` is greater than 0.75, the `review_category` should be 'Positive'. If it's less than 0.25, it should be 'Negative'. Otherwise, it should be 'Neutral'. Which PySpark function or method should the data engineer use to achieve this?

  1. Adf.selectExpr('*', 'IF(sentiment_score > 0.75, "Positive", IF(sentiment_score < 0.25, "Negative", "Neutral")) AS review_category')
  2. Bdf.apply(lambda row: 'Positive' if row['sentiment_score'] > 0.75 else ('Negative' if row['sentiment_score'] < 0.25 else 'Neutral'), axis=1)
  3. Cdf.withColumn('review_category', case(col('sentiment_score') > 0.75, 'Positive').case(col('sentiment_score') < 0.25, 'Negative').default('Neutral'))
  4. Ddf.withColumn('review_category', when(col('sentiment_score') > 0.75, 'Positive').when(col('sentiment_score') < 0.25, 'Negative').otherwise('Neutral'))
Show answer & explanation

Correct answer: D. df.withColumn('review_category', when(col('sentiment_score') > 0.75, 'Positive').when(col('sentiment_score') < 0.25, 'Negative').otherwise('Neutral'))

The PySpark `when().otherwise()` construct is the standard and most efficient way to implement conditional logic for creating new columns in a DataFrame based on multiple conditions. It allows for chained `when()` clauses and a final `otherwise()` for the default case.

Why the other options are wrong

  • A. This uses Spark SQL syntax (`selectExpr` with `IF` statements) which is not PySpark, though it could achieve the result if the question asked for Spark SQL.
  • B. Using `apply` with a lambda function on a DataFrame is generally inefficient for large datasets in PySpark as it processes row-by-row, defeating Spark's distributed nature.
  • C. This syntax is incorrect for PySpark. `case` and `default` are not the functions used for this purpose.

PySpark when().otherwise()

A PySpark function used with `withColumn()` to apply conditional logic (if-else if-else) when creating or updating a DataFrame column.

  • Chains `when()` conditions followed by an `otherwise()` for a default value.
  • Requires importing `when` and `col` from `pyspark.sql.functions`.
  • Efficient for conditional logic on large datasets compared to UDFs or `apply`.

Memory trick: When the light changes, otherwise you stop.

More Explore and analyze data (15-20%) questions