Microsoft Certified: Fabric Analytics Engineer AssociateExplore and analyze data (15-20%)Medium

A data engineer is optimizing a PySpark script that processes a large Delta table named `event_logs`. They frequently need to extract the year and month from a `timestamp` column for filtering and grouping. Which PySpark function is the most efficient and idiomatic way to achieve this for both year and month extraction?

  1. Adf.withColumn('event_year', F.expr("YEAR(timestamp)")).withColumn('event_month', F.expr("MONTH(timestamp)"))
  2. Bdf.rdd.map(lambda row: (row['timestamp'].year, row['timestamp'].month)).toDF(['event_year', 'event_month'])
  3. Cdf.withColumn('event_year', year('timestamp')).withColumn('event_month', month('timestamp'))
  4. Ddf.selectExpr("*, YEAR(timestamp) as event_year, MONTH(timestamp) as event_month")
Show answer & explanation

Correct answer: C. df.withColumn('event_year', year('timestamp')).withColumn('event_month', month('timestamp'))

Using `df.withColumn()` with directly imported PySpark SQL functions like `year()` and `month()` from `pyspark.sql.functions` is the most idiomatic and performant way to add new columns based on transformations of existing columns in PySpark. These functions are optimized for Spark's distributed execution.

Why the other options are wrong

  • A. Using `F.expr()` works but is generally less type-safe and potentially less performant than dedicated functions when direct PySpark functions exist.
  • B. Converting to RDD and mapping is inefficient and should be avoided for column-wise operations where DataFrame API functions exist, as it bypasses Spark's Catalyst optimizer.
  • D. While `selectExpr()` can achieve this, `withColumn()` is often preferred for adding individual columns to retain existing ones without explicit 'select *'.

PySpark Date Part Extraction

The process of deriving specific components (like year, month, day) from a date or timestamp column in a PySpark DataFrame using optimized built-in functions.

  • Use `pyspark.sql.functions` for date/time functions.
  • `year()` and `month()` are common functions.
  • `withColumn()` is used to add or update columns.
  • Avoid RDD operations for DataFrame transformations when possible.

Memory trick: Functions Extract Dates, Columns Get New Names, All in PySpark's Flow.

More Explore and analyze data (15-20%) questions