Microsoft Certified: Fabric Analytics Engineer AssociateExplore and analyze data (15-20%)Medium

A data analyst is preparing a report on customer demographics. They have a Spark DataFrame named `customer_df` with columns `CustomerID`, `Age`, and `City`. They need to group customers by `City` and then calculate the average `Age` for each city. Which PySpark operation should be used after `groupBy('City')` to compute the average?

  1. A.filter(col('Age') > 0)
  2. B.select('City', 'Age')
  3. C.agg(avg('Age'))
  4. D.sort('City')
Show answer & explanation

Correct answer: C. .agg(avg('Age'))

After grouping a DataFrame using `groupBy()`, the `.agg()` method is used to apply one or more aggregate functions (like `avg`, `sum`, `count`) to the grouped data, producing summary statistics.

Why the other options are wrong

  • A. .filter() is for row-level filtering, not aggregation on groups.
  • B. .select() is for column projection, not aggregation.
  • D. .sort() is for ordering the results, not for calculating aggregates.

PySpark agg() function

The PySpark `.agg()` function is used with grouped DataFrames to apply one or more aggregate functions to the grouped data, producing summary statistics.

  • Follows a `groupBy()` operation.
  • Takes aggregate functions (e.g., `avg`, `sum`, `count`).
  • Returns a DataFrame with aggregated results.

Memory trick: Group your data, then Aggregate the stats.

More Explore and analyze data (15-20%) questions