Microsoft Certified: Fabric Analytics Engineer AssociateExplore and analyze data (15-20%)Medium
A data analyst is preparing a report on customer demographics. They have a Spark DataFrame named `customer_df` with columns `CustomerID`, `Age`, and `City`. They need to group customers by `City` and then calculate the average `Age` for each city. Which PySpark operation should be used after `groupBy('City')` to compute the average?
- A.filter(col('Age') > 0)
- B.select('City', 'Age')
- C.agg(avg('Age'))
- D.sort('City')
Show answer & explanationAnswer & explanation
Correct answer: C. .agg(avg('Age'))
After grouping a DataFrame using `groupBy()`, the `.agg()` method is used to apply one or more aggregate functions (like `avg`, `sum`, `count`) to the grouped data, producing summary statistics.
Why the other options are wrong
- A. .filter() is for row-level filtering, not aggregation on groups.
- B. .select() is for column projection, not aggregation.
- D. .sort() is for ordering the results, not for calculating aggregates.
PySpark agg() function
The PySpark `.agg()` function is used with grouped DataFrames to apply one or more aggregate functions to the grouped data, producing summary statistics.
- Follows a `groupBy()` operation.
- Takes aggregate functions (e.g., `avg`, `sum`, `count`).
- Returns a DataFrame with aggregated results.
Memory trick: Group your data, then Aggregate the stats.