AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisMedium

A data scientist is analyzing a dataset containing customer ratings for various products on an e-commerce platform. The distribution of ratings is heavily skewed towards positive values, and there are a few extremely low ratings that could be considered outliers. They need to calculate a measure of central tendency that is least affected by this skewness and these outliers. Which statistical measure should they choose?

  1. AMean
  2. BMedian
  3. CMode
  4. DStandard Deviation
Show answer & explanation

Correct answer: B. Median

The median is a robust statistic that represents the middle value of a dataset. Unlike the mean, it is not influenced by extreme values (outliers) or the skewness of the distribution, making it the most appropriate measure of central tendency in this scenario.

Why the other options are wrong

  • A. The mean is highly sensitive to outliers and skewed distributions, meaning it would be pulled towards the extreme low ratings, misrepresenting the central tendency.
  • C. The mode represents the most frequent value. While it's not affected by outliers, it might not accurately represent the 'central tendency' for continuous or ordinal data unless there's a very clear peak, and it doesn't account for spread.
  • D. The standard deviation measures the spread or dispersion of data, not its central tendency. It is also sensitive to outliers.

Robust Statistics for Skewed Data

Statistical measures that are less affected by outliers or deviations from normality, such as the median for central tendency or the interquartile range for dispersion.

  • Median is robust to outliers and skew for central tendency.
  • IQR is robust to outliers and skew for dispersion.
  • Used when data does not meet assumptions of parametric tests or contains extreme values.

Memory trick: Median: 'Middle' ground, even when data is 'messy'.

More Exploratory Data Analysis questions