AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisMedium

A data scientist is analyzing a dataset of customer demographics. They want to visualize the distribution of 'age' and 'income' simultaneously to identify any potential clusters or relationships. The 'age' feature is normally distributed, while 'income' is heavily skewed to the right. Which data visualization technique is most appropriate to display the joint distribution of these two features effectively, considering their different distributional properties?

  1. ABox plot for 'age' and a separate box plot for 'income'.
  2. BHistogram for 'age' and a separate histogram for 'income'.
  3. CScatter plot with a 2D kernel density estimate (KDE) overlay.
  4. DBar chart showing the counts of 'age' and 'income' categories.
Show answer & explanation

Correct answer: C. Scatter plot with a 2D kernel density estimate (KDE) overlay.

A scatter plot is excellent for showing the relationship between two continuous variables. Adding a 2D KDE overlay allows visualization of the joint density, revealing areas of higher concentration and potential clusters, which is particularly useful when one or both variables have complex distributions like skewness.

Why the other options are wrong

  • A. Incorrect. Separate box plots show individual distributions and potential outliers, but not the joint distribution or relationship between the variables.
  • B. Incorrect. Separate histograms show individual distributions, but not their joint distribution or relationship.
  • D. Incorrect. A bar chart is for categorical data or binned continuous data, and would not effectively show the continuous joint distribution or relationship.

2D Kernel Density Estimate (KDE)

A non-parametric way to estimate the probability density function of two continuous random variables, often visualized as contours or a heatmap on a scatter plot, showing areas of higher data concentration.

  • Visualizes joint distribution of two continuous variables.
  • Does not assume a specific distribution shape.
  • Useful for identifying clusters and density variations.

Memory trick: Scatter points tell the story, KDE shows the density's glory.

More Exploratory Data Analysis questions