AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisHard
A data scientist is analyzing a dataset of customer survey responses. One question asks for 'CustomerSatisfaction' on a scale of 1 to 5. They notice that the responses are heavily concentrated at 4 and 5, with very few responses at 1, 2, or 3. They want to check if this observed distribution significantly differs from a hypothetical uniform distribution (where each rating 1-5 has an equal chance). Which statistical test is most appropriate for this comparison?
- AChi-squared goodness-of-fit test
- BPaired t-test
- CANOVA
- DKolmogorov-Smirnov test
Show answer & explanationAnswer & explanation
Correct answer: A. Chi-squared goodness-of-fit test
The Chi-squared goodness-of-fit test is used to determine whether an observed frequency distribution differs significantly from an expected (hypothetical) distribution. In this case, comparing the observed 'CustomerSatisfaction' ratings to a uniform distribution (where each category has an equal expected frequency) is precisely what this test is designed for.
Why the other options are wrong
- B. A paired t-test compares means of two related groups.
- C. ANOVA compares means of a continuous variable across three or more independent groups.
- D. The Kolmogorov-Smirnov test compares a sample distribution to a reference distribution or two sample distributions, but it's typically for continuous data and might be less direct for categorical frequency comparison than chi-squared.
Chi-squared Goodness-of-Fit Test
A statistical test used to determine if an observed frequency distribution for a single categorical variable differs significantly from an expected theoretical distribution.
- Compares observed frequencies to expected frequencies.
- Null hypothesis: observed distribution fits the expected distribution.
- Alternative hypothesis: observed distribution does not fit the expected distribution.
- Requires categorical data and expected frequencies for each category.
Memory trick: Chi-squared goodness-of-fit checks if the observed fits the 'should'.