CompTIA Data+ (DA0-002)Data AnalysisHard

A data analyst is working with a large dataset of customer demographics and purchase history. They want to identify distinct groups of customers based on their similarities in age, income, and purchasing behavior without any prior knowledge of what these groups might be. Which unsupervised machine learning technique is best suited for this task?

  1. ADecision Tree Classification
  2. BLinear Regression
  3. CLogistic Regression
  4. DK-Means Clustering
Show answer & explanation

Correct answer: D. K-Means Clustering

K-Means Clustering is an unsupervised learning algorithm used to partition 'n' observations into 'k' clusters, where each observation belongs to the cluster with the nearest mean. It is ideal for identifying natural groupings within data when there is no predefined target variable, which matches the scenario of finding distinct customer groups based on similarities.

Why the other options are wrong

  • A. Decision Tree Classification is a supervised learning technique for predicting categorical target variables.
  • B. Linear Regression is a supervised learning technique for predicting continuous target variables.
  • C. Logistic Regression is a supervised learning technique for predicting binary or categorical target variables.

K-Means Clustering

An unsupervised machine learning algorithm that partitions a dataset into 'k' distinct, non-overlapping subgroups (clusters), where each data point belongs to the cluster with the closest mean (centroid).

  • Unsupervised: no labeled target variable required.
  • Aims to minimize intra-cluster variance and maximize inter-cluster variance.
  • Requires specifying the number of clusters (k) beforehand.

Memory trick: K-Means is like telling a group of kids to sort themselves into K teams based on how similar they feel.

More Data Analysis questions