CompTIA Data+ (DA0-002)Data AnalysisEasy

A data analyst is working with a large dataset of customer demographics and purchase history. They want to identify distinct groups of customers based on their purchasing behavior (e.g., high-value frequent buyers, occasional bargain hunters, loyal brand advocates). The analyst does not have pre-defined labels for these customer segments but wants to discover natural groupings within the data. Which unsupervised machine learning technique is most suitable for this task?

  1. AK-Means Clustering
  2. BLogistic Regression
  3. CLinear Regression
  4. DDecision Tree Analysis
Show answer & explanation

Correct answer: A. K-Means Clustering

K-Means Clustering is an unsupervised learning algorithm used to partition 'n' observations into 'k' clusters, where each observation belongs to the cluster with the nearest mean (centroid). It is ideal for discovering natural groupings in data without prior labels, which perfectly matches the scenario.

Why the other options are wrong

  • B. Logistic Regression is a supervised learning technique used for binary classification (predicting a categorical outcome).
  • C. Linear Regression is a supervised learning technique used for predicting a continuous target variable.
  • D. Decision Tree Analysis is a supervised learning technique used for both classification and regression, requiring labeled data.

K-Means Clustering

An unsupervised machine learning algorithm that partitions a dataset into 'k' distinct, non-overlapping subgroups (clusters). It assigns each data point to the cluster whose mean (centroid) is closest.

  • Unsupervised learning algorithm.
  • Requires the number of clusters (k) to be specified.
  • Aims to minimize the sum of squared distances between data points and their cluster centroids.

Memory trick: No labels? K-Means finds groups, PCA reduces loops!

More Data Analysis questions