CompTIA Data+ (DA0-002)Data AnalysisHard

A data scientist is analyzing a large dataset of customer transactions. They want to group customers based on their purchasing behavior without any prior knowledge of customer segments. The goal is to discover natural groupings within the data. Which unsupervised learning technique is most appropriate for this task?

  1. ASupport Vector Machine (SVM)
  2. BDecision Tree Classification
  3. CK-Means Clustering
  4. DLinear Regression
Show answer & explanation

Correct answer: C. K-Means Clustering

K-Means Clustering is an unsupervised learning algorithm that partitions n observations into k clusters, where each observation belongs to the cluster with the nearest mean (centroid). It is ideal for discovering natural groupings in data without prior labels or knowledge of segments.

Why the other options are wrong

  • A. Support Vector Machine (SVM) is a supervised technique for classification or regression.
  • B. Decision Tree Classification is a supervised technique for predicting a categorical outcome.
  • D. Linear Regression is a supervised technique for predicting a continuous outcome.

K-Means Clustering

An unsupervised machine learning algorithm used to partition data into K distinct, non-overlapping subgroups (clusters) where each data point belongs to the cluster with the nearest mean (centroid).

  • Unsupervised learning (no labeled data needed).
  • Requires specifying the number of clusters (K) beforehand.
  • Aims to minimize within-cluster variance.
  • Sensitive to initial centroid placement and outliers.

Memory trick: Clustering 'C'reates groups without a guide.

More Data Analysis questions