A data scientist is analyzing a large dataset of customer reviews. They want to identify groups of customers with similar sentiment patterns without pre-defining the number of groups. The sentiment scores are continuous values between -1 (negative) and 1 (positive). Which unsupervised machine learning technique is best suited for this task?
- AK-Means Clustering
- BHierarchical Clustering
- CLinear Regression
- DPrincipal Component Analysis (PCA)
Show answer & explanationAnswer & explanation
Correct answer: B. Hierarchical Clustering
Hierarchical clustering is an unsupervised technique that does not require pre-specifying the number of clusters. It builds a hierarchy of clusters, represented by a dendrogram, allowing the data scientist to visually determine the optimal number of groups based on the desired level of granularity. K-Means requires a predefined 'k', PCA is for dimensionality reduction, and Linear Regression is supervised.
Why the other options are wrong
- A. K-Means requires the number of clusters (K) to be specified beforehand.
- C. Linear Regression is a supervised learning algorithm for predicting continuous values, not for unsupervised grouping.
- D. PCA is a dimensionality reduction technique, not for grouping data.
Hierarchical Clustering
An unsupervised clustering algorithm that builds a hierarchy of clusters, either by starting with individual data points and merging them (agglomerative) or by starting with one large cluster and splitting it (divisive).
- Does not require pre-specifying the number of clusters.
- Results are often visualized with a dendrogram.
- Useful for exploring natural groupings in data.
- Can be computationally intensive for very large datasets.
Memory trick: Hierarchical clustering helps you harvest clusters from a growing tree.