AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisEasy

A data scientist is analyzing a large dataset of customer reviews. They want to identify common themes and topics discussed in the reviews without predefining categories. Which unsupervised learning technique is most appropriate for this task?

  1. AK-Means Clustering
  2. BSupport Vector Machine (SVM)
  3. CPrincipal Component Analysis (PCA)
  4. DLatent Dirichlet Allocation (LDA)
Show answer & explanation

Correct answer: D. Latent Dirichlet Allocation (LDA)

Latent Dirichlet Allocation (LDA) is specifically designed for topic modeling in text data, identifying underlying themes in documents without prior labels.

Why the other options are wrong

  • A. K-Means Clustering is used for numerical data clustering, not directly for topic modeling in text.
  • B. SVM is a supervised learning algorithm used for classification and regression, requiring labeled data.
  • C. PCA is a dimensionality reduction technique for numerical data, not for topic extraction from text.

Latent Dirichlet Allocation (LDA)

A generative probabilistic model for collections of discrete data such as text corpora. It models documents as mixtures of topics, and topics as mixtures of words.

  • Unsupervised learning technique.
  • Used for topic modeling in text data.
  • Identifies latent (hidden) themes in documents.

Memory trick: LDA: Let's Discover Archives of topics!

More Exploratory Data Analysis questions