AWS Certified Machine Learning – SpecialtyExploratory Data AnalysisHard
A data scientist is analyzing a large dataset of customer feedback text. They want to identify groups of reviews that express similar themes or topics without any pre-labeled data. The goal is to uncover latent topics within the text to better understand customer concerns. Which unsupervised technique is most suitable for this task?
- ALatent Dirichlet Allocation (LDA)
- BWord Embeddings (e.g., Word2Vec)
- CNamed Entity Recognition (NER)
- DSentiment Analysis
Show answer & explanationAnswer & explanation
Correct answer: A. Latent Dirichlet Allocation (LDA)
Latent Dirichlet Allocation (LDA) is an unsupervised topic modeling technique specifically designed to discover abstract "topics" that occur in a collection of documents. It assumes that documents are a mixture of various topics and that each topic is a mixture of words. This directly addresses the goal of uncovering latent themes or topics in customer feedback without pre-labeled data.
Why the other options are wrong
- B. Word Embeddings represent words as vectors and capture semantic relationships, but they don't directly perform topic extraction across documents.
- C. Named Entity Recognition (NER) identifies and classifies named entities (e.g., people, organizations) in text, not for topic discovery.
- D. Sentiment Analysis is a supervised or lexicon-based technique for classifying emotional tone, not for discovering latent topics.
Latent Dirichlet Allocation (LDA)
An unsupervised generative probabilistic model used for topic modeling, which presumes that documents are combinations of topics and that topics are combinations of words.
- Discovers abstract 'topics' from a collection of documents.
- Documents are modeled as mixtures of topics.
- Topics are modeled as mixtures of words.
- Requires specifying the number of topics (K) beforehand.
Memory trick: LDA lets documents discover their latent topics.