AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium
A data scientist is preparing a dataset of customer reviews for sentiment analysis using a generative AI model. The model needs to understand the semantic relationships between words, such as 'good' being closer to 'excellent' than to 'terrible'. Which data representation technique should be used to capture these semantic relationships?
- ATerm Frequency-Inverse Document Frequency (TF-IDF)
- BBag-of-Words (BoW)
- CWord Embeddings
- DOne-Hot Encoding
Show answer & explanationAnswer & explanation
Correct answer: C. Word Embeddings
Word Embeddings are dense vector representations of words that capture semantic and syntactic relationships. Words with similar meanings or contexts are mapped to nearby points in the vector space, allowing models to understand relationships like 'good' being closer to 'excellent' than to 'terrible'. One-Hot Encoding, TF-IDF, and Bag-of-Words are sparser representations that do not inherently capture such semantic relationships.
Why the other options are wrong
- A. TF-IDF measures word importance based on frequency but does not capture semantic similarity between different words.
- B. Bag-of-Words counts word occurrences, ignoring word order and semantic relationships.
- D. One-Hot Encoding creates sparse, orthogonal vectors that do not capture semantic relationships between words.
Word Embeddings
Dense vector representations of words that capture their semantic and syntactic relationships, allowing words with similar meanings to be represented by vectors that are close in a high-dimensional space.
- Represent words as continuous vectors.
- Capture semantic and syntactic similarities.
- Learned from large text corpora (e.g., Word2Vec, GloVe, FastText).
Memory trick: Embeddings understand words, not just count them.