AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium

A data scientist is preparing a dataset of customer reviews for sentiment analysis using a machine learning model. The reviews are free-form text. To enable the model to understand the semantic relationships between words (e.g., 'good' is similar to 'excellent', 'bad' is similar to 'terrible'), which technique should be applied?

  1. ATF-IDF
  2. BWord Embeddings
  3. COne-Hot Encoding
  4. DStemming
Show answer & explanation

Correct answer: B. Word Embeddings

Word embeddings (like Word2Vec, GloVe, FastText) represent words as dense vectors in a continuous vector space where words with similar meanings are located close to each other. This allows models to capture semantic relationships and context, which is crucial for understanding sentiment.

Why the other options are wrong

  • A. TF-IDF (Term Frequency-Inverse Document Frequency) assigns numerical weights to words based on their frequency in a document and across documents, indicating importance, but it does not capture semantic similarity between different words.
  • C. One-Hot Encoding represents each word as a unique binary vector, but it doesn't capture any semantic relationships between words; all words are equally 'distant'.
  • D. Stemming reduces words to their root form (e.g., 'running' to 'run') to reduce dimensionality and normalize words, but it does not capture semantic relationships between different words.

Word Embeddings

Numerical representations of words in a vector space, where words with similar meanings have similar vector representations, thereby capturing semantic relationships.

  • Represent words as dense, real-valued vectors.
  • Capture semantic and syntactic relationships.
  • Learned from large text corpora.
  • Examples: Word2Vec, GloVe, FastText.

Memory trick: Text to numbers, meaning must transfer.

More AI/ML and Generative AI Fundamentals questions