AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium
A research team is exploring how to enable an AI model to understand the hierarchical relationships and semantic similarities between words, such as knowing that 'apple' and 'banana' are both 'fruits', and 'fruit' is a type of 'food'. Which AI/ML concept is best suited for representing these complex linguistic relationships?
- AWord Embeddings
- BBag-of-Words (BoW)
- CTerm Frequency-Inverse Document Frequency (TF-IDF)
- DN-grams
Show answer & explanationAnswer & explanation
Correct answer: A. Word Embeddings
Word embeddings represent words as dense vectors in a continuous vector space where words with similar meanings or relationships are closer together. This allows models to capture semantic and syntactic relationships, directly supporting the understanding of 'apple' and 'banana' being 'fruits'.
Why the other options are wrong
- B. Bag-of-Words represents text as a collection of word counts, ignoring word order and semantic relationships.
- C. TF-IDF measures the importance of a word in a document relative to a corpus but doesn't inherently capture semantic relationships like 'apple' and 'banana' being 'fruits'.
- D. N-grams are sequences of N words, capturing some local context, but they don't inherently represent deeper semantic or hierarchical relationships between words in a continuous space.
Word Embeddings
Word embeddings are dense vector representations of words in a continuous vector space, where words with similar meanings or relationships are located closer to each other, enabling AI models to understand semantic and syntactic contexts.
- Words as dense vectors.
- Captures semantic/syntactic relationships.
- Similar words are closer in vector space.
Memory trick: Turn words into numbers, keep the meaning.