AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium
A data scientist is working with a large text corpus to train a generative AI model. They need to prepare the text data to be suitable for machine learning algorithms, which typically require numerical input. Which of the following techniques would be most appropriate for converting words into numerical representations while capturing semantic meaning?
- APadding
- BTokenization
- COne-hot encoding
- DWord embeddings
Show answer & explanationAnswer & explanation
Correct answer: D. Word embeddings
Word embeddings are dense vector representations of words that capture semantic relationships. Words with similar meanings will have similar vector representations, which is crucial for generative AI models to understand context and generate coherent text. While tokenization and one-hot encoding are steps in text processing, they don't capture semantic meaning in the same rich way.
Why the other options are wrong
- A. Padding is used to make sequences of varying lengths uniform, but it does not convert words into numerical representations or capture semantic meaning.
- B. Tokenization is the process of breaking text into smaller units (tokens), which is a prerequisite but not the final step for numerical representation with semantic meaning.
- C. One-hot encoding creates sparse binary vectors and does not capture semantic relationships between words.
Word Embeddings
Word embeddings are dense vector representations of words that encode their semantic meaning and relationships. Words with similar meanings are mapped to nearby points in the vector space.
- Capture semantic and syntactic relationships.
- Lower-dimensional and denser than one-hot encodings.
- Learned from large text corpora (e.g., Word2Vec, GloVe, FastText).
Memory trick: Words to numbers, retaining meaning.