AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium
A research team is developing a generative AI model to create novel protein structures based on a sequence of amino acids. They need an architecture that can understand long-range dependencies within the sequence and process the input in parallel for efficiency. Which neural network architecture is best suited for this task?
- AMultilayer Perceptron (MLP)
- BTransformer
- CRecurrent Neural Network (RNN)
- DConvolutional Neural Network (CNN)
Show answer & explanationAnswer & explanation
Correct answer: B. Transformer
The Transformer architecture is specifically designed to handle sequential data with long-range dependencies and allows for parallel processing of input, which makes it highly efficient and effective for tasks like generating protein structures from amino acid sequences. RNNs struggle with long-range dependencies and parallelization, while CNNs are primarily for spatial data and MLPs lack sequence understanding.
Why the other options are wrong
- A. MLPs are simple feedforward networks that do not inherently handle sequential data or long-range dependencies well.
- C. RNNs process sequences sequentially, making them slow and prone to vanishing/exploding gradients over long sequences.
- D. CNNs are generally used for grid-like data (e.g., images) and are not ideal for understanding long-range dependencies in sequences.
Transformer Architecture
A neural network architecture that relies on self-attention mechanisms to weigh the importance of different parts of the input sequence, enabling parallel processing and effective capture of long-range dependencies.
- Uses self-attention (multi-head attention).
- Processes input elements in parallel.
- Excellent for sequence-to-sequence tasks (e.g., translation, text generation).
Memory trick: Transformers transform sequences with attention.