AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium

A research team is developing a generative AI model to create novel protein structures based on a sequence of amino acids. They need an architecture that can understand long-range dependencies within the sequence and process the input in parallel for efficiency. Which neural network architecture is best suited for this task?

  1. AMultilayer Perceptron (MLP)
  2. BTransformer
  3. CRecurrent Neural Network (RNN)
  4. DConvolutional Neural Network (CNN)
Show answer & explanation

Correct answer: B. Transformer

The Transformer architecture is specifically designed to handle sequential data with long-range dependencies and allows for parallel processing of input, which makes it highly efficient and effective for tasks like generating protein structures from amino acid sequences. RNNs struggle with long-range dependencies and parallelization, while CNNs are primarily for spatial data and MLPs lack sequence understanding.

Why the other options are wrong

  • A. MLPs are simple feedforward networks that do not inherently handle sequential data or long-range dependencies well.
  • C. RNNs process sequences sequentially, making them slow and prone to vanishing/exploding gradients over long sequences.
  • D. CNNs are generally used for grid-like data (e.g., images) and are not ideal for understanding long-range dependencies in sequences.

Transformer Architecture

A neural network architecture that relies on self-attention mechanisms to weigh the importance of different parts of the input sequence, enabling parallel processing and effective capture of long-range dependencies.

  • Uses self-attention (multi-head attention).
  • Processes input elements in parallel.
  • Excellent for sequence-to-sequence tasks (e.g., translation, text generation).

Memory trick: Transformers transform sequences with attention.

More AI/ML and Generative AI Fundamentals questions