AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsMedium

A research team is developing a generative AI model to create novel protein structures based on specific functional requirements. They want the model to understand the long-range dependencies within protein sequences and how changes in one part of the sequence can affect distant parts. Which neural network architecture is best suited for this task?

  1. AMultilayer Perceptron (MLP)
  2. BRecurrent Neural Network (RNN)
  3. CTransformer
  4. DConvolutional Neural Network (CNN)
Show answer & explanation

Correct answer: C. Transformer

Transformers are particularly effective at capturing long-range dependencies in sequential data, such as protein sequences, due to their self-attention mechanism. This allows them to weigh the importance of different parts of the input sequence when processing each element, overcoming limitations of RNNs with very long sequences.

Why the other options are wrong

  • A. MLPs are feedforward networks that treat inputs independently and cannot effectively model sequential dependencies or context.
  • B. RNNs can handle sequential data and some dependencies, but they struggle with very long sequences due to vanishing/exploding gradients and limited memory.
  • D. CNNs are primarily designed for spatial data (like images) but can be adapted for sequences, though they are less efficient at capturing long-range dependencies than Transformers.

Transformer Architecture

A neural network architecture that relies on a self-attention mechanism to weigh the importance of different parts of the input data, excelling in sequential data and long-range dependencies.

  • Uses self-attention layers.
  • Processes input sequences in parallel.
  • Excellent for capturing long-range dependencies.
  • Foundation for many large language models (LLMs).

Memory trick: Nets learn patterns, some better for sequences.

More AI/ML and Generative AI Fundamentals questions