CompTIA PenTest+ (PT0-003)Attacks and ExploitsHard

A red team is assessing an organization's Machine Learning (ML) model used for fraud detection. Without any access to the model's training data or internal architecture, the team repeatedly submits carefully varied inputs to the model and observes its outputs. Their goal is to create new inputs that are subtly altered from legitimate transactions but are incorrectly classified as non-fraudulent by the ML model. Which type of AI attack are they performing?

  1. AData Poisoning Attack
  2. BAdversarial Example (Evasion) Attack
  3. CModel Inversion Attack
  4. DModel Extraction Attack
Show answer & explanation

Correct answer: B. Adversarial Example (Evasion) Attack

This scenario describes an Adversarial Example (Evasion) Attack. The red team is crafting inputs that are designed to bypass the ML model's detection capabilities by exploiting its weaknesses. The key characteristics are modifying legitimate inputs subtly and observing output to cause misclassification, without internal knowledge of the model.

Why the other options are wrong

  • A. Data Poisoning involves injecting malicious data into the training set, which is not the case here as they have no access to training data.
  • C. Model Inversion aims to reconstruct sensitive training data from the model's outputs, not to evade detection.
  • D. Model Extraction aims to steal the model's parameters or architecture, not to evade its current classifications.

Adversarial Example (Evasion) Attack

An AI attack where an attacker crafts subtly perturbed inputs that are intentionally designed to fool a machine learning model, causing it to make incorrect predictions or classifications, often in a black-box setting.

  • Subtly alters legitimate inputs.
  • Aims to cause misclassification or evade detection.
  • Can be performed in black-box (no model info) or white-box settings.
  • Exploits model's blind spots or decision boundaries.

Memory trick: AI attacks: Poison data, Invert models, Evade with examples, Extract secrets.

More Attacks and Exploits questions