CompTIA PenTest+ (PT0-003)Attacks and ExploitsMedium

A penetration tester is assessing an organization's Machine Learning (ML) model used for fraud detection. The tester repeatedly submits carefully crafted, slightly perturbed inputs to the model and observes the model's predictions. The goal is to understand the model's decision boundaries and potentially create inputs that cause misclassifications. Which type of AI attack is BEST described by this scenario?

  1. AModel Extraction Attack
  2. BData Poisoning Attack
  3. CIndirect Prompt Injection
  4. DAdversarial Example (Evasion) Attack
Show answer & explanation

Correct answer: D. Adversarial Example (Evasion) Attack

The scenario describes an attacker making 'carefully crafted, slightly perturbed inputs' to 'cause misclassifications' in an ML model. This is the definition of an Adversarial Example (Evasion) Attack, where inputs are designed to fool the model into making incorrect predictions without necessarily understanding its internal parameters or corrupting its training data.

Why the other options are wrong

  • A. Model Extraction (or Model Stealing) aims to replicate a proprietary model's functionality or architecture by querying it.
  • B. Data Poisoning involves injecting malicious data into the training set to corrupt the model's future behavior.
  • C. Indirect Prompt Injection targets large language models (LLMs) by embedding malicious instructions in external data that the LLM processes, which is not the focus here.

Adversarial Example (Evasion) Attack

An attack against machine learning models where an attacker creates subtly altered inputs (adversarial examples) that are designed to be misclassified by the model, while appearing normal to humans.

  • Targets the inference phase of an ML model.
  • Aims to cause incorrect predictions (evasion).
  • Often involves small, imperceptible perturbations to input data.
  • Can be 'black-box' (querying model) or 'white-box' (accessing model internals).

Memory trick: Adversarial examples evade, subtly, with slight changes.

More Attacks and Exploits questions