A red team assesses a facial-recognition-based access control system. Without any access to the model's training pipeline, the team crafts subtle, carefully calculated pixel perturbations to a photo of an unauthorized team member's face. When presented to the camera, the modified image is misclassified by the model as an authorized employee, even though the perturbations are imperceptible to a human observer. What type of attack does this describe?
- AModel extraction attack
- BAdversarial example (evasion) attack
- CIndirect prompt injection
- DData poisoning attack
Show answer & explanationAnswer & explanation
Correct answer: B. Adversarial example (evasion) attack
An adversarial example (evasion) attack crafts inputs with small, often imperceptible perturbations specifically designed to cause a trained model to misclassify at inference time, without altering the model's training data or internals. Data poisoning corrupts the training data itself, model extraction attempts to steal or replicate model functionality, and indirect prompt injection manipulates an LLM through untrusted content it processes.
Why the other options are wrong
- A. Model extraction attempts to reconstruct or steal the model's parameters/logic through repeated queries, not to cause misclassification.
- C. Indirect prompt injection targets LLMs processing untrusted text content, not image-based classifiers.
- D. Data poisoning corrupts the training dataset over time, which is different from manipulating a single inference-time input.
Adversarial Example (Evasion) Attack
An AI attack technique where an attacker crafts input data with small, often imperceptible perturbations specifically designed to cause a trained machine learning model to misclassify or behave incorrectly at inference time.
- Occurs at inference time, not during training
- Perturbations are often invisible or negligible to humans but significant to the model's decision boundary
- Commonly demonstrated against image classifiers and facial recognition systems
Memory trick: Adversarial example = a magic disguise only the model sees