CompTIA PenTest+ (PT0-003)Attacks and ExploitsHard

A red team repeatedly submits carefully varied inputs to a company's proprietary fraud-detection API and records the corresponding classification outputs. Using thousands of these input-output pairs, the team trains their own local model that closely mimics the target's decision boundaries, without ever accessing its source code or training data. Which attack technique does this describe?

  1. AAdversarial evasion attack
  2. BData poisoning attack
  3. CPrompt injection attack
  4. DModel extraction (model stealing) attack
Show answer & explanation

Correct answer: D. Model extraction (model stealing) attack

Model extraction attacks abuse an exposed prediction API by systematically querying it and using the outputs to train a substitute model, effectively stealing the intellectual property and internal logic of the original model.

Why the other options are wrong

  • A. Adversarial evasion crafts inputs to cause misclassification at inference time, not to clone the model.
  • B. Data poisoning corrupts the training data of the target model, not replicating it externally.
  • C. Prompt injection manipulates an LLM's instructions, unrelated to reconstructing a classifier's decision logic.

Model Extraction Attack

An attack where an adversary queries a machine learning API extensively and uses the input-output pairs to train a substitute model that replicates the target's functionality, stealing intellectual property.

  • Also called model stealing or model cloning
  • Exploits publicly accessible prediction/inference APIs
  • Mitigated by rate limiting, query monitoring, and output perturbation

Memory trick: Model extraction = photocopying a locked book one page-query at a time

More Attacks and Exploits questions