AWS Certified AI PractitionerFoundation ModelsHard

A machine learning engineer is tasked with fine-tuning a pre-trained foundation model for a highly specialized legal domain. The goal is to enhance the model's performance on legal jargon and specific case law without losing its broad general knowledge. The available legal dataset for fine-tuning is relatively small compared to the original pre-training data. Which fine-tuning technique is most appropriate for efficiently adapting the model while preserving its general capabilities?

  1. AQuantization-aware training
  2. BDomain-Adaptive Pre-training (DAPT)
  3. CPrompt engineering without fine-tuning
  4. DFull fine-tuning of all model parameters
Show answer & explanation

Correct answer: B. Domain-Adaptive Pre-training (DAPT)

Domain-Adaptive Pre-training (DAPT) is a technique where a pre-trained foundation model undergoes an additional pre-training phase on a large, unlabeled dataset specific to the target domain (e.g., legal texts). This allows the model to learn domain-specific vocabulary and patterns without losing its general knowledge, making it more specialized for the legal domain before any task-specific fine-tuning. Given the small labeled dataset for fine-tuning, DAPT helps bridge the gap by leveraging unlabeled domain data first.

Why the other options are wrong

  • A. Quantization-aware training is for model compression and efficiency, not for domain adaptation or preserving general knowledge.
  • C. Prompt engineering alone might not be sufficient for deep domain specialization and understanding legal jargon without model adaptation.
  • D. Full fine-tuning with a small dataset risks overfitting and catastrophic forgetting of general knowledge.

Domain-Adaptive Pre-training (DAPT)

An intermediate pre-training step where a foundation model is further pre-trained on a large, unlabeled dataset specific to a target domain, before task-specific fine-tuning.

  • Helps models specialize in a domain while retaining general knowledge
  • Leverages readily available unlabeled domain data
  • Improves performance on domain-specific tasks with less labeled fine-tuning data

Memory trick: General wisdom, then domain deep dive.

More Foundation Models questions