AWS Certified AI PractitionerAI/ML and Generative AI FundamentalsHard

A team is training a large language model (LLM) for a specific domain, such as legal documents. They have a vast amount of general text data but only a limited, highly specialized dataset of legal texts. To adapt the LLM effectively to the legal domain, what is the most appropriate technique, assuming the base LLM is already pre-trained on general data?

  1. AIncreasing the number of layers in the pre-trained LLM.
  2. BTraining from scratch with only the legal dataset.
  3. CPerforming unsupervised clustering on the legal dataset.
  4. DFine-tuning the pre-trained LLM with the legal dataset.
Show answer & explanation

Correct answer: D. Fine-tuning the pre-trained LLM with the legal dataset.

Fine-tuning is the most effective approach in this scenario. It involves taking a pre-trained model (which has learned general language patterns from a large corpus) and further training it on a smaller, specific dataset. This allows the model to adapt its learned knowledge to the nuances of the legal domain without requiring massive amounts of domain-specific data or the computational cost of training from scratch.

Why the other options are wrong

  • A. Increasing model layers without proper re-training or fine-tuning, especially with limited data, would be ineffective and could even worsen performance or lead to instability.
  • B. Training from scratch with a limited specialized dataset would likely lead to underfitting and poor performance, as the model would not learn general language patterns effectively.
  • C. Unsupervised clustering would group similar legal documents but would not adapt the LLM to generate or understand legal language; it's a data analysis technique, not a model adaptation technique.

Fine-tuning (LLMs)

The process of taking a pre-trained large language model and further training it on a smaller, domain-specific dataset to adapt its capabilities to a particular task or domain.

  • Requires a base pre-trained model.
  • Uses a smaller, specialized dataset.
  • Adapts model to specific tasks/domains.
  • More efficient than training from scratch.

Memory trick: LLMs 'learn' general, then get 'fine-tuned' for 'specific' expertise.

More AI/ML and Generative AI Fundamentals questions