AWS Certified AI PractitionerFoundation ModelsMedium

An AI solutions architect is evaluating different approaches for a client who wants to build a custom intelligent assistant for their specialized legal firm. The assistant needs to answer complex legal questions based on the firm's private document repository. The architect decides to use a pre-trained Large Language Model (LLM) and augment it with the firm's data. Which technique is most appropriate for integrating the firm's private legal documents into the LLM's knowledge base without retraining the entire model?

  1. AFull fine-tuning of the LLM on the private legal documents.
  2. BCompletely training a new LLM from scratch using only the firm's data.
  3. CPrompt engineering combined with Retrieval Augmented Generation (RAG).
  4. DReducing the LLM's parameter count to fit the specialized dataset.
Show answer & explanation

Correct answer: C. Prompt engineering combined with Retrieval Augmented Generation (RAG).

Retrieval Augmented Generation (RAG) is a highly effective technique where an LLM's response is generated based on information retrieved from an external knowledge base (like the firm's private legal documents) in real-time, guided by prompt engineering. This avoids costly and time-consuming full model retraining while providing up-to-date and specific information.

Why the other options are wrong

  • A. Full fine-tuning is expensive and may still lead to hallucination if the model isn't continuously updated with new legal info. RAG is more efficient for dynamic, external knowledge.
  • B. Training an LLM from scratch is prohibitively expensive and time-consuming for most organizations and unnecessary when a powerful pre-trained LLM exists.
  • D. Reducing parameter count would likely degrade the LLM's general capabilities and wouldn't directly integrate new knowledge effectively for a specialized task.

Retrieval Augmented Generation (RAG)

An AI technique that enhances the capabilities of a Large Language Model (LLM) by allowing it to retrieve relevant information from an external knowledge base before generating a response.

  • Combats LLM hallucinations by grounding responses in facts.
  • Enables LLMs to access and use up-to-date or private information.
  • Reduces the need for continuous model retraining for new data.

Memory trick: RAG Retrieves Answers, Grounding Generations.

More Foundation Models questions