Microsoft Certified: Azure AI Engineer AssociateImplement generative AI solutionsMedium

A company is developing an internal knowledge base chatbot using Azure OpenAI Service. The chatbot needs to provide answers based on a large collection of proprietary documents, but it frequently generates responses that are factually incorrect or hallucinated, even when the user's query is directly related to the documents. The company wants to improve the factual accuracy of the chatbot's responses by ensuring it always references the provided documents. Which architectural pattern is most suitable for this requirement?

  1. ARetrieval Augmented Generation (RAG)
  2. BDirect Prompting with increased temperature
  3. CFine-tuning the base model with proprietary data
  4. DUsing a larger, more general-purpose model
Show answer & explanation

Correct answer: A. Retrieval Augmented Generation (RAG)

Retrieval Augmented Generation (RAG) is specifically designed to address the problem of hallucinations and improve factual accuracy by retrieving relevant information from an external knowledge base and feeding it to the LLM as context before generating a response. This grounds the model's output in verifiable data.

Why the other options are wrong

  • B. Increasing temperature would make the model more creative and prone to hallucinations, not less.
  • C. Fine-tuning can teach the model a domain-specific style or terminology, but it doesn't guarantee factual accuracy or prevent hallucinations for new, unseen information.
  • D. A larger model might have more general knowledge but is still prone to hallucinations if not grounded in specific, external data for the query.

Retrieval Augmented Generation (RAG)

An architectural pattern where an LLM is augmented with a retrieval system to fetch relevant information from a knowledge base before generating a response.

  • Reduces model hallucinations by grounding responses in external data.
  • Improves factual accuracy and allows for citation of sources.
  • Involves a retriever (e.g., vector search) and a generator (LLM).

Memory trick: To make LLMs truthful, retrieve facts and then generate.

More Implement generative AI solutions questions