Microsoft Certified: Azure AI Engineer AssociateImplement generative AI solutionsHard

A software development team is using Azure OpenAI Service to generate code documentation. They find that the model sometimes generates documentation that is plausible but contains subtle errors or outdated information that could lead to incorrect usage. They also notice that the model occasionally 'hallucinates' API endpoints that don't exist. To improve the factual accuracy and reduce hallucinations, which advanced technique should they implement?

  1. AImplementing Retrieval Augmented Generation (RAG) with their internal, authoritative documentation.
  2. BIncreasing the 'presence_penalty' to discourage unique but incorrect information.
  3. CSetting the 'temperature' parameter to a very low value (e.g., 0.1) to make output more deterministic.
  4. DFine-tuning the model on a large corpus of correct and up-to-date documentation.
Show answer & explanation

Correct answer: A. Implementing Retrieval Augmented Generation (RAG) with their internal, authoritative documentation.

Retrieval Augmented Generation (RAG) is specifically designed to ground the LLM's responses in authoritative, external knowledge. By retrieving relevant information from the company's internal documentation and providing it to the LLM as context, RAG significantly improves factual accuracy and reduces hallucinations.

Why the other options are wrong

  • B. Presence penalty reduces repetition but doesn't guarantee factual accuracy or prevent hallucinations of new, incorrect facts.
  • C. Low temperature makes output more deterministic but doesn't prevent factual errors or hallucinations if the underlying 'knowledge' in the model is incorrect or incomplete.
  • D. Fine-tuning can help with style and domain-specific language, but it's expensive, time-consuming, and doesn't guarantee the model learns all facts or prevents hallucinations when confronted with new queries outside its training data; it essentially 'burns in' knowledge, which can still become outdated.

Retrieval Augmented Generation (RAG)

An architecture that enhances LLMs by retrieving relevant information from an external knowledge base and feeding it as context into the prompt.

  • Reduces hallucinations and improves factual accuracy.
  • Keeps information current without retraining the model.
  • Combines the generative power of LLMs with reliable data sources.

Memory trick: RAG Rescues Against Generative Ghosts: RAG fights the 'ghosts' of hallucinations.

More Implement generative AI solutions questions