Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsHard

A financial institution is building a chatbot for internal employees to answer questions about company policies and procedures. The chatbot needs to provide accurate, up-to-date answers by referencing a large repository of internal documents. It should also be able to generate human-like responses and cite its sources. Which architecture pattern, combining Azure AI services, is most suitable for this scenario?

  1. ADirect integration of Azure OpenAI Service with a static knowledge base.
  2. BRetrieval Augmented Generation (RAG) using Azure AI Search and Azure OpenAI Service.
  3. CAzure AI Language's QnA Maker integrated with a pre-defined FAQ document.
  4. DCustom Text Classification model to categorize queries and provide canned responses.
Show answer & explanation

Correct answer: B. Retrieval Augmented Generation (RAG) using Azure AI Search and Azure OpenAI Service.

The RAG pattern, combining Azure AI Search for retrieving relevant documents and Azure OpenAI Service for generating responses based on those documents, directly addresses the need for accurate, up-to-date, cited, and human-like answers from a large document repository. This pattern prevents hallucination and ensures grounded responses.

Why the other options are wrong

  • A. Direct OpenAI integration might hallucinate or not cite sources if not explicitly prompted/engineered, and wouldn't efficiently search large document repositories.
  • C. QnA Maker with a pre-defined FAQ is too limited for a 'large repository of internal documents' and dynamic querying; it's best for static FAQs.
  • D. Custom Text Classification and canned responses lack the generative capability and ability to dynamically reference documents for up-to-date answers.

Retrieval Augmented Generation (RAG)

An AI architecture pattern that enhances large language models (LLMs) by first retrieving relevant information from a knowledge base and then using that information to generate more accurate, grounded, and citable responses.

  • Combats LLM 'hallucination' by grounding responses in retrieved data.
  • Enables LLMs to access and utilize up-to-date, proprietary, or specific data.
  • Typically involves two main components: a retriever (e.g., search engine) and a generator (e.g., LLM).

Memory trick: Generate smart answers from vast knowledge.

More Implement natural language processing solutions questions