Microsoft Certified: Azure AI Engineer AssociateImplement generative AI solutionsHard
A company is building a knowledge base chatbot using Azure OpenAI Service. Users frequently ask questions that require retrieving specific information from a large, constantly updated internal document repository. To provide accurate and up-to-date answers without retraining the model, which architecture pattern, leveraging embeddings, is most appropriate?
- AManually updating the model's knowledge base through prompt engineering.
- BUsing a standard GPT model with zero-shot prompting.
- CImplementing Retrieval Augmented Generation (RAG) with embeddings for document search.
- DFine-tuning the GPT model with the entire document repository.
Show answer & explanationAnswer & explanation
Correct answer: C. Implementing Retrieval Augmented Generation (RAG) with embeddings for document search.
Retrieval Augmented Generation (RAG) is the ideal pattern for this scenario. It involves using embeddings to search a knowledge base for relevant documents, then providing those retrieved documents as context to an LLM (like GPT) to generate an informed answer. This avoids costly retraining and ensures up-to-date information.
Why the other options are wrong
- A. Manually updating the model through prompt engineering is not scalable or efficient for a large, constantly updated repository.
- B. Zero-shot prompting with a standard GPT model would not have access to the specific, up-to-date internal document repository.
- D. Fine-tuning is expensive, time-consuming, and would require frequent re-fine-tuning for constantly updated information, making it impractical.
Retrieval Augmented Generation (RAG)
An AI architecture pattern where a large language model (LLM) is augmented by retrieving relevant information from an external knowledge base (often using embeddings) before generating a response, ensuring up-to-date and factual answers.
- Combines information retrieval with LLM generation.
- Addresses LLM's knowledge cutoff and hallucination issues.
- Uses embeddings for efficient semantic search in external data.
Memory trick: RAG is like giving the AI a 'Research Assistant' for fresh info.