Microsoft Certified: Azure AI Engineer AssociateImplement generative AI solutionsHard
A team is developing a generative AI application using Azure OpenAI Service for a regulated industry. The application will process sensitive client information to generate personalized reports. Due to strict compliance requirements, it is crucial that the raw sensitive data is never directly exposed to the large language model (LLM) or stored persistently by the service. However, the LLM still needs to understand the context of the sensitive information to generate accurate reports. Which prompt engineering technique can best address this challenge while adhering to data privacy and compliance?
- AFew-shot prompting with example sensitive data
- BZero-shot prompting with anonymized placeholders
- CInstruction-based prompting with direct sensitive data insertion
- DStructured output prompting with sensitive data fields
Show answer & explanationAnswer & explanation
Correct answer: B. Zero-shot prompting with anonymized placeholders
Using anonymized placeholders (e.g., '[ClientName]', '[AccountBalance]') in a zero-shot prompt allows the LLM to understand the structure and type of information required without ever seeing or processing the actual sensitive data. The placeholders are replaced with real data only after the LLM generates the report structure, ensuring compliance.
Why the other options are wrong
- A. Few-shot prompting with example sensitive data would directly expose sensitive information to the model, violating compliance.
- C. Direct sensitive data insertion is explicitly what needs to be avoided due to compliance requirements.
- D. While structured output is useful, directly including sensitive data fields in the prompt still exposes the data to the LLM.
Data Anonymization for LLMs
The process of modifying sensitive data before it is input to an LLM to remove or obscure personally identifiable information, while retaining its utility for the model.
- Crucial for privacy and compliance in regulated industries.
- Can involve tokenization, generalization, or using placeholders.
- Allows LLMs to process context without directly handling sensitive raw data.
Memory trick: Privacy first: anonymize data before it reaches the generative mind.