Microsoft Certified: Azure AI Engineer AssociateImplement generative AI solutionsMedium
A healthcare provider is using Azure OpenAI Service to assist doctors in drafting patient summaries. Due to the sensitive nature of health information, it is critical to prevent the model from generating any personally identifiable information (PII) that might inadvertently appear in the prompt. Which responsible AI practice should be prioritized to address this specific concern?
- AImplementing robust content filtering on the model's output.
- BAnonymizing or de-identifying input data before sending it to Azure OpenAI.
- CRegularly auditing the model's responses for factual accuracy.
- DConfiguring the model's 'temperature' parameter to a low value.
Show answer & explanationAnswer & explanation
Correct answer: B. Anonymizing or de-identifying input data before sending it to Azure OpenAI.
The most effective way to prevent PII leakage from the model is to ensure that PII is never sent to the model in the first place. Anonymizing or de-identifying input data is a crucial pre-processing step for sensitive information.
Why the other options are wrong
- A. Content filtering on output is important but is a reactive measure; preventing PII from entering the model is proactive and more secure.
- C. Factual accuracy is important but doesn't directly address the risk of generating sensitive PII.
- D. Temperature affects randomness but doesn't prevent the generation of specific types of content like PII if it was in the input.
Data Anonymization for LLMs
The process of removing or encrypting personally identifiable information (PII) from data before it is used as input for large language models.
- Crucial for protecting privacy and complying with regulations (e.g., HIPAA, GDPR).
- Prevents models from inadvertently generating or retaining PII.
- A proactive measure to mitigate data leakage risks.
Memory trick: Anonymize All Inputs, No PII Leaks: Make sure your inputs are anonymous to prevent any PII leaks.