Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsMedium
A developer is building a voice-controlled application for a specialized medical device. The application needs to accurately recognize specific medical terminology and commands that are not commonly found in general language models. The device operates in a quiet, controlled environment, so background noise is minimal. Which Azure AI Speech customization approach should the developer prioritize to achieve the highest accuracy for this unique vocabulary?
- ANeural Text-to-Speech (TTS) voice customization.
- BCustom acoustic model training.
- CCustom language model training.
- DSpeaker diarization model customization.
Show answer & explanationAnswer & explanation
Correct answer: C. Custom language model training.
Custom language model training is essential when dealing with specialized vocabulary or domain-specific terminology. Since the environment is quiet, the primary challenge is recognizing specific words, which a custom language model directly addresses by increasing the probability of these terms being recognized correctly.
Why the other options are wrong
- A. Neural Text-to-Speech (TTS) customization is for generating speech output, not for improving input speech recognition.
- B. Custom acoustic model training is for improving recognition in unique acoustic environments or with specific accents, which is not the primary issue here given the 'quiet, controlled environment'.
- D. Speaker diarization separates different speakers in an audio stream, which is irrelevant to recognizing specialized vocabulary from a single speaker.
Custom Language Model (Speech)
A specialized model trained to recognize domain-specific vocabulary, phrases, and grammar, improving the accuracy of speech recognition for particular applications or industries.
- Enhances recognition of unique terminology.
- Improves accuracy for specific use cases (e.g., medical, legal).
- Complements the base acoustic model.
Memory trick: Words are the language model's domain, sounds are the acoustic's reign.