Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsMedium
A developer is building a mobile application that allows users to dictate notes and reminders. The application needs to accurately convert spoken language into text, even in environments with moderate background noise, and support various accents. Which Azure AI Speech feature should the developer use, and what is a crucial consideration for improving its accuracy?
- ASpeech-to-Text; providing custom acoustic models
- BNeural Text-to-Speech (TTS); providing custom voice fonts
- CLanguage Understanding (LUIS); providing intent examples
- DCustom Neural Voice; providing speaker diarization data
Show answer & explanationAnswer & explanation
Correct answer: A. Speech-to-Text; providing custom acoustic models
Speech-to-Text is the core feature for converting spoken language to text. To improve accuracy in noisy environments and for various accents, providing custom acoustic models (which adapt to specific acoustic conditions or speaking styles) is crucial. TTS is for generating speech, Custom Neural Voice is for creating unique voices, and LUIS is for understanding intent, not transcription accuracy.
Why the other options are wrong
- B. Neural TTS converts text to speech, and custom voice fonts are not relevant for improving speech-to-text accuracy.
- C. LUIS is for understanding user intent from text/speech, not for improving the accuracy of speech-to-text transcription itself.
- D. Custom Neural Voice creates a synthetic voice, and speaker diarization identifies speakers, neither is for improving transcription accuracy.
Custom Acoustic Model (Speech-to-Text)
A specialized model trained within Azure AI Speech-to-Text to improve transcription accuracy for specific acoustic environments, speaker accents, or speaking styles.
- Enhances transcription accuracy for challenging audio.
- Requires audio data for training.
- Complements custom language models for domain-specific vocabulary.
Memory trick: To hear correctly, you need a good ear (acoustic model) for the speech.