Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsMedium
A developer is building a mobile application that allows users to dictate notes and reminders. The application needs to transcribe the audio into text accurately, even in environments with background noise and for users with diverse accents. Which Azure AI Speech feature, combined with a custom model, would best address these requirements?
- ACustom Speech
- BBatch Speech-to-Text
- CNeural Text-to-Speech
- DSpeaker Diarization
Show answer & explanationAnswer & explanation
Correct answer: A. Custom Speech
Custom Speech allows developers to create custom speech-to-text models tailored to specific vocabularies, acoustic environments, and accents. This directly addresses the need for accurate transcription in noisy environments and for diverse accents, which the base model might struggle with.
Why the other options are wrong
- B. Batch Speech-to-Text is for transcribing large audio files asynchronously, but doesn't inherently improve accuracy for challenging audio without custom models.
- C. Neural Text-to-Speech converts text to speech, not speech to text.
- D. Speaker Diarization identifies different speakers in an audio file, but doesn't improve transcription accuracy.
Azure Custom Speech
A feature of Azure AI Speech that allows users to create custom speech-to-text models by providing audio and text data to improve recognition accuracy for specific domains, accents, and acoustic environments.
- Improves accuracy for domain-specific vocabulary and jargon.
- Adapts to various accents and speech patterns.
- Can be trained with acoustical data (audio files) and linguistic data (text transcripts).
Memory trick: Customize speech for better listening and speaking.