Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsEasy
A developer is creating a mobile application that needs to accept short voice commands (e.g., 'open app', 'play music') from users. The application also needs to handle variations in pronunciation and speaking styles without requiring extensive custom training data. Which Azure AI Speech feature provides robust, pre-trained speech-to-text capabilities suitable for common commands?
- ASpeech-to-Text (Base Model)
- BNeural Text-to-Speech
- CCustom Speech
- DSpeaker Diarization
Show answer & explanationAnswer & explanation
Correct answer: A. Speech-to-Text (Base Model)
The base Speech-to-Text model from Azure AI Speech is highly robust and pre-trained to handle a wide range of common commands, pronunciations, and speaking styles without the need for custom training, making it suitable for this scenario where extensive customization isn't explicitly required.
Why the other options are wrong
- B. Neural Text-to-Speech converts text to speech, not speech to text.
- C. Custom Speech is for specialized, highly accurate transcription in unique environments or with specific vocabulary, which isn't strictly necessary for 'common commands' without 'extensive custom training'.
- D. Speaker Diarization identifies different speakers, not for transcribing commands.
Azure AI Speech-to-Text (Base Model)
The default, general-purpose speech recognition model provided by Azure AI Speech, which is pre-trained on vast amounts of data to transcribe common speech with high accuracy across various languages, accents, and speaking styles.
- Ready-to-use without any custom training.
- Robust against variations in pronunciation and speaking style.
- Suitable for a wide range of general speech recognition tasks.
Memory trick: Hear and write common words.