A developer is building a voice-controlled application for factory workers that operates in a loud, industrial environment. The application needs to accurately transcribe short voice commands (e.g., 'stop machine', 'start conveyor') despite significant background machinery noise. The commands use specific, domain-specific terminology. What combination of Azure AI Speech features should the developer prioritize to achieve high accuracy?
- ANeural Text-to-Speech and Language Understanding (LUIS)
- BBatch Speech-to-Text and Key Phrase Extraction
- CCustom Speech (Acoustic Model) and Custom Speech (Language Model)
- DCustom Neural Voice and Speaker Diarization
Show answer & explanationAnswer & explanation
Correct answer: C. Custom Speech (Acoustic Model) and Custom Speech (Language Model)
To achieve high accuracy in a noisy, domain-specific environment, both a custom acoustic model and a custom language model are essential. A custom acoustic model (part of Custom Speech) adapts to the specific noise and speaking styles of the factory. A custom language model (also part of Custom Speech) helps the speech-to-text engine recognize domain-specific terminology more accurately. This combination directly addresses both challenges.
Why the other options are wrong
- A. Neural Text-to-Speech generates speech, and LUIS understands intent, not transcription accuracy for specific audio/vocabulary.
- B. Batch Speech-to-Text is for processing large audio files, and Key Phrase Extraction is for text analysis; neither directly improves real-time transcription accuracy in noisy, domain-specific contexts.
- D. Custom Neural Voice generates speech, and Speaker Diarization identifies speakers; neither improves transcription accuracy in noise or for domain terms.
Custom Speech (Acoustic & Language Models)
Azure AI Speech features that allow training speech-to-text models to adapt to specific acoustic conditions (noise, accents) and domain-specific vocabulary, respectively.
- Custom acoustic models improve recognition in noisy environments.
- Custom language models enhance recognition of domain-specific words.
- Both are crucial for high accuracy in specialized, challenging scenarios.
Memory trick: To understand factory talk, you need a custom ear for the noise AND a custom dictionary for the jargon.