Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsMedium
A media studio is creating an interactive audio experience where users can ask questions and receive responses in a consistent, branded voice that is unique to their intellectual property. The studio wants to generate highly realistic, expressive speech from text using a voice that sounds like a specific character. Which Azure AI Speech feature should the studio use?
- ABatch Speech-to-Text
- BNeural Text-to-Speech (TTS)
- CCustom Neural Voice
- DSpeaker Diarization
Show answer & explanationAnswer & explanation
Correct answer: C. Custom Neural Voice
Custom Neural Voice allows creating a unique, high-quality synthetic voice that matches a specific character or brand. This requires training with audio samples of the desired voice. While Neural TTS provides high-quality speech, it uses pre-built voices, not a custom one.
Why the other options are wrong
- A. Batch Speech-to-Text converts audio to text, which is not relevant for generating speech.
- B. Neural TTS generates high-quality speech but uses pre-defined voices, not a custom, branded voice.
- D. Speaker Diarization identifies who spoke when in an audio file, not for generating speech.
Custom Neural Voice
An Azure AI Speech feature that allows organizations to create a unique, high-quality synthetic voice that represents their brand or a specific character.
- Requires a significant amount of high-quality audio data of the target voice.
- Generates highly expressive and natural-sounding speech.
- Provides brand consistency for voice experiences.
Memory trick: To make your unique voice, you need Custom Neural Voice, not just any robotic sound.