Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsMedium

A media studio is creating an interactive audio experience where users can ask questions and receive responses in a consistent, branded voice that is unique to their intellectual property. The studio wants to generate highly realistic, expressive speech from text using a voice that sounds like a specific character. Which Azure AI Speech feature should the studio use?

  1. ABatch Speech-to-Text
  2. BNeural Text-to-Speech (TTS)
  3. CCustom Neural Voice
  4. DSpeaker Diarization
Show answer & explanation

Correct answer: C. Custom Neural Voice

Custom Neural Voice allows creating a unique, high-quality synthetic voice that matches a specific character or brand. This requires training with audio samples of the desired voice. While Neural TTS provides high-quality speech, it uses pre-built voices, not a custom one.

Why the other options are wrong

  • A. Batch Speech-to-Text converts audio to text, which is not relevant for generating speech.
  • B. Neural TTS generates high-quality speech but uses pre-defined voices, not a custom, branded voice.
  • D. Speaker Diarization identifies who spoke when in an audio file, not for generating speech.

Custom Neural Voice

An Azure AI Speech feature that allows organizations to create a unique, high-quality synthetic voice that represents their brand or a specific character.

  • Requires a significant amount of high-quality audio data of the target voice.
  • Generates highly expressive and natural-sounding speech.
  • Provides brand consistency for voice experiences.

Memory trick: To make your unique voice, you need Custom Neural Voice, not just any robotic sound.

More Implement natural language processing solutions questions