Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsEasy

A retail company wants to implement a voice-enabled search feature on its e-commerce website. Users should be able to speak product names or categories, and the system should accurately transcribe their speech into text for searching. The company expects users to speak in various accents and at different paces. Which Azure AI Speech service component is primarily responsible for converting spoken words into written text?

  1. AText-to-speech
  2. BCustom Voice
  3. CSpeech Synthesis Markup Language (SSML)
  4. DSpeech-to-text
Show answer & explanation

Correct answer: D. Speech-to-text

Speech-to-text is the core component of Azure AI Speech that converts spoken audio into written text, which is essential for a voice-enabled search feature. It is designed to handle various accents and speech patterns.

Why the other options are wrong

  • A. Text-to-speech converts written text into spoken audio, the opposite of the requirement.
  • B. Custom Voice allows creating a unique voice model for text-to-speech, not transcribing speech to text.
  • C. SSML is used to control speech synthesis (text-to-speech) output, not speech recognition.

Speech-to-text

A core feature of Azure AI Speech that converts spoken language into written text. It is crucial for applications requiring voice input, such as voice assistants, transcription services, and voice-enabled search.

  • Converts audio input into written text.
  • Supports various languages and accents.
  • Essential for voice interfaces and transcription.

Memory trick: Speech-to-text listens and writes.

More Implement natural language processing solutions questions