Microsoft Certified: Azure AI Engineer AssociateImplement natural language processing solutionsMedium
A call center wants to implement a real-time voice assistant that can understand customer queries and provide instant responses. The assistant needs to process spoken language, identify the customer's intent, extract relevant entities, and then generate a spoken response. Which sequence of Azure AI services best supports this real-time interaction?
- AAzure AI Speech (STT) -> Azure AI Language (NER) -> Azure AI Language (Sentiment) -> Azure AI Speech (TTS)
- BAzure AI Speech (STT) -> Azure AI Translator (Text) -> Azure AI Speech (TTS)
- CAzure AI Language (Key Phrase Extraction) -> Azure AI Speech (TTS) -> Azure OpenAI Service
- DAzure AI Speech (STT) -> Language Understanding (LUIS) -> Azure AI Speech (TTS)
Show answer & explanationAnswer & explanation
Correct answer: D. Azure AI Speech (STT) -> Language Understanding (LUIS) -> Azure AI Speech (TTS)
This sequence directly addresses the core requirements: Speech-to-Text (STT) to convert spoken query, Language Understanding (LUIS) to identify intent and extract entities, and Text-to-Speech (TTS) to generate a spoken response. LUIS is specifically designed for intent and entity recognition in conversational AI scenarios.
Why the other options are wrong
- A. While using relevant services, it includes Sentiment Analysis and NER separately, which are often part of or secondary to LUIS's core function for intent/entity in conversational flows.
- B. Azure AI Translator is for language translation, not intent recognition, making this sequence incorrect for the scenario.
- C. Key Phrase Extraction and Azure OpenAI are not typically used in this order for real-time conversational intent and response generation, and it misses the initial STT.
Real-time Conversational AI Flow
The typical sequence of operations for a voice-enabled conversational AI system, involving converting speech to text, understanding the text's meaning (intent and entities), and then converting a generated text response back into speech.
- Relies on seamless integration of Speech-to-Text (STT), Natural Language Understanding (NLU), and Text-to-Speech (TTS).
- NLU component (like LUIS) is critical for identifying user intent and extracting relevant information.
- Must be optimized for low latency to enable real-time interaction.
Memory trick: Speak, Understand, Respond.