Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsEasy

A municipal government is developing an application to help visually impaired citizens navigate public spaces. The application needs to provide real-time audio descriptions of text encountered in their environment, such as street signs, building names, and informational placards. This requires immediate processing of camera input to extract text and convert it to speech. Which Azure Computer Vision API is the most suitable for the text extraction component?

  1. ADescribe Image
  2. BRead API
  3. CObject Detection
  4. DAnalyze Image (Categories)
Show answer & explanation

Correct answer: B. Read API

The Read API in Azure Computer Vision is specifically designed for high-accuracy OCR, efficiently extracting text from diverse real-world images, including signs and placards, which is crucial for a real-time navigation aid.

Why the other options are wrong

  • A. Describe Image generates a descriptive sentence about an image, not specific text extraction.
  • C. Object Detection identifies objects but does not extract text from them.
  • D. Analyze Image (Categories) classifies images into broad categories, not text extraction.

Computer Vision Read API

An advanced OCR capability within Azure Computer Vision for high-accuracy text extraction from various image types, including signs, documents, and labels.

  • Supports both printed and handwritten text
  • Handles various orientations and image qualities
  • Provides text lines and word bounding box locations
  • Ideal for real-time text-to-speech scenarios

Memory trick: Read API 'reads' the world for you, so your app can speak it.

More Implement image and video processing solutions questions