Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureMedium

A company is developing an application that helps visually impaired individuals navigate their surroundings by providing real-time audio descriptions of what's in front of them. This requires the AI service to interpret the entire visual scene and generate a natural language sentence describing its content. Which computer vision capability is essential for this functionality?

  1. AImage Captioning
  2. BImage Classification
  3. COptical Character Recognition (OCR)
  4. DObject Detection
Show answer & explanation

Correct answer: A. Image Captioning

Image Captioning is specifically designed to generate a natural language description of the entire content of an image, which directly fulfills the requirement of providing real-time audio descriptions of surroundings.

Why the other options are wrong

  • B. Image Classification assigns a single label or a few labels to an image, which is insufficient for a detailed, natural language description.
  • C. OCR extracts text from images, which is not the primary need for describing visual surroundings.
  • D. Object Detection identifies and locates individual objects but does not generate a coherent sentence describing the entire scene.

Image Captioning

A computer vision task that generates a natural language description (caption) of the content within an image.

  • Combines computer vision and natural language processing.
  • Aims to provide a human-like understanding of image content.
  • Useful for accessibility, content indexing, and storytelling.

Memory trick: Captioning gives images a voice.

More Describe features of computer vision workloads on Azure questions