Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureMedium

A startup is developing an application that helps visually impaired users understand their surroundings. The application needs to describe the content of an image in natural language, for example, 'A young woman holding a cup of coffee at a cafe.' Which Azure AI Computer Vision capability would be most suitable?

  1. AObject Detection
  2. BCustom Vision
  3. CImage Captioning
  4. DImage Tagging
Show answer & explanation

Correct answer: C. Image Captioning

Image Captioning generates a descriptive sentence in natural language that summarizes the content of an image, which directly matches the requirement for describing image content to visually impaired users.

Why the other options are wrong

  • A. Object Detection identifies and locates objects within an image, but does not describe the overall scene in natural language.
  • B. Custom Vision is a tool for training custom image classification or object detection models, not a direct capability for generating captions.
  • D. Image Tagging assigns keywords or tags to an image, but does not generate a full descriptive sentence.

Image Captioning

An Azure AI Computer Vision capability that generates a natural language description (caption) of the content within an image.

  • Provides a concise summary of the visual content.
  • Useful for accessibility, content indexing, and search.
  • Differs from tagging, which provides keywords, and object detection, which identifies specific items.

Memory trick: Images can be tagged, detected, or fully described.

More Describe features of computer vision workloads on Azure questions