Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsMedium

A global e-commerce company wants to enhance its product catalog by automatically generating descriptive captions for images uploaded by vendors. These captions should be natural-sounding sentences summarizing the visual content of the product. For example, an image of a 'red leather handbag on a wooden table' should generate a caption like 'A red leather handbag sits on a wooden table.' Which Azure Computer Vision capability should they use?

  1. AAnalyze Image (Tags)
  2. BRead API
  3. CObject Detection
  4. DDescribe Image
Show answer & explanation

Correct answer: D. Describe Image

The Describe Image capability of Azure Computer Vision is specifically designed to generate human-readable, natural language sentences that describe the content of an image, making it ideal for creating product captions.

Why the other options are wrong

  • A. Analyze Image (Tags) returns a list of keywords, not a descriptive sentence.
  • B. Read API extracts text from images, which is not the goal here.
  • C. Object Detection identifies and localizes specific objects but does not generate a descriptive sentence.

Computer Vision Describe Image

An Azure Computer Vision capability that generates a human-readable, natural language sentence describing the visual content of an image.

  • Creates concise, descriptive captions
  • Identifies main objects, actions, and scenes
  • Can return multiple caption candidates with confidence scores
  • Useful for accessibility and content management

Memory trick: Describe Image 'tells' you what's in the picture, like a good storyteller.

More Implement image and video processing solutions questions