Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsMedium
A global e-commerce company wants to enhance its product catalog by automatically generating descriptive captions for images uploaded by vendors. These captions should be natural-sounding sentences summarizing the visual content of the product. For example, an image of a 'red leather handbag on a wooden table' should generate a caption like 'A red leather handbag sits on a wooden table.' Which Azure Computer Vision capability should they use?
- AAnalyze Image (Tags)
- BRead API
- CObject Detection
- DDescribe Image
Show answer & explanationAnswer & explanation
Correct answer: D. Describe Image
The Describe Image capability of Azure Computer Vision is specifically designed to generate human-readable, natural language sentences that describe the content of an image, making it ideal for creating product captions.
Why the other options are wrong
- A. Analyze Image (Tags) returns a list of keywords, not a descriptive sentence.
- B. Read API extracts text from images, which is not the goal here.
- C. Object Detection identifies and localizes specific objects but does not generate a descriptive sentence.
Computer Vision Describe Image
An Azure Computer Vision capability that generates a human-readable, natural language sentence describing the visual content of an image.
- Creates concise, descriptive captions
- Identifies main objects, actions, and scenes
- Can return multiple caption candidates with confidence scores
- Useful for accessibility and content management
Memory trick: Describe Image 'tells' you what's in the picture, like a good storyteller.