Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureMedium
A company is developing an application that helps visually impaired individuals navigate their surroundings by providing real-time audio descriptions of what's in front of them. This requires the AI service to interpret the entire visual scene and generate a natural language sentence describing its content. Which computer vision capability is essential for this functionality?
- AImage Captioning
- BImage Classification
- COptical Character Recognition (OCR)
- DObject Detection
Show answer & explanationAnswer & explanation
Correct answer: A. Image Captioning
Image Captioning is specifically designed to generate a natural language description of the entire content of an image, which directly fulfills the requirement of providing real-time audio descriptions of surroundings.
Why the other options are wrong
- B. Image Classification assigns a single label or a few labels to an image, which is insufficient for a detailed, natural language description.
- C. OCR extracts text from images, which is not the primary need for describing visual surroundings.
- D. Object Detection identifies and locates individual objects but does not generate a coherent sentence describing the entire scene.
Image Captioning
A computer vision task that generates a natural language description (caption) of the content within an image.
- Combines computer vision and natural language processing.
- Aims to provide a human-like understanding of image content.
- Useful for accessibility, content indexing, and storytelling.
Memory trick: Captioning gives images a voice.