Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureMedium
A smart city initiative aims to analyze urban traffic flow by understanding the density and movement of vehicles and pedestrians at various intersections. This requires not just identifying objects, but also understanding the specific areas occupied by different classes (e.g., car, pedestrian, bicycle) within an image or video frame. Which computer vision capability is BEST suited for this detailed spatial understanding?
- AImage Classification
- BSemantic Segmentation
- CObject Detection
- DOptical Character Recognition (OCR)
Show answer & explanationAnswer & explanation
Correct answer: B. Semantic Segmentation
Semantic Segmentation precisely classifies each pixel in an image to a predefined class, enabling detailed understanding of the exact shape and area occupied by different objects like vehicles and pedestrians, which is crucial for density and movement analysis.
Why the other options are wrong
- A. Image Classification assigns a single label to an entire image, which is insufficient for detailed object area and density analysis.
- C. Object Detection draws bounding boxes around objects, identifying their general location but not their precise shape or pixel-level coverage.
- D. OCR is for extracting text from images, which is unrelated to analyzing traffic object density.
Semantic Segmentation
A computer vision task that involves classifying each pixel in an image into a predefined category, providing a detailed, pixel-level understanding of the scene.
- Assigns a class label to every pixel.
- Provides precise object boundaries and shapes.
- Useful for tasks requiring detailed spatial analysis, such as autonomous driving, medical imaging, and urban planning.
Memory trick: Segmentation slices up pixels for meaning.