Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureMedium
A researcher is studying wildlife populations and needs to count specific animal species in aerial photographs. The task requires not only identifying the species (e.g., 'deer', 'bear') but also drawing a bounding box around each individual animal to get an accurate count and location. Which type of computer vision model should the researcher build or utilize?
- AObject Detection model
- BImage Captioning model
- CImage Classification model
- DSemantic Segmentation model
Show answer & explanationAnswer & explanation
Correct answer: A. Object Detection model
An Object Detection model is necessary because it can identify multiple instances of different objects within a single image and provide their precise locations with bounding boxes, which is crucial for counting individual animals.
Why the other options are wrong
- B. An Image Captioning model generates a descriptive sentence for the entire image, not individual object counts or locations.
- C. An Image Classification model would only tell if an image *contains* a deer or bear, not how many or where they are.
- D. A Semantic Segmentation model classifies every pixel, outlining the shape of objects, but an object detection model is more direct and efficient for counting and bounding distinct instances.
Object Detection vs. Classification
Object detection identifies and locates multiple objects within an image, while image classification assigns a single label to the entire image.
- Classification: 'This image contains a dog.'
- Detection: 'There is a dog at [x,y,w,h] and another dog at [x2,y2,w2,h2].'
- Both are fundamental computer vision tasks.
Memory trick: To count individuals, you must detect them first.