Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureMedium
A developer is creating a mobile application that allows users to quickly identify famous landmarks by simply pointing their phone camera at them. The application needs to return the name of the landmark (e.g., 'Eiffel Tower', 'Statue of Liberty'). Which computer vision capability is most appropriate for this task?
- AOptical Character Recognition (OCR)
- BObject Detection
- CFace Detection
- DImage Classification
Show answer & explanationAnswer & explanation
Correct answer: D. Image Classification
Image Classification is the most appropriate capability as it takes an image as input and assigns a single category label to the entire image, which in this case would be the name of the famous landmark.
Why the other options are wrong
- A. OCR extracts text from images, which is not the primary goal here.
- B. Object Detection identifies and locates *multiple* distinct objects within an image; while a landmark is an object, the primary goal is to classify the *entire image* as that landmark.
- C. Face Detection is specialized for human faces and irrelevant for landmark identification.
Image Classification
A computer vision task that assigns a single category label to an entire input image.
- Answers the question 'What is in this image?' at a high level.
- Outputs a probability distribution over a set of predefined classes.
- Used for content moderation, visual search, and organizing image libraries.
Memory trick: Classify the whole image to name a landmark.