Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsMedium
A retail company wants to analyze customer behavior in their stores. They need to automatically detect when a customer enters a specific product aisle and estimate how long they stay there. They also want to identify if the customer picks up a product. Which combination of Azure Computer Vision capabilities should be used?
- ARead API and Object Detection
- BImage Analysis (Tags) and Read API
- CSpatial Analysis and Object Detection
- DForm Recognizer and Spatial Analysis
Show answer & explanationAnswer & explanation
Correct answer: C. Spatial Analysis and Object Detection
Azure Computer Vision's Spatial Analysis can track people's movements within defined zones, ideal for detecting entry/exit and dwell time. Object Detection can identify specific actions like picking up a product when trained on relevant images.
Why the other options are wrong
- A. Read API is for text extraction; not suitable for behavior analysis.
- B. Image Analysis (Tags) provides general image descriptions, not specific behavioral tracking, and Read API is for text.
- D. Form Recognizer is for document processing; not relevant here.
Computer Vision Spatial Analysis + Object Detection
Combining Spatial Analysis for tracking people in spaces with Object Detection for identifying specific items or actions.
- Spatial Analysis tracks movement in defined zones (e.g., enter, exit, dwell time).
- Object Detection identifies objects within an image (e.g., a hand picking up a product).
- Together they enable comprehensive behavior analysis in physical spaces.
Memory trick: Spatial Analysis watches where they go, Object Detection sees what they grab.