Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsMedium

A retail company wants to analyze customer behavior in their stores. They need to automatically detect when a customer enters a specific product aisle and estimate how long they stay there. They also want to identify if the customer picks up a product. Which combination of Azure Computer Vision capabilities should be used?

  1. ARead API and Object Detection
  2. BImage Analysis (Tags) and Read API
  3. CSpatial Analysis and Object Detection
  4. DForm Recognizer and Spatial Analysis
Show answer & explanation

Correct answer: C. Spatial Analysis and Object Detection

Azure Computer Vision's Spatial Analysis can track people's movements within defined zones, ideal for detecting entry/exit and dwell time. Object Detection can identify specific actions like picking up a product when trained on relevant images.

Why the other options are wrong

  • A. Read API is for text extraction; not suitable for behavior analysis.
  • B. Image Analysis (Tags) provides general image descriptions, not specific behavioral tracking, and Read API is for text.
  • D. Form Recognizer is for document processing; not relevant here.

Computer Vision Spatial Analysis + Object Detection

Combining Spatial Analysis for tracking people in spaces with Object Detection for identifying specific items or actions.

  • Spatial Analysis tracks movement in defined zones (e.g., enter, exit, dwell time).
  • Object Detection identifies objects within an image (e.g., a hand picking up a product).
  • Together they enable comprehensive behavior analysis in physical spaces.

Memory trick: Spatial Analysis watches where they go, Object Detection sees what they grab.

More Implement image and video processing solutions questions