Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureHard
A company is developing an augmented reality (AR) application for interior design. The application needs to understand the layout of a room in real-time, identifying surfaces like walls, floors, and ceilings, and distinguishing them from furniture. This requires a detailed, pixel-by-pixel understanding of the scene to place virtual objects realistically. Which computer vision capability is most crucial for this level of scene understanding?
- AOptical Character Recognition (OCR)
- BObject Detection
- CSemantic Segmentation
- DImage Classification
Show answer & explanationAnswer & explanation
Correct answer: C. Semantic Segmentation
Semantic Segmentation is crucial because it provides a pixel-level understanding of the scene, classifying each pixel as belonging to a specific category like 'wall', 'floor', 'ceiling', or 'furniture'. This detailed understanding is essential for realistically placing virtual objects in an AR environment.
Why the other options are wrong
- A. OCR is for text extraction and irrelevant to scene understanding for AR.
- B. Object Detection would draw bounding boxes around furniture but wouldn't provide the detailed, pixel-level understanding of surfaces like walls and floors needed for realistic AR placement.
- D. Image Classification would assign a single label to the entire room (e.g., 'living room'), which is too broad for AR object placement.
Computer Vision for AR/VR
Computer vision techniques used to enable augmented and virtual reality experiences, often involving real-time scene understanding.
- Key capabilities include 3D reconstruction, object tracking, and semantic segmentation.
- Enables realistic placement and interaction of virtual objects in real environments.
- Requires high precision and real-time performance.
Memory trick: For AR, segment the scene semantically.