Microsoft Azure AI Fundamentals (AI-900)Describe features of computer vision workloads on AzureHard

A company is developing an augmented reality (AR) application for interior design. The application needs to understand the layout of a room in real-time, identifying surfaces like walls, floors, and ceilings, and distinguishing them from furniture. This requires a detailed, pixel-by-pixel understanding of the scene to place virtual objects realistically. Which computer vision capability is most crucial for this level of scene understanding?

  1. AOptical Character Recognition (OCR)
  2. BObject Detection
  3. CSemantic Segmentation
  4. DImage Classification
Show answer & explanation

Correct answer: C. Semantic Segmentation

Semantic Segmentation is crucial because it provides a pixel-level understanding of the scene, classifying each pixel as belonging to a specific category like 'wall', 'floor', 'ceiling', or 'furniture'. This detailed understanding is essential for realistically placing virtual objects in an AR environment.

Why the other options are wrong

  • A. OCR is for text extraction and irrelevant to scene understanding for AR.
  • B. Object Detection would draw bounding boxes around furniture but wouldn't provide the detailed, pixel-level understanding of surfaces like walls and floors needed for realistic AR placement.
  • D. Image Classification would assign a single label to the entire room (e.g., 'living room'), which is too broad for AR object placement.

Computer Vision for AR/VR

Computer vision techniques used to enable augmented and virtual reality experiences, often involving real-time scene understanding.

  • Key capabilities include 3D reconstruction, object tracking, and semantic segmentation.
  • Enables realistic placement and interaction of virtual objects in real environments.
  • Requires high precision and real-time performance.

Memory trick: For AR, segment the scene semantically.

More Describe features of computer vision workloads on Azure questions