Microsoft Certified: Azure AI Engineer AssociatePlan and manage an Azure AI solutionMedium
A company is developing an Azure AI solution that includes a custom vision model trained in Azure Custom Vision. The model needs to be deployed to an IoT Edge device for real-time inference on images captured locally, with minimal latency and offline capabilities. Which deployment target should you choose for the Custom Vision model?
- AExport the model for a compact device (e.g., ONNX, TensorFlow Lite) and deploy via Azure IoT Edge.
- BDeploy the model as a web service to an Azure App Service.
- CExport the model as a Docker container and deploy to Azure Kubernetes Service (AKS).
- DDeploy the model to an Azure Machine Learning endpoint.
Show answer & explanationAnswer & explanation
Correct answer: A. Export the model for a compact device (e.g., ONNX, TensorFlow Lite) and deploy via Azure IoT Edge.
Azure Custom Vision allows exporting models in formats optimized for edge devices (like ONNX or TensorFlow Lite). Deploying these via Azure IoT Edge enables real-time inference on the device itself, providing minimal latency and offline capabilities, which directly addresses the requirements.
Why the other options are wrong
- B. Deploying to Azure App Service is for cloud-based web applications and doesn't support edge deployment or offline inference.
- C. Deploying to AKS is suitable for cloud-based, scalable web services but doesn't meet the 'IoT Edge device', 'minimal latency', or 'offline capabilities' requirements.
- D. Deploying to an Azure Machine Learning endpoint is for cloud-based inference and doesn't meet the requirements for edge deployment or offline capabilities.
Azure Custom Vision Export Options
Azure Custom Vision allows exporting trained models in various formats optimized for different deployment targets, including compact devices (ONNX, TensorFlow Lite) and Docker containers.
- Export for compact devices enables edge deployment.
- Docker container export is for cloud or container environments.
- Choice depends on inference location and latency requirements.
Memory trick: For 'Edge' and 'Offline', think 'IoT Edge' and 'Export Compact'.