Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsHard
A logistics company wants to automate the processing of shipping labels. These labels often have varying formats, contain both printed and handwritten text, and are sometimes partially obscured or wrinkled. They need to extract key information such as sender, recipient, tracking number, and weight. Which Azure AI service is the most effective for this complex OCR task?
- AAzure Form Recognizer's Custom Models
- BAzure Custom Vision
- CAzure Computer Vision's General OCR
- DAzure Computer Vision's Read API
Show answer & explanationAnswer & explanation
Correct answer: D. Azure Computer Vision's Read API
The Read API in Azure Computer Vision is specifically optimized for extracting text from complex images, including mixed printed and handwritten content, and handling various orientations and image quality issues, which aligns perfectly with the challenges of shipping labels. While Form Recognizer can do this, the Read API is a core Computer Vision offering for robust OCR on general documents, and often a foundational step.
Why the other options are wrong
- A. While Form Recognizer's custom models *could* be trained, the Read API is the underlying powerful OCR engine for general complex documents, and for shipping labels, it's often sufficient and more direct than building a custom Form Recognizer model unless the *structure* is highly consistent and needs key-value pairs specifically.
- B. Custom Vision is for object detection and image classification, not for advanced text extraction from documents.
- C. General OCR is older and less capable than the Read API for complex scenarios like mixed handwritten/printed text and image imperfections.
Azure Computer Vision Read API
A highly accurate and robust Optical Character Recognition (OCR) API designed for extracting text from diverse images and documents, including handwritten and mixed content.
- Supports both printed and handwritten text
- Handles varying orientations and image quality
- Provides text lines, words, and bounding box locations
- More advanced than general OCR for complex scenarios
Memory trick: Read API Reads All Rough Document Text.