Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsHard
A museum is digitizing its archive of ancient, fragile manuscripts. Many of these documents are in various historical scripts, contain mixed languages (e.g., Latin, Greek, Old English), and often have complex layouts with columns, annotations, and illustrations. They need to extract all textual content accurately, preserving the layout information and identifying script/language where possible. Which Azure Computer Vision API is the most appropriate choice?
- AAnalyze Image (Description)
- BOCR (Optical Character Recognition) API (Legacy)
- CForm Recognizer - Layout Model
- DRead API with multi-page and language support
Show answer & explanationAnswer & explanation
Correct answer: D. Read API with multi-page and language support
The Read API is the most advanced OCR in Azure Computer Vision, specifically designed to handle complex documents, including multiple languages, historical scripts, and intricate layouts, while preserving the structural information of the text. The legacy OCR API is inferior, and Form Recognizer is for structured forms, not free-form manuscripts.
Why the other options are wrong
- A. Analyze Image (Description) generates a high-level description of an image, not detailed text extraction.
- B. The legacy OCR API has limited language support and struggles with complex layouts and historical scripts.
- C. Form Recognizer's Layout model is for structured documents with forms, not free-form, historical manuscripts with highly variable content.
Computer Vision Read API Multi-language & Layout
An advanced OCR feature of Azure Computer Vision that accurately extracts text from images with mixed languages, complex layouts, and historical scripts, while preserving structural information.
- Handles printed, handwritten, and mixed text
- Supports a wide range of languages, including historical ones
- Extracts text with bounding box information for lines and words
- Preserves reading order and structural elements like paragraphs and pages
Memory trick: Read API 'reads' ancient scrolls, understanding every language and squiggle.