Microsoft Certified: Azure AI Engineer AssociateImplement image and video processing solutionsHard

A cultural heritage organization is digitizing ancient manuscripts. These manuscripts contain text in multiple historical languages, some of which are no longer commonly spoken, and feature complex, ornate page layouts with columns, illustrations, and marginalia. The goal is to extract all text for archival and research purposes, preserving the original reading order and structural information. Which Azure Cognitive Service, with its specific capabilities, offers the best solution for this highly complex OCR challenge?

  1. AAzure Form Recognizer's prebuilt document models.
  2. BAzure Computer Vision's Read API with multi-language and layout understanding.
  3. CAzure Custom Vision trained on individual characters.
  4. DAzure Computer Vision's OCR (Legacy) API with language hints.
Show answer & explanation

Correct answer: B. Azure Computer Vision's Read API with multi-language and layout understanding.

Azure Computer Vision's Read API is specifically designed for robust and accurate OCR on complex documents. It excels at handling multiple languages (including less common ones), intricate layouts with columns and mixed content, and accurately preserving the reading order and structural information, making it ideal for ancient manuscripts.

Why the other options are wrong

  • A. Form Recognizer's prebuilt models are for common business documents; custom models would be needed, but Read API is still superior for arbitrary historical text and layouts.
  • C. Custom Vision is for object detection/classification, not OCR of complex text documents.
  • D. The legacy OCR API is less accurate, has limited language support, and struggles with complex layouts.

Read API Multi-language & Layout

The Azure Computer Vision Read API excels at extracting text from documents with multiple languages and complex layouts, accurately determining reading order and structural elements.

  • Supports over 100 languages for both print and handwriting.
  • Intelligently detects and preserves reading order.
  • Handles complex layouts like columns, tables, and mixed content.

Memory trick: When the text is old and wild, Read API keeps it mild (organized and understood).

More Implement image and video processing solutions questions