A company is digitizing its archive of historical documents, many of which are old, faded, and contain a mix of machine-printed text, handwritten notes, and even some stamps. The goal is to extract all textual content accurately, regardless of its format or condition, and preserve its spatial layout. Which Azure Cognitive Service feature provides the most robust solution for this challenging OCR task?
- AAzure Custom Vision's classification model
- BAzure Form Recognizer's Layout API
- CAzure Computer Vision's OCR (Legacy) API
- DAzure Computer Vision's Read API
Show answer & explanationAnswer & explanation
Correct answer: D. Azure Computer Vision's Read API
Azure Computer Vision's Read API (part of the Computer Vision service) is specifically designed for general-purpose OCR on text-heavy images and documents, handling both printed and handwritten text, mixed languages, and various image conditions (low resolution, faded). It also preserves the spatial layout and provides line- and word-level bounding boxes, making it superior to the legacy OCR API and more general than Form Recognizer's Layout API for arbitrary documents.
Why the other options are wrong
- A. Custom Vision is for training custom image classification/object detection, not OCR.
- B. Form Recognizer's Layout API is good for structured documents but less flexible for 'arbitrary historical documents' with varying layouts.
- C. The legacy OCR API is older, less accurate, and doesn't handle handwritten text or complex layouts as well as the Read API.
Computer Vision Read API
The Azure Computer Vision Read API is an advanced OCR capability designed for accurately extracting printed and handwritten text from images and documents, handling complex scenarios and preserving text structure.
- Handles printed and handwritten text.
- Robust across various image qualities (faded, skewed).
- Provides word and line bounding boxes, preserving layout.
Memory trick: When documents are ancient, Read API will find the words, however faint.