AWS Certified AI PractitionerAWS Services for AI/ML and Generative AIEasy
A financial services company processes millions of customer documents daily, including scanned forms and PDFs. They need to extract specific data fields, such as account numbers, names, and transaction amounts, from these documents for regulatory compliance and automated processing. The documents often have varying layouts and are sometimes handwritten. Which AWS service should they use?
- AAmazon Textract
- BAmazon Comprehend
- CAmazon Transcribe
- DAmazon Rekognition
Show answer & explanationAnswer & explanation
Correct answer: A. Amazon Textract
Amazon Textract is specifically designed for automatically extracting text, handwriting, and data from scanned documents and PDFs, including structured data from forms and tables, even with varying layouts and handwriting.
Why the other options are wrong
- B. Amazon Comprehend is for understanding natural language text, not for extracting data from document images.
- C. Amazon Transcribe converts speech to text, which is unrelated to document data extraction.
- D. Amazon Rekognition is for image and video analysis, not primarily for text and data extraction from documents.
Amazon Textract
A machine learning service that automatically extracts text, handwriting, and data from scanned documents, forms, and tables.
- Extracts text, forms, and tables.
- Supports various document types and layouts.
- Can process handwritten text.
Memory trick: Textract pulls out the exact text from documents.