Microsoft Certified: Azure AI Engineer AssociateImplement knowledge mining solutionsMedium
A publishing company has a large archive of digital books in PDF format. They are building an Azure AI Search solution to allow users to search the content of these books. Some books are very large, exceeding 100 pages, and the default extraction process is failing for these documents. You need to configure the AI Search indexer to successfully process these large PDFs. Which configuration setting should you adjust?
- AEnable `imageAction` for OCR processing.
- BIncrease the `maxPageCount` in the indexer's `parameters`.
- CSet the `parsingMode` to `json`.
- DSet the `documentRoot` to a specific JSON path.
Show answer & explanationAnswer & explanation
Correct answer: B. Increase the `maxPageCount` in the indexer's `parameters`.
The `maxPageCount` parameter within the indexer's `parameters` is used to specify the maximum number of pages to process for documents like PDFs. Increasing this value will allow the indexer to process larger multi-page documents that exceed the default limit.
Why the other options are wrong
- A. `imageAction` enables or disables OCR for images, which is not the primary issue for large PDF page counts.
- C. `parsingMode` to `json` is for parsing JSON files, not for handling large PDFs.
- D. `documentRoot` is used when extracting content from complex JSON structures, not for PDF page limits.
Azure AI Search Indexer maxPageCount
A configuration parameter for Azure AI Search indexers that specifies the maximum number of pages to extract from multi-page documents like PDFs.
- Default value is often low (e.g., 5 pages) to prevent excessive processing.
- Needs to be increased for large documents to ensure full content extraction.
- Configured within the `parameters` section of the indexer definition.
Memory trick: To process many pages, increase the max page count.