Microsoft Certified: Azure AI Engineer AssociateImplement generative AI solutionsHard

An AI engineer is developing a generative AI application that summarizes lengthy financial reports using Azure OpenAI Service. The reports can be very long, often exceeding the maximum input token limit of the LLM. The engineer needs a strategy to ensure that the entire report is processed and summarized effectively, without losing critical information due to truncation. Which approach is most appropriate for handling these long inputs?

  1. AUse a larger, more expensive model that has a slightly higher token limit.
  2. BInstruct the model to ignore less important sections of the report.
  3. CSplit the report into smaller chunks and summarize each chunk independently, then combine.
  4. DIncrease the `max_tokens` parameter for the model's output.
Show answer & explanation

Correct answer: C. Split the report into smaller chunks and summarize each chunk independently, then combine.

Splitting the long report into smaller, manageable chunks and summarizing each independently, then combining these summaries, is a common and effective strategy for handling inputs that exceed the LLM's token limit. This ensures all critical information is processed and contributes to the final summary.

Why the other options are wrong

  • A. While larger models might have slightly higher limits, they often don't solve the problem for extremely long documents and are more expensive. It's a temporary fix, not a robust strategy for arbitrary length.
  • B. Instructing the model to ignore sections risks losing critical information, which is precisely what needs to be avoided for financial reports.
  • D. `max_tokens` controls the output length, not the input length.

Long Document Processing (Chunking & Summarization)

A technique to handle documents exceeding an LLM's token limit by breaking them into smaller, overlapping chunks, summarizing or processing each chunk, and then combining or recursively summarizing the results.

  • Prevents information loss due to LLM context window constraints.
  • Enables processing of arbitrarily long texts.
  • Can involve 'map-reduce' or hierarchical summarization patterns.

Memory trick: Long text, chop it, map it, then reduce it to a summary.

More Implement generative AI solutions questions