AWS Certified AI PractitionerFoundation ModelsMedium

A startup is developing an application that uses a foundation model to summarize long technical reports. They are encountering issues where the model sometimes generates plausible-sounding but factually incorrect information, a phenomenon known as 'hallucination.' Which fundamental limitation of foundation models contributes most directly to this issue?

  1. ATheir inability to process complex mathematical equations accurately.
  2. BTheir exclusive reliance on supervised learning, preventing creative generation.
  3. CTheir small parameter count, limiting their knowledge capacity.
  4. DTheir training objective to predict the next token, not to verify factual accuracy.
Show answer & explanation

Correct answer: D. Their training objective to predict the next token, not to verify factual accuracy.

Foundation models, especially Large Language Models, are trained with an objective to predict the next most probable token based on their vast training data. This probabilistic generation does not inherently include a mechanism for factual verification, leading to instances where the model generates coherent but incorrect information (hallucinations).

Why the other options are wrong

  • A. While FMs might struggle with precise math, hallucination is more about factual incorrectness in text, not just mathematical error.
  • B. FMs are often trained using unsupervised or self-supervised methods on massive datasets, and they are highly capable of creative generation, which can sometimes manifest as hallucination.
  • C. FMs typically have very large parameter counts, which gives them immense knowledge, but doesn't prevent hallucination.

LLM Hallucination

The phenomenon where a Large Language Model generates information that is plausible-sounding but factually incorrect, nonsensical, or fabricated.

  • Arises from the probabilistic nature of next-token prediction.
  • A significant challenge in deploying LLMs for factual tasks.
  • Mitigated by techniques like Retrieval Augmented Generation (RAG).

Memory trick: Next-Token Prediction Powers Plausible Problems.

More Foundation Models questions