paper-with-me

홈 › Papers

Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?

2025-08-28 · Yang Nan, Pengfei He, Ravi Tandon, Han Xu arxiv

Large language models (LLMs) have delivered significant breakthroughs across diverse domains but can still produce unreliable or misleading outputs, posing critical challenges for real-world applications. While many recent studies focus on quantifying model uncertainty, relatively little work has been devoted to \textit{diagnosing the source of uncertainty}. In this study, we show that, when an LLM is uncertain, the patterns of disagreement among its multiple generated responses contain rich clues about the underlying cause of uncertainty. To illustrate this point, we collect multiple responses from a target LLM and employ an auxiliary LLM to analyze their patterns of disagreement. The auxiliary model is tasked to reason about the likely source of uncertainty, such as whether it stems from ambiguity in the input question, a lack of relevant knowledge, or both. In cases involving knowledge gaps, the auxiliary model also identifies the specific missing facts or concepts contributing to the uncertainty. In our experiment, we validate our framework on AmbigQA, OpenBookQA, and MMLU-Pro, confirming its generality in diagnosing distinct uncertainty sources. Such diagnosis shows the potential for relevant manual interventions that improve LLM performance and reliability.

📄 PDF Abstract BibTeX arXiv:2509.04464

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Black-Box Hallucination Detection via Consistency Under the Uncertain Expression

2025-09-26 · Seongho Joo, Kyungmin Min, Jahyun Koo, Kyomin Jung arxiv

Despite the great advancement of Language modeling in recent days, Large Language Models (LLMs) such as GPT3 are notorious for generating non-factual responses, so-called "hallucination" problems. Existing methods for de…

Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation

2024-12-10 · Qinhong Lin, Linna Zhou, Zhongliang Yang, Yuang Cai

Large Language Models (LLMs) display formidable capabilities in generative tasks but also pose potential risks due to their tendency to generate hallucinatory responses. Uncertainty Quantification (UQ), the evaluation of…

Text GenerationUncertainty Quantification

LUQ: Long-text Uncertainty Quantification for LLMs

2024-03-29 · Caiqi Zhang, Fangyu Liu, Marco Basaldella, Nigel Collier

Large Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks. However, LLMs are also prone to generate nonfactual content. Uncertainty Quantification (UQ) is pivotal in enhancing our und…

Text GenerationUncertainty Quantification

Escaping the Mode Lottery: Multi-Response Training Improves Language Model Generalization

2026-05-30 · Hasan Amin, Kian Ahrabian, Ming Yin, Rajiv Khanna arxiv

Modern language-model fine-tuning typically pairs each prompt with a single response, even though many prompts admit multiple valid completions. This effectively reduces a multi-modal conditional distribution to a one-sa…

Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness

2023-08-30 · Jiuhai Chen, Jonas Mueller

We introduce BSDetector, a method for detecting bad and speculative answers from a pretrained Large Language Model by estimating a numeric confidence score for any output it generated. Our uncertainty quantification tech…

Language ModelingLanguage ModellingLarge Language ModelUncertainty Quantification