paper-with-me

홈 › Papers

Innovation: An Almost Characterization of Hallucination

2026-05-26 · Nishant P. Das, Piyush Srivastava arxiv

Hallucination is a central limitation of large language models (LLMs), and substantial effort has been devoted to understanding and mitigating it. Towards this, Kalai and Vempala (STOC 2024) introduced a probabilistic framework formalizing calibration and hallucination, and showed that, with high probability, calibrated LLMs hallucinate roughly at the rate of the "missing mass", a measure of how incomplete the training data is relative to its source. This raises two fundamental questions: (i) what property of a calibrated LLM makes hallucinations unavoidable? and (ii) can hallucinations be avoided by giving up calibration? We answer these questions by introducing a simpler property we call innovation that measures the tendency of a model to produce outputs outside the training data. We show that innovation is implied by the condition for hallucination identified by Kalai and Vempala, and, further, that it is an almost characterization of hallucination: hallucination implies innovation, and conversely, innovation implies hallucination with high probability. We also provide lower bounds on the hallucination rate based on the "innovation rate", and by relating innovation rate back to missing mass, we obtain new hallucination rate lower bounds based on missing mass that extend the results of Kalai and Vempala.

📄 PDF Abstract BibTeX arXiv:2605.26808

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations

2025-12-25 · Chengxu Yang, Jingling Yuan, Siqi Cai, Jiawei Jiang 외 arxiv

Hallucinations in large language models (LLMs) are commonly regarded as errors to be minimized. However, recent perspectives suggest that some hallucinations may encode creative or epistemically valuable content, a dimen…

Auction Design using Value Prediction with Hallucinations

2025-02-12 · Ilan Lobel, Humberto Moreira, Omar Mouchtaki

We investigate a Bayesian mechanism design problem where a seller seeks to maximize revenue by selling an indivisible good to one of n buyers, incorporating potentially unreliable predictions (signals) of buyers' private…

PredictionValue prediction

Global Positioning: the Uniqueness Question and a New Solution Method

2023-10-13 · Mireille Boutin, Gregor Kemper

We provide a new algebraic solution procedure for the global positioning problem in $n$ dimensions using $m$ satellites. We also give a geometric characterization of the situations in which the problem does not have a un…

All

HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems

2026-06-22 · Mateusz Barański, Jan Jasiński, Julitta Bartolewska, Marcin Witkowski 외 arxiv

End-to-end Automatic Speech Recognition (ASR) systems hallucinate on natural speech, yet existing mitigation methods are typically evaluated on non-speech or artificially corrupted audio. We introduce HALAS, the first hu…

Speech Recognition

From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models

2024-10-09 · Yuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang 외

Hallucinations in large vision-language models (LVLMs) are a significant challenge, i.e., generating objects that are not presented in the visual input, which impairs their reliability. Recent studies often attribute hal…

AttributeHallucination