paper-with-me

Papers

Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity

2026-01-31 · Prakhar Ganesh, Reza Shokri, Golnoosh Farnadi arxiv

Large language models (LLMs) are known to "hallucinate" by generating false or misleading outputs. Hallucinations pose various harms, from erosion of trust to widespread misinformation. Existing hallucination evaluation, however, focuses only on correctness and often overlooks consistency, necessary to distinguish and address these harms. To bridge this gap, we introduce prompt multiplicity, a framework for quantifying consistency in LLM evaluations. Our analysis reveals significant multiplicity (over 50% inconsistency in benchmarks like Med-HALT), suggesting that hallucination-related harms have been severely misunderstood. Furthermore, we study the role of consistency in hallucination detection and mitigation. We find that: (a) detection techniques detect consistency, not correctness, and (b) mitigation techniques like RAG, while beneficial, can introduce additional inconsistencies. By integrating prompt multiplicity into hallucination evaluation, we provide an improved framework of potential harms and uncover critical limitations in current detection and mitigation strategies.

📄 PDF Abstract BibTeX arXiv:2602.00723

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Heaven-Sent or Hell-Bent? Benchmarking the Intelligence and Defectiveness of LLM Hallucinations

2025-12-25 · Chengxu Yang, Jingling Yuan, Siqi Cai, Jiawei Jiang 외 arxiv

Hallucinations in large language models (LLMs) are commonly regarded as errors to be minimized. However, recent perspectives suggest that some hallucinations may encode creative or epistemically valuable content, a dimen…

Transferring Visual Attributes from Natural Language to Verified Image Generation

2023-05-24 · Rodrigo Valerio, Joao Bordalo, Michal Yarom, Yonatan Bitton 외

Text to image generation methods (T2I) are widely popular in generating art and other creative artifacts. While visual hallucinations can be a positive factor in scenarios where creativity is appreciated, such artifacts …

Image GenerationText to Image GenerationText-to-Image GenerationVisual Question Answering (VQA)

Hallucinations Live in Variance

2026-01-11 · Aaron R. Flouro, Shawn P. Chadwick arxiv

Benchmarks measure whether a model is correct. They do not measure whether a model is reliable. This distinction is largely academic for single-shot inference, but becomes critical for agentic AI systems, where a single …

Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification

2024-07-02 · Pritish Sahu, Karan Sikka, Ajay Divakaran

Large Visual Language Models (LVLMs) struggle with hallucinations in visual instruction following task(s), limiting their trustworthiness and real-world applicability. We propose Pelican -- a novel framework designed to …

Claim VerificationHallucinationInstruction Followingvisual instruction following

Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs

2024-07-04 · Faisal Hamman, Pasan Dissanayake, Saumitra Mishra, Freddy Lecue 외

Fine-tuning LLMs on tabular classification tasks can lead to the phenomenon of fine-tuning multiplicity where equally well-performing models make conflicting predictions on the same input. Fine-tuning multiplicity can ar…

Decision Makingtabular-classification