paper-with-me

홈 › Papers

Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

2025-02-20 · Yihao Xue, Kristjan Greenewald, Youssef Mroueh, Baharan Mirzasoleiman

Large Language Models (LLMs) suffer from hallucination problems, which hinder their reliability in sensitive applications. In the black-box setting, several self-consistency-based techniques have been proposed for hallucination detection. We empirically study these techniques and show that they achieve performance close to that of a supervised (still black-box) oracle, suggesting little room for improvement within this paradigm. To address this limitation, we explore cross-model consistency checking between the target model and an additional verifier LLM. With this extra information, we observe improved oracle performance compared to purely self-consistency-based methods. We then propose a budget-friendly, two-stage detection algorithm that calls the verifier model only for a subset of cases. It dynamically switches between self-consistency and cross-consistency based on an uncertainty interval of the self-consistency classifier. We provide a geometric interpretation of consistency-based hallucination detection methods through the lens of kernel mean embeddings, offering deeper theoretical insights. Extensive experiments show that this approach maintains high detection performance while significantly reducing computational cost.

📄 PDF Abstract BibTeX arXiv:2502.15845

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Similar Papers 제목 키워드 기반

Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction

2026-02-02 · Enes Altinisik, Masoomali Fatehkia, Fatih Deniz, Nadir Durrani 외 arxiv

Factual hallucination remains a central challenge for large language models (LLMs). Existing mitigation approaches primarily rely on either external post-hoc verification or mapping uncertainty directly to abstention dur…

EVADE: Evidence-Verified Agentic Diagnosis with Escape

2026-08-19 · Mohaimenul Azam Khan Raiaan, Nur Mohammad Fahad arxiv

Medical vision-language models (VLMs) can achieve high accuracy but remain unreliable: they are systematically overconfident, benefit little from test-time reasoning, and lack the ability to reliably calibrate trust in t…

Uncertainty-aware Unsupervised Multi-Object Tracking

2023-07-28 · ICCV 2023 1 · Kai Liu, Sheng Jin, Zhihang Fu, Ze Chen 외

Without manually annotated identities, unsupervised multi-object trackers are inferior to learning reliable feature embeddings. It causes the similarity-based inter-frame association stage also be error-prone, where an u…

Multi-Object TrackingObjectObject Tracking

How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures

2026-06-04 · Tanvi Thoria, Kiana Jafari, Marc R. Schlichting, Mykel J. Kochenderfer arxiv

Failures in language model reasoning emerge through distinct processes that leave identifiable signatures in the reasoning trace. We characterize these failures using token-level uncertainty signals, finding they arise t…

Scribble-Supervised Semantic Segmentation by Uncertainty Reduction on Neural Representation and Self-Supervision on Neural Eigenspace

2021-02-19 · ICCV 2021 10 · Zhiyi Pan, Peng Jiang, Yunhai Wang, Changhe Tu 외

Scribble-supervised semantic segmentation has gained much attention recently for its promising performance without high-quality annotations. Due to the lack of supervision, confident and consistent predictions are usuall…

SegmentationSemantic Segmentation