paper-with-me

홈 › Papers

HTDC: Hesitation-Triggered Differential Calibration for Mitigating Hallucination in Large Vision-Language Models

2026-04-13 · Xinyun Liu arxiv

Large vision-language models (LVLMs) achieve strong multimodal performance, but still suffer from hallucinations caused by unstable visual grounding and over-reliance on language priors. Existing training-free decoding methods typically apply calibration at every decoding step, introducing unnecessary computation and potentially disrupting stable predictions. We address this problem by identifying layer-wise hesitation, a simple signal of grounding instability reflected by fluctuations in token preference across intermediate layers. Based on this observation, we propose Hesitation-Triggered Differential Calibration (HTDC), a training-free decoding framework that preserves standard full-branch inference and activates calibration only at hesitation-prone steps. When triggered, HTDC contrasts the full branch with two lightweight probes, a visual-nullification probe and a semantic-nullification probe, to suppress hallucination-prone candidates while avoiding unnecessary intervention on stable steps. Experiments on representative hallucination benchmarks show that HTDC consistently reduces hallucinations while maintaining strong task accuracy, achieving a favorable trade-off between effectiveness and computational overhead.

📄 PDF Abstract BibTeX arXiv:2604.12115

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Towards Interactive Annotation for Hesitation in Conversational Speech

2020-05-01 · LREC 2020 5 · Jane Wottawa, Marie Tahon, Apolline Marin, Nicolas Audibert

Manual annotation of speech corpora is expensive in both human resources and time. Furthermore, recognizing affects in spontaneous, non acted speech presents a challenge for humans and machines. The aim of the present st…

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

2026-08-27 · Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma arxiv

Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown obj…

Open World Object Detection

Synthesis of Event-triggered Controllers for SIRS Epidemic Models

2023-10-14 · Lichen Ding, Kazumune Hashimoto, Shigemasa Takai

In this paper, we investigate the problem of mitigating epidemics by applying an event-triggered control strategy. We consider a susceptible-infected-removed-susceptible (SIRS) model, which builds upon the foundational S…

HESITA(te) in Portuguese

2014-05-01 · LREC 2014 5 · C, Sara eias, Dirce Celorico, Jorge Proen{\c{c}}a 외

Hesitations, so-called disfluencies, are a characteristic of spontaneous speech, playing a primary role in its structure, reflecting aspects of the language production and the management of inter-communication. In this p…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Management+3

A Closer Look at the Calibration of Differentially Private Learners

2022-10-15 · HANLIN ZHANG, Xuechen Li, Prithviraj Sen, Salim Roukos 외

We systematically study the calibration of classifiers trained with differentially private stochastic gradient descent (DP-SGD) and observe miscalibration across a wide range of vision and language tasks. Our analysis id…