paper-with-me

홈 › Papers

ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models

2026-08-19 · Jihae Jeong, Junha Choi, Hwanjo Yu arxiv

Large vision-language models (LVLMs) often hallucinate, generating content that the input image does not support. Preventing such content during decoding calls for a candidate-specific measure of how strongly the image supports the token under consideration. The model's visual-token states offer a natural source of this evidence because projecting each state through the output head reveals which vocabulary items that position favors. These position-wise readouts cannot be pooled directly because their probability magnitudes are not comparable across visual positions. Vocabulary ranks provide a scale-invariant basis for pooling, but tokens still differ systematically in their typical rank-based evidence. We propose ReWEIGH, a training-free decoding intervention that aggregates these ranks across visual positions and compares each candidate with a token-specific reference estimated from unlabeled images. At inference, ReWEIGH caches the image evidence during prefill and applies a bounded penalty only to candidates that fall below their reference. On four 7B backbones, ReWEIGH reduces hallucinated object mentions by up to 21.3% while largely preserving or improving descriptive and general performance. With evidence cached, the average added latency is 1.33% per token, and the reductions extend across six architecture families to 32B parameters.

📄 PDF Abstract BibTeX arXiv:2608.19075

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs

2026-08-21 · Haiming Li, Yingsheng Liu, Jingmin Zhu, Siyuan Yan 외 arxiv

Multimodal LLMs apply the language model interface to visual inputs, where ordinal regression tasks such as age estimation, image quality assessment, and disease grading require autoregressive decisions over ordered clas…

Image Quality AssessmentAge Estimation

VORD: Visual Ordinal Calibration for Mitigating Object Hallucinations in Large Vision-Language Models

2024-12-20 · Dexter Neo, Tsuhan Chen

Large Vision-Language Models (LVLMs) have made remarkable developments along with the recent surge of large language models. Despite their advancements, LVLMs have a tendency to generate plausible yet inaccurate or incon…

SkillSight: Calibrating Generic Content Bias for Skill Retrieval

2026-07-21 · Jinying Xiao, Bin Li, Xiaopeng Li, Jianling Li 외 arxiv

As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents…

LORE: Jointly Learning the Intrinsic Dimensionality and Relative Similarity Structure From Ordinal Data

2026-02-04 · Vivek Anand, Alec Helbling, Mark A. Davenport, Gordon J. Berman 외 arxiv

Learning the intrinsic dimensionality of subjective perceptual spaces such as taste, smell, or aesthetics from ordinal data is a challenging problem. We introduce LORE (Low Rank Ordinal Embedding), a scalable framework t…

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

2026-06-17 · Haipeng Luo, Qingfeng Sun, Songli Wu, Can Xu 외 arxiv

Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training. We …

Reinforcement Learning