paper-with-me

Papers

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models

2025-05-26 · Jihoon Lee, Min Song

Despite significant advancements in Large Vision-Language Models, Object Hallucination (OH) remains a persistent challenge. Building upon prior studies on contrastive decoding that address this issue without requiring additional model training, we introduce RVCD (Retrieval Visual Contrastive Decoding), an advanced method to suppress OH. RVCD leverages both negative and positive images at the logit level, explicitly referencing AI-generated images designed to represent a single concept. Our approach demonstrates substantial improvements over existing decoding-based methods.

📄 PDF Abstract BibTeX arXiv:2505.20569

Code (1)

jihoonlee9898/rvcd 공식 구현 jax

Tasks

HallucinationObject HallucinationRetrieval

Similar Papers 제목 키워드 기반

Multi-Level Bidirectional Biomimetic Learning for EEG-Based Visual Decoding

2026-05-06 · Jingtao Liu, Peiliang Gong, Chuhang Zheng, Yiheng Liu 외 arxiv

EEG-based visual neural decoding aims to align neural responses with visual stimuli for tasks such as image retrieval. However, limited paired data and a fundamental mismatch between high-fidelity digital images and biol…

Representation LearningContrastive LearningImage Retrieval

Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

2023-11-28 · CVPR 2024 1 · Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li 외

Large Vision-Language Models (LVLMs) have advanced considerably, intertwining visual recognition and language understanding to generate content that is not only coherent but also contextually attuned. Despite their succe…

HallucinationObjectObject Hallucination

Spatial-Functional awareness Transformer-based graph archetype contrastive learning for Decoding Visual Neural Representations from EEG

2025-09-29 · Yueming Sun, Long Yang arxiv

Decoding visual neural representations from Electroencephalography (EEG) signals remains a formidable challenge due to their high-dimensional, noisy, and non-Euclidean nature. In this work, we propose a Spatial-Functiona…

Contrastive LearningBrain DecodingEeg Decoding

Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding

2025-02-03 · Chao Wang, Xuancheng Zhou, Weiwei Fu, Yang Zhou

Large Visual Language Models (LVLMs) integrate visual and linguistic modalities, exhibiting exceptional performance across various multimodal tasks. Nevertheless, LVLMs remain vulnerable to the issue of object hallucinat…

AttributeMMEObject

CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs

2026-05-22 · Xiaoyi Huang, Kejia Zhang, Zhiming Luo arxiv

Large Vision-Language Models have shown strong multimodal reasoning capabilities, yet they remain susceptible to object hallucinations when language priors dominate insufficient or misaligned visual evidence. Training-fr…

Multimodal Reasoning