paper-with-me

홈 › Papers

Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding

2025-09-27 · Lin Long, Changdae Oh, Seongheon Park, Sharon Li arxiv

Large vision-language models (LVLMs) achieve strong performance on multimodal tasks, yet they often default to their language prior (LP) -- memorized textual patterns from pre-training while under-utilizing visual evidence. Prior analyses of LP mostly rely on input-output probing, which fails to reveal the internal mechanisms governing when and how vision influences model behavior. To address this gap, we present the first systematic analysis of language prior through the lens of chain-of-embedding, which examines the layer-wise representation dynamics within LVLMs. Our analysis reveals a universal phenomenon: each model exhibits a Visual Integration Point (VIP), a critical layer at which visual information begins to meaningfully reshape hidden representations and influence decoding for multimodal reasoning. Building on this observation, we introduce the Total Visual Integration (TVI) estimator, which aggregates representational discrepancy beyond the VIP to quantify how strongly visual query influences response generation. Across 60 model-dataset combinations spanning 10 contemporary LVLMs and 6 benchmarks, we demonstrate that VIP consistently emerges, and that TVI reliably predicts the strength of language prior. This offers a principled toolkit for diagnosing and understanding language prior in LVLMs.

📄 PDF Abstract BibTeX arXiv:2509.23050

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningResponse Generation

Similar Papers 제목 키워드 기반

DRIVINGVQA: Analyzing Visual Chain-of-Thought Reasoning of Vision Language Models in Real-World Scenarios with Driving Theory Tests

2025-01-08 · Charles Corbière, Simon Roburin, Syrielle Montariol, Antoine Bosselut 외

Large vision-language models (LVLMs) augment language models with visual understanding, enabling multimodal reasoning. However, due to the modality gap between textual and visual data, they often face significant challen…

Multimodal ReasoningMultiple-choiceVisual Reasoning

VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors

2025-12-07 · Wenbo Lyu, Yingjun Du, Jinglin Zhao, Xianton Zhen 외 arxiv

Understanding multi-image, multi-turn scenarios is a critical yet underexplored capability for Large Vision-Language Models (LVLMs). Existing benchmarks predominantly focus on static or horizontal comparisons -- e.g., sp…

Visual Reasoning

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

2025-02-15 · Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue 외

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-cont…

Question AnsweringStreaming video understandingVideo Understanding

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

2026-04-22 · Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr, Federico Tombari 외 arxiv

Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they still struggle with object hallucinations, i.e., the claim of nonexiste…

Response Generation

Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

2024-03-19 · Zuyan Liu, Yuhao Dong, Yongming Rao, Jie zhou 외

In the realm of vision-language understanding, the proficiency of models in interpreting and reasoning over visual content has become a cornerstone for numerous applications. However, it is challenging for the visual enc…

Instruction Followingvisual instruction followingVisual Question Answering