paper-with-me

홈 › Papers

Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models

2025-09-16 · Jianfei Zhao, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan arxiv

Large Vision-Language Models (LVLMs) can accurately locate key objects in images, yet their attention to these objects tends to be very brief. Motivated by the hypothesis that sustained focus on key objects can improve LVLMs' visual capabilities, we propose Cross-Layer Vision Smoothing (CLVS). The core idea of CLVS is to incorporate a vision memory that smooths the attention distribution across layers. Specifically, we initialize this vision memory with position-unbiased visual attention in the first layer. In subsequent layers, the model's visual attention jointly considers the vision memory from previous layers, while the memory is updated iteratively, thereby maintaining smooth attention on key objects. Given that visual understanding primarily occurs in the early and middle layers of the model, we use uncertainty as an indicator of completed visual understanding and terminate the smoothing process accordingly. Experiments on four benchmarks across three LVLMs confirm the effectiveness and generalizability of our method. CLVS achieves state-of-the-art overall performance across a variety of visual understanding tasks and attains comparable results to the leading approaches on image captioning benchmarks.

📄 PDF Abstract BibTeX arXiv:2509.12897

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

Deep Graph Attention Networks

2024-10-21 · Jun Kato, Airi Mita, Keita Gobara, Akihiro Inokuchi

Graphs are useful for representing various realworld objects. However, graph neural networks (GNNs) tend to suffer from over-smoothing, where the representations of nodes of different classes become similar as the number…

Graph Attention

Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization

2026-06-29 · Haoming Meng, Anton Sugolov, Vardan Papyan arxiv

Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that emerge during training. Motivated by this observation, we introduce \emph{Depth-wis…

Image Classification

Weak Alignment Supervision from Hybrid Model Improves End-to-end ASR

2023-11-24 · Jintao Jiang, Yingbo Gao, Zoltan Tuske

In this paper, we aim to create weak alignment supervision from an existing hybrid system to aid the end-to-end modeling of automatic speech recognition. Towards this end, we use the existing hybrid ASR system to produce…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Characterizing visual cortical magnification with topological smoothing and optimal transportation

2024-04-09 · Yujian Xiong, Yanshuai Tu, Zhong-Lin Lu, Yalin Wang

Human vision has different concentration on visual fields. Cortical magnification factor (CMF) is a popular measurement on visual acuity and cortex concentration. In order to achieve thorough measurement of CMF across th…

Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability

2026-04-19 · Lijie Zhou arxiv

Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-utilizing visual evidence -- a phenomenon termed ``text shortcut learn…

Data Augmentation