paper-with-me

홈 › Papers

Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability

2026-04-19 · Lijie Zhou arxiv

Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-utilizing visual evidence -- a phenomenon termed ``text shortcut learning.'' We propose an adversarial evaluation framework that quantifies this cross-modal dependency by measuring accuracy degradation (Drop) when semantically conflicting text is paired with unchanged images. Four adversarial strategies -- shape\_swap, color\_swap, position\_swap, and random\_text -- are applied to a controlled geometric-shapes dataset ($n{=}1{,}000$). We compare three configurations: Baseline CLIP (ViT-B/32), LoRA fine-tuning, and LoRA Optimized (integrating Hard Negative Mining, Label Smoothing, layer-wise learning rates, Cosine Restarts, curriculum learning, and data augmentation). The optimized model reduces average Drop from 27.5\% to 9.8\% (64.4\% relative improvement, $p{<}0.001$) while maintaining 97\% normal accuracy. Attention visualization and embedding-space analysis confirm that the optimized model attends more to visual features and achieves tighter cross-modal alignment.

📄 PDF Abstract BibTeX arXiv:2604.17217

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection

2025-08-03 · Chengming Wang, Guodong Fan, Jinjiang Li, Min Gan 외 arxiv

With the advancement of remote sensing satellite technology and the rapid progress of deep learning, remote sensing change detection (RSCD) has become a key technique for regional monitoring. Traditional change detection…

Change Detection

Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

2021-10-20 · Reuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin 외

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompa…

Look at What I’m Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

2021-12-01 · NeurIPS 2021 12 · Reuben Tan, Bryan Plummer, Kate Saenko, Hailin Jin 외

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompa…

Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis

2022-04-17 · ACL 2022 5 · Yan Ling, Jianfei Yu, Rui Xia

As an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years. However, previous approaches either (i) use separately pre-trained visual…

Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)DecoderSentiment Analysis

Cross-Probe BERT for Efficient and Effective Cross-Modal Search

2021-01-01 · Tan Yu, Hongliang Fei, Ping Li

Inspired by the great success of BERT in NLP tasks, many text-vision BERT models emerged recently. Benefited from cross-modal attention, text-vision BERT models have achieved excellent performance in many language-visio…

Image RetrievalRetrieval