paper-with-me

홈 › Papers

Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models

2025-10-26 · Aya Nakayama, Brian Wong, Yuji Nishimura, Kaito Tanaka arxiv

The "style trap" poses a significant challenge for Large Vision-Language Models (LVLMs), hindering robust semantic understanding across diverse visual styles, especially in in-context learning (ICL). Existing methods often fail to effectively decouple style from content, hindering generalization. To address this, we propose the Semantic-Preserving Cross-Style Visual Reasoner (SP-CSVR), a novel framework for stable semantic understanding and adaptive cross-style visual reasoning. SP-CSVR integrates a Cross-Style Feature Encoder (CSFE) for style-content disentanglement, a Semantic-Aligned In-Context Decoder (SAICD) for efficient few-shot style adaptation, and an Adaptive Semantic Consistency Module (ASCM) employing multi-task contrastive learning to enforce cross-style semantic invariance. Extensive experiments on a challenging multi-style dataset demonstrate SP-CSVR's state-of-the-art performance across visual captioning, visual question answering, and in-context style adaptation. Comprehensive evaluations, including ablation studies and generalization analysis, confirm SP-CSVR's efficacy in enhancing robustness, generalization, and efficiency across diverse visual styles.

📄 PDF Abstract BibTeX arXiv:2510.22838

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringContrastive LearningVisual Reasoning

Similar Papers 제목 키워드 기반

CPST: Comprehension-Preserving Style Transfer for Multi-Modal Narratives

2023-12-14 · Yi-Chun Chen, Arnav Jhala

We investigate the challenges of style transfer in multi-modal visual narratives. Among static visual narratives such as comics and manga, there are distinct visual styles in terms of presentation. They include style fea…

Style Transfer

MAST: Mask-Guided Attention Mass Allocation for Training-Free Multi-Style Transfer

2026-04-14 · Dongkyung Kang, Jaeyeon Hwang, Junseo Park, Minji Kang 외 arxiv

Style transfer aims to render a content image with the visual characteristics of a reference style while preserving its underlying semantic layout and structural geometry. While recent diffusion-based models demonstrate …

Style Transfer

CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer

2026-02-16 · Wenbo Nie, Zixiang Li, Renshuai Tao, Bin Wu 외 arxiv

Transferring visual style between images while preserving semantic correspondence between similar objects remains a central challenge in computer vision. While existing methods have made great strides, most of them opera…

Semantic correspondenceStyle Transfer

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

2026-06-25 · Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto 외 arxiv

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually ground…

Continual Pretraining

EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning

2024-04-22 · Mingjie Ma, zhihuan yu, Yichao Ma, GuoHui Li

Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions requiring human commonsense, and to provide rationales explaining why the answers are correct. With emergence of Large…

Visual Commonsense Reasoning